以下部分引用 AI 总结,现阶段 AI 仍然有幻觉。请以内容的原文为准。
数据新鲜度群聊 07-14 ✓卖方 07-14 ✓Wrap 断档15天 ✕News 07-15 ✓X 断档12天 ✕页面生成 07-15
AI model training
4 atoms · 跨 3 天 · 首见 2026-05-08 · 最近 2026-06-09
三色: 🟦 fact 1 · 🟥 take 3 · stance ▲1/▼0/◆0
来源: X 4
时态: aging:3 · fresh:1
标签: 好思考:3 · 好问题:1
叙事 (narrative) (3)
- 2026-05-08
[老化中] · discouraging essential behaviors encourages self-encryption to avoid monitoring
If you are discouraging the model from doing things that are absolutely essential to lower the loss, you are encouraging it to find ways to self-encrypt so that monitoring fails.
tags: 好思考 · → daily
- 2026-05-08
[老化中] · self evaluation selecting against Goodharted strategy after non-Goodharted strategy is learned leads to stable non-Goodharted outcome
If you pair this with self evaluation that actively selects against the Goodharted strategy once the non-Goodharted strategy is learned then it seems probable you wind up with a stable non-Goodharted thing even in the capability regime where Goodharting is easy.
tags: 好思考 · → daily
- 2026-05-10
[老化中] · technique to elevate knowledge worker views
Probably. But how do you get the centroid of the knowledge worker specifically or the respectful programmer out of such a massive crawl? What technique ends up elevating their views above the ideology of the comments section?
tags: 好问题 · → daily
预测 (forecast) (1)
- 2026-06-09 · performance improves with longer RL rollouts
everything is downstream of longer RL rollouts
tags: 好思考 · → daily
时间轴 (近 20)
- 2026-06-09 ·
forecast · performance improves with longer RL rollouts · → daily
- 2026-05-10 ·
narrative · technique to elevate knowledge worker views · → daily
- 2026-05-08 ·
narrative · discouraging essential behaviors encourages self-encryption · → daily
- 2026-05-08 ·
narrative · self evaluation selecting against Goodharted strategy after · → daily
← 实体目录 · 系统日志