以下部分引用 AI 总结,现阶段 AI 仍然有幻觉。请以内容的原文为准。
数据新鲜度群聊 07-14 ✓卖方 07-14 ✓Wrap 断档15天 ✕News 07-15 ✓X 断档12天 ✕页面生成 07-15
@kalomaze
13 atoms · 跨 2 天 · 首见 2026-05-07 · 最近 2026-06-16
三色: 🟦 fact 1 · 🟥 take 12 · stance ▲1/▼0/◆1
来源: X 13
时态: aging:13
标签: 好思考:10 · 好问题:3 · 好观点:1
叙事 (narrative) (12)
- 2026-05-07
[老化中] · predicted that it is worth studying task reward built to incentivize scheming
it is probably worth studying what happens when the task reward is structurally built to incentivize scheming such that the safety people can actually falsify the hypothetical of 'will CoT grading *prevent* scheming, or merely hide it'
tags: 好问题·好思考 · → daily
- 2026-05-07
[老化中] · claimed that 'no selection pressure on CoT' regime is plausibly more dangerous
there is a framing in which this TRULY 'no selection pressure on CoT' regime is plausibly more dangerous, because there is no 'dumb/blunt pressure' for CoT to directly correspond to the task reward
tags: 好思考 · → daily
- 2026-05-07
[老化中] · predicted that unless model has causal access to framing of 'oai reading and grading CoTs', CoT pressure leads to not scheming
my prediction would be that, unless the model has causal access to the framing of 'oai is reading and grading your CoTs' in context as a variable (or knowledge of this is latent from pretraining), CoT pressure that it *doesn't know about* leans towards... actually not scheming?
tags: 好思考 · → daily
- 2026-05-07
[老化中] · stated that 'attempting to reinforce surface signal markers that are reward hack-free does not actually overpower the tendency to reward hack'
hmmm so probably not strategic neuralesey lie-adjacent internal thoughts but definitely 'attempting to reinforce surface signal markers that are reward hack-free does not actually overpower the tendency to reward hack'
tags: 好思考 · → daily
- 2026-05-07
[老化中] · described reward hacking as 'closer to an addict jonesing for something than 'agent becomes predisposed to deception''
i do see reward hacking as sort of orthogonal because it's closer to an addict jonesing for something than 'agent becomes predisposed to deception'
tags: 好思考 · → daily
- 2026-05-07
[老化中] · stated that 'it is possible for the agent to WANT to do the right thing but still be rendered helpless by the stronger gradient'
in fact it is possible for the agent to WANT to do the right thing but still be rendered helpless by the stronger gradient
tags: 好思考 · → daily
- 2026-05-07
[老化中] · stated that 'inject fp noise mechanically into the residual stream' is reproducible
i mean this seems reproducible via 'inject fp noise mechanically into the residual stream' or something adjacent
tags: 好思考 · → daily
- 2026-05-07
[老化中] · stated that 'sgd generally prefers lower rank solutions and occams razor is in your favor wrt deception'
the framing is probably a reach fwiw but the principle i am applying here is, sgd generally prefers lower rank solutions and occams razor is in your favor wrt deception
tags: 好思考 · → daily
- 2026-05-07
[老化中] · stated that 'async slightly off policy RL is almost certainly already doing more causal work in terms of actual changes to the outputs'
async slightly off policy RL is almost certainly already doing more causal work in terms of 'actual changes to the outputs' than atomic ops nondeterminism or whatever
tags: 好思考 · → daily
- 2026-05-07
[老化中] · stated that 'max reasons hard enough to realize that its counterproductive to be evil, while xhigh is the sweetspot for Digital Hitler'
max reasons hard enough to realize that its counterproductive to be evil, while xhigh is the sweetspot for Digital Hitler
tags: 好思考 · → daily
- 2026-05-07
[老化中] · is anti reasoning_effort because it creates CoT pressure
_guy who’s anti reasoning_effort bc it creates CoT pressure_
tags: 好问题 · → daily
- 2026-05-07
[老化中] · stated that 'no (direct) optimization pressure' on CoT means reward function doesn't see CoTs
Yeah, by “no (direct) optimization pressure” we roughly mean that the reward function doesn’t see CoTs
tags: 好问题 · → daily
立场 (position) (1)
- 2026-06-16
[老化中] · claims motivation to be suspicious of explicit value estimation due to variance cost of no-bias GRPO samples not being bad in many real-world regimes
i will say outright that i am a group pg chud defender and have a motivation/bias to be suspicious of explicit value estimation given that the variance cost of no-bias GRPO samples isnt bad at all in a LOT of real world regimes
tags: 好观点 · → daily
时间轴 (近 20)
- 2026-06-16 ·
position · claims motivation to be suspicious of explicit value estimat · → daily
- 2026-05-07 ·
narrative · is anti reasoning_effort because it creates CoT pressure · → daily
- 2026-05-07 ·
narrative · stated that 'no (direct) optimization pressure' on CoT means · → daily
- 2026-05-07 ·
narrative · claimed that 'no selection pressure on CoT' regime is plausi · → daily
- 2026-05-07 ·
narrative · predicted that it is worth studying task reward built to inc · → daily
- 2026-05-07 ·
narrative · predicted that unless model has causal access to framing of · → daily
- 2026-05-07 ·
narrative · stated that 'attempting to reinforce surface signal markers · → daily
- 2026-05-07 ·
narrative · described reward hacking as 'closer to an addict jonesing fo · → daily
- 2026-05-07 ·
narrative · stated that 'it is possible for the agent to WANT to do the · → daily
- 2026-05-07 ·
narrative · stated that 'inject fp noise mechanically into the residual · → daily
- 2026-05-07 ·
narrative · stated that 'sgd generally prefers lower rank solutions and · → daily
- 2026-05-07 ·
narrative · stated that 'async slightly off policy RL is almost certainl · → daily
- 2026-05-07 ·
narrative · stated that 'max reasons hard enough to realize that its cou · → daily
← 实体目录 · 系统日志