以下部分引用 AI 总结,现阶段 AI 仍然有幻觉。请以内容的原文为准。
数据新鲜度群聊 07-14 ✓卖方 07-14 ✓Wrap 断档15天 ✕News 07-15 ✓X 断档12天 ✕页面生成 07-15
RLVR
5 atoms · 跨 4 天 · 首见 2026-05-07 · 最近 2026-06-22
三色: 🟦 fact 0 · 🟥 take 5 · stance ▲1/▼1/◆3
来源: X 5
时态: fresh:4 · aging:1
标签: 好思考:3 · 好观点:2
叙事 (narrative) (5)
- 2026-06-22 · can encourage novel ideas given a good goal
rlvr (verifiable rewards) can encourage novel ideas given a good goal
tags: 好思考 · → daily
- 2026-05-24 · 影响方向
more RLVR will further warp agents around task-completion
tags: 好思考 · → daily
value: direction=正向
- 2026-05-21 · 训练产生低秩更新
it was known for awhile that RLVR produces low rank updates
tags: 好思考 · → daily
- 2026-06-22 · may not generalize capability to other domains without verifiable rewards
i am not sure if this can produce a model that generalize this capability to other domains without verifiable rewards
tags: 好观点 · → daily
- 2026-05-07
[老化中] · RLVR is not indefinitely improving and loses entropy
true but frankly this shows we're still bad at RLVR and lose entropy. in theory RL should be infinitely improving, and pass@k should be going up for high values of k
tags: 好观点 · → daily
时间轴 (近 20)
- 2026-06-22 ·
narrative · can encourage novel ideas given a good goal · → daily
- 2026-06-22 ·
narrative · may not generalize capability to other domains without verif · → daily
- 2026-05-24 ·
narrative · 影响方向 · → daily
- 2026-05-21 ·
narrative · 训练产生低秩更新 · → daily
- 2026-05-07 ·
narrative · RLVR is not indefinitely improving and loses entropy · → daily
← 实体目录 · 系统日志