以下部分引用 AI 总结,现阶段 AI 仍然有幻觉。请以内容的原文为准。
数据新鲜度群聊 07-14 ✓卖方 07-14 ✓Wrap 断档15天 ✕News 07-15 ✓X 断档12天 ✕页面生成 07-15
PPO
7 atoms · 跨 1 天 · 首见 2026-07-01 · 最近 2026-07-01
三色: 🟦 fact 2 · 🟥 take 5 · stance ▲3/▼1/◆0
来源: X 7
时态: fresh:7
标签: 好观点:5 · 好数字:2 · 好信源:2 · 好思考:1
叙事 (narrative) (5)
- 2026-07-01 · 与 GRPO 算法上在 lambda=1 时非常相似
lambda=1 PPO is very similar to GRPO algorithmically
tags: 好观点·好思考 · → daily
- 2026-07-01 · 在短 horizon 下引入额外延迟
Short horizon PPo has extra latency has it has to wait for the critic
tags: 好观点 · → daily
- 2026-07-01 · 在长 horizon 下 advantage function 可以排序轨迹
Advantage function can rank at least and pick out ones that looked like they were going somewhere
tags: 好观点 · → daily
- 2026-07-01 · 在 tool calls/环境中 value function 通过子目标奖励提供优势
The moment there are few tool calls/ environment returning state I think value model's take over. They can smear reward in between the sub goals
tags: 好观点 · → daily
- 2026-07-01 · 在 MDP 状态变化时 value model 可以从不同轨迹共享知识
Tool calls or code exec update the state of the MDP. If you have many states then maybe you know how to value at least some of them. This is where the sharing of knowledge happens between trajectories
tags: 好观点 · → daily
事实 (fact) (2)
- 2026-07-01 · 短 horizon 定义
Short: 32K or less. Few pages long
tags: 好数字·好信源 · → daily
value: qty=32K or less · date=N/A
- 2026-07-01 · 长 horizon 定义
Long: 64/128K and higher, multiple rounds of tool calls or code execution
tags: 好数字·好信源 · → daily
value: qty=64/128K and higher · date=N/A
时间轴 (近 20)
- 2026-07-01 ·
narrative · 在短 horizon 下引入额外延迟 · → daily
- 2026-07-01 ·
narrative · 在长 horizon 下 advantage function 可以排序轨迹 · → daily
- 2026-07-01 ·
fact · 短 horizon 定义 · → daily
- 2026-07-01 ·
fact · 长 horizon 定义 · → daily
- 2026-07-01 ·
narrative · 在 tool calls/环境中 value function 通过子目标奖励提供优势 · → daily
- 2026-07-01 ·
narrative · 在 MDP 状态变化时 value model 可以从不同轨迹共享知识 · → daily
- 2026-07-01 ·
narrative · 与 GRPO 算法上在 lambda=1 时非常相似 · → daily
← 实体目录 · 系统日志