以下部分引用 AI 总结,现阶段 AI 仍然有幻觉。请以内容的原文为准。
数据新鲜度群聊 07-14 ✓卖方 07-14 ✓Wrap 断档15天 ✕News 07-15 ✓X 断档12天 ✕页面生成 07-15
ZhihuFrontier
X · 19 atoms · 2026-06-24 → 2026-06-26 · 信用档先验 cred3
🟦 fact 8 · 🟥 take 11 · stance ▲3/▼5/◆3
覆盖实体 (top 10)
话题分布 (top 8)
训练方法 (2) · 算力成本 (2) · 推理能力 (2) · 技术路线 (1) · 训练算法 (1) · 推理模型 (1) · 算法适用性 (1) · 短任务验证 (1)
样例 atoms
- 2026-06-24 · GLM-5.2 · GLM-5.2 dropping GRPO does not mean GRPO is 'bad.' It means the assumptions that made GRPO attractive for short LLM RL tasks may no longer hold for long-horizon agentic tasks.
- 2026-06-25 · GLM-5.2 · GLM-5.2 moving away from GRPO is more like a practical correction than a rejection of GRPO itself
- 2026-06-25 · GRPO · GRPO is still a good algorithm for short, verifiable tasks
- 2026-06-25 · GRPO · GRPO has limits in long-horizon tasks
- 2026-06-25 · Zhipu · Zhipu brings the critic network back, so token-level advantage can be measured more carefully
← Author 索引 · 实体目录