以下部分引用 AI 总结,现阶段 AI 仍然有幻觉。请以内容的原文为准。
数据新鲜度群聊 07-14 ✓卖方 07-14 ✓Wrap 断档15天 ✕News 07-15 ✓X 断档12天 ✕页面生成 07-15
Claude-Opus-4.8
4 atoms · 跨 2 天 · 首见 2026-05-28 · 最近 2026-05-29
三色: 🟦 fact 4 · 🟥 take 0 · stance ▲0/▼3/◆1
来源: X 4
时态: fresh:4
标签: 好数字:3 · 好思考:1
叙事 (narrative) (1)
- 2026-05-29 · 所有模型普遍过于乐观,低估剩余预算
2/ All models are universally too optimistic. Most of 20 model-task pairs underestimate remaining budget. Weaker models are MORE optimistic. The bias doesn't shrink with task progress.
tags: 好思考 · → daily
value: direction=过于乐观
事实 (fact) (3)
- 2026-05-29 · 任务成功率高于 Gemini 和 GPT-5.2
On SWE-bench: Opus leads task success, Gemini leads feasibility F₁, GPT-5.2 leads interval coverage. Three different winners, three different capabilities.
tags: 好数字 · → daily
value: direction=高于
- 2026-05-29 · 在消耗60%预算后,模型仍以>70%概率预测可行
Models predict 'feasible' above 70% even after 60% of budget is consumed. The alarm fires only in the final 20%.
tags: 好数字 · → daily
value: qty=>70% · date=消耗60%预算后
- 2026-05-28 · medium 模型物理渲染能力
如果使用 medium, 是无法完成这个测试的, 写的 shader 有问题直接炸了
tags: 好数字 · → daily
value: qty=失败 · date=2026-05-28
时间轴 (近 20)
- 2026-05-29 ·
fact · 任务成功率高于 Gemini 和 GPT-5.2 · → daily
- 2026-05-29 ·
narrative · 所有模型普遍过于乐观,低估剩余预算 · → daily
- 2026-05-29 ·
fact · 在消耗60%预算后,模型仍以>70%概率预测可行 · → daily
- 2026-05-28 ·
fact · medium 模型物理渲染能力 · → daily
← 实体目录 · 系统日志