以下部分引用 AI 总结,现阶段 AI 仍然有幻觉。请以内容的原文为准。
数据新鲜度群聊 07-14 ✓卖方 07-14 ✓Wrap 断档15天 ✕News 07-15 ✓X 断档12天 ✕页面生成 07-15
METR graph
6 atoms · 跨 2 天 · 首见 2026-05-09 · 最近 2026-05-22
三色: 🟦 fact 0 · 🟥 take 6 · stance ▲0/▼4/◆2
来源: X 6
时态: aging:4 · fresh:2
标签: 好观点:6
叙事 (narrative) (6)
- 2026-05-22 · 知名度
none of them had heard of the METR graph
tags: 好观点 · → daily
- 2026-05-22 · 知名度
you have to show it to people for them to recognize it, they don’t know it as “the METR graph”
tags: 好观点 · → daily
- 2026-05-09
[老化中] · 衡量标准:50% 成功率,非100%
it is about achieving *50%* success. Not 100 or 99 or even 90. The key problem with GenAI has been reliability; this graph does not address reliable performance. At all.
tags: 好观点 · → daily
value: qty=50% · direction=low
- 2026-05-09
[老化中] · 仅限于软件任务,非通用智能
If you read carefully, it is only about software tasks. Not general intelligence.
tags: 好观点 · → daily
- 2026-05-09
[老化中] · 未证明人类16小时可完成的任务大部分能被可靠完成
It certainly doesn’t tell you that *most* (let alone) all things that humans can do in 16 hours can be done in Mythos, let alone reliably
tags: 好观点 · → daily
- 2026-05-09
[老化中] · 对神经符号AI的验证,非LLM可无限扩展的证明
this a vindication of neurosymbolic AI – but not a proof that LLMs themselves can be perpetually scaled. As such it’s not a proof that another trillion dollars will continue the graph.
tags: 好观点 · → daily
时间轴 (近 20)
- 2026-05-22 ·
narrative · 知名度 · → daily
- 2026-05-22 ·
narrative · 知名度 · → daily
- 2026-05-09 ·
narrative · 衡量标准:50% 成功率,非100% · → daily
- 2026-05-09 ·
narrative · 仅限于软件任务,非通用智能 · → daily
- 2026-05-09 ·
narrative · 未证明人类16小时可完成的任务大部分能被可靠完成 · → daily
- 2026-05-09 ·
narrative · 对神经符号AI的验证,非LLM可无限扩展的证明 · → daily
← 实体目录 · 系统日志