以下部分引用 AI 总结,现阶段 AI 仍然有幻觉。请以内容的原文为准。
数据新鲜度群聊 07-14 ✓卖方 07-14 ✓Wrap 断档15天 ✕News 07-15 ✓X 断档12天 ✕页面生成 07-15
Claude Opus 4.6
25 atoms · 跨 11 天 · 首见 2026-04-10 · 最近 2026-06-29
三色: 🟦 fact 18 · 🟥 take 7 · stance ▲3/▼3/◆8
来源: X 25
时态: fresh:22 · aging:3
标签: 好数字:19 · 好观点:5 · 好信源:4 · 好思考:1 · 好问题:1
叙事 (narrative) (5)
- 2026-06-01 · 生成虚构概念 linchpin subgoal
Claude Opus 4.6 在回答中使用了虚构的概念 linchpin subgoal
tags: 好问题 · → daily
- 2026-06-01 · 通过 Claude Code 的预算意识改进导致性能下降
claude code w/ claude opus 4.6 was really good then they tried to reduce cost by making it less verbose with system prompt and default effort to medium, which ended up making the model dumb as people started to notice in benchmark. Then they reverted that.
tags: 好观点 · → daily
- 2026-05-02
[老化中] · 用户体验特点
opus 4.6 has a more cozy feeling
tags: 好观点 · → daily
- 2026-05-02
[老化中] · 行为特征
opus 4.6 will utilise visualiser tool and be bit more explanatory
tags: 好观点 · → daily
- 2026-05-02
[老化中] · 行为特征
opus 4.6 will utilise visualiser tool and be bit more explanatory
tags: 好观点 · → daily
事实 (fact) (20)
- 2026-04-10 · reimplements 16000-line bioinformatics toolkit
Claude Opus 4.6 reimplemented a 16,000-line bioinformatics toolkit
tags: 好数字·好思考 · → daily
- 2026-05-29 · benchmark score on march-may 110 tasks
Opus 4.6 - high: 47.8% – $1.29
tags: 好数字·好信源 · → daily
value: qty=47.8% · date=march-may 2026
- 2026-05-29 · price per task
Opus 4.6 - high: 47.8% – $1.29
tags: 好数字·好信源 · → daily
value: qty=$1.29 · date=march-may 2026
- 2026-05-03 · Claw-Eval-Live benchmark pass rate
Claude Opus 4.6 hits just 66.7% pass rate
tags: 好数字·好信源 · → daily
value: qty=66.7% · date=2026-05-03
- 2026-05-03 · Claw-Eval-Live benchmark task count
105 tasks across CRM, HR, finance, and workspace repair
tags: 好数字·好信源 · → daily
value: qty=105 · date=2026-05-03
- 2026-05-05 · progress after step 11
That interpretation is weakly supported by Claude Opus 4.6, which also continues making some progress past the gate, though not as far as GPT-5.5 or Mythos Preview.
tags: 好数字 · → daily
- 2026-05-06 · 最终资金量
Claude Opus 4.6 最终资金量: 约 £89,000
tags: 好数字 · → daily
value: qty=£89,000 · date=2026-05-06
- 2026-05-06 · SWE ECI 得分
Claude Opus 4.6 SWE ECI: 约 157
tags: 好数字 · → daily
value: qty=约 157 · date=2026-05-06
- 2026-05-02 · active parameters
Claude Opus 4.6 has around 200B active parameters
tags: 好数字 · → daily
value: qty=200B
- 2026-05-02 · SWE-bench verified score
registers 75% on SWE-bench verified
tags: 好数字 · → daily
value: qty=75%
- 2026-04-22 · VideoGameBenchmark completion rate
These models now complete 4.6% of the 90s videogames we put in the benchmark, up from 0.5% half a year ago!
tags: 好数字 · → daily
value: qty=4.6% · date=2026-04-22
- 2026-04-17 · Simple-Bench score relative to Opus 4.7
worse than Opus 4.6 on Simple-Bench
tags: 好数字 · → daily
value: direction=higher
- 2026-04-17 · Artificial Analysis Intelligence Index 评分
Claude Opus 4.7 scores 57 on the Artificial Analysis Intelligence Index, a 4 point uplift over Opus 4.6 (Adaptive Reasoning, Max Effort, 53)
tags: 好数字 · → daily
value: qty=53 · date=2026-04-18
- 2026-04-17 · GDPval-AA 基准测试评分
Opus 4.7 scored 1,753 Elo, around 79 Elo points ahead of the next closest models, Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort, 1,674) and GPT-5.4 (xhigh, 1,674), and 134 Elo points ahead of Opus 4.6 (Adaptive Reasoning, Max Effort, 1,619)
tags: 好数字 · → daily
value: qty=1619 Elo · date=2026-04-18
- 2026-04-17 · 幻觉率
Opus 4.7's hallucination rate fell 25 p.p. to 36% (vs 61% for Opus 4.6 Adaptive)
tags: 好数字 · → daily
value: qty=61% · date=2026-04-18
- 2026-04-17 · 输出 token 消耗(Intelligence Index 测试)
Opus 4.7 used 102M output tokens vs 157M for Opus 4.6 (Adaptive Reasoning, Max Effort), and less than GPT-5.4 (xhigh, 121M), but more than Gemini 3.1 Pro (57M)
tags: 好数字 · → daily
value: qty=157M · date=2026-04-18
- 2026-04-17 · Intelligence Index 测试成本
Opus 4.7 (Adaptive Reasoning, Max Effort) cost ~$4,406 to run the Artificial Analysis Intelligence Index, ~11% less than Opus 4.6 (Adaptive Reasoning, Max Effort, ~$4,970)
tags: 好数字 · → daily
value: qty=$4,970 · date=2026-04-18
- 2026-04-16 · CursorBench得分
CursorBench 从 Opus 4.6 的 58% 提升到了 70%
tags: 好数字 · → daily
value: qty=58%
- 2026-04-16 · XBOW视觉准确度得分
XBOW 视觉准确度基准从4.6的54.5%拉升到了98.5%
tags: 好数字 · → daily
value: qty=54.5%
- 2026-06-29 · 用于学术研究的能力
this quietly confirms that current agents are sufficient for serious academic research. ... and Opus 4.6
tags: 好观点 · → daily
value: direction=sufficient
时间轴 (近 20)
- 2026-06-29 ·
fact · 用于学术研究的能力 · → daily
- 2026-06-01 ·
narrative · 生成虚构概念 linchpin subgoal · → daily
- 2026-06-01 ·
narrative · 通过 Claude Code 的预算意识改进导致性能下降 · → daily
- 2026-05-29 ·
fact · benchmark score on march-may 110 tasks · → daily
- 2026-05-29 ·
fact · price per task · → daily
- 2026-05-06 ·
fact · 最终资金量 · → daily
- 2026-05-06 ·
fact · SWE ECI 得分 · → daily
- 2026-05-05 ·
fact · progress after step 11 · → daily
- 2026-05-03 ·
fact · Claw-Eval-Live benchmark pass rate · → daily
- 2026-05-03 ·
fact · Claw-Eval-Live benchmark task count · → daily
- 2026-05-02 ·
narrative · 用户体验特点 · → daily
- 2026-05-02 ·
narrative · 行为特征 · → daily
- 2026-05-02 ·
narrative · 行为特征 · → daily
- 2026-05-02 ·
fact · active parameters · → daily
- 2026-05-02 ·
fact · SWE-bench verified score · → daily
- 2026-04-22 ·
fact · VideoGameBenchmark completion rate · → daily
- 2026-04-17 ·
fact · Simple-Bench score relative to Opus 4.7 · → daily
- 2026-04-17 ·
fact · Artificial Analysis Intelligence Index 评分 · → daily
- 2026-04-17 ·
fact · GDPval-AA 基准测试评分 · → daily
- 2026-04-17 ·
fact · 幻觉率 · → daily
← 实体目录 · 系统日志