以下部分引用 AI 总结,现阶段 AI 仍然有幻觉。请以内容的原文为准。
数据新鲜度群聊 07-14 ✓卖方 07-14 ✓Wrap 断档15天 ✕News 07-15 ✓X 断档12天 ✕页面生成 07-15
Opus 4.8
118 atoms · 跨 27 天 · 首见 2026-05-28 · 最近 2026-07-02
三色: 🟦 fact 69 · 🟥 take 49 · stance ▲48/▼19/◆26
叙事状态: 🍂 fading (退潮) · 💤 沉寂 · 策展近14d 0 atom · 🐦 X 近14d 16
状态只算低频策展源 (群/卖方); X firehose 仅作背景音量
来源: 群 1 · X 117
时态: fresh:116 · stale:2
标签: 好数字:63 · 好观点:42 · 好信源:13 · 好思考:7 · 好问题:1
别名 (合并): opus 4.8
🟨 AI 综合 · junior analyst 概览
展开 AI 综合 (灰色 · 非市场结论 · 点击数字溯源到原 atom)
model: deepseek-chat · 2026-07-12 · 默认折叠
Opus 4.8 在基准测试中碾压 GPT-5.5,编程能力排名第一并成为首个完全解决2个ProgramBench任务模型
¹¹¹。GDPval-AA 得分较 Opus 4.7 猛增 +137 分,领先 GPT-5.5 xhigh 达 +121 分
¹¹。但它在 DeepSWE 上被 GPT-5.5 在分数、时间和Token三方面全面超越,指令遵循仍差于 GPT-5.5 xhigh
¹¹。与 Opus 4.7 性能比较虽在误差范围内却得分更低,且被用户反馈“感觉不同而且经常令人恶心”
¹¹。
🧭 拥挤度 (一人一票): ▲ 14 位作者 (KOL14) + 群 1 条 vs ▼ 10 位作者 (KOL10)
⚖️ 多头 14/14 来自KOL
💢 核心分歧 (5 轴)
1. 编码与数学推理能力突破 vs 仍有不足
*topic: 模型性能/技术路线/模型对比/产品发布/模型能力 · 20 bull vs 3 bear*
🟢 bullish 侧:
- <a id="atom-7bfca5d71fb1593f"></a>2026-05-28
X·@scaling01 · 百万token长上下文表现 · 技术路线/模型发布 +同日1条
Opus 4.8@1 million now almost as good as GPT-5.5's 256K score
tags: 好观点 · → daily
value: qty=几乎等同于GPT-5.5的256K得分
📷 原图
- <a id="atom-c84b43d5adeb4cc0"></a>2026-05-28
X·@eliebakouch · benchmark performance compared to Mythos and previous generation · 技术路线/模型发布
opus 4.8 benchmarks vs mythos and previous generation graphwalks (long context), USAMO (math) are the biggest improvements. vending bench score is insanely bad
tags: 好数字 · → daily
value: qty=improved · direction=up
📷 原图
- <a id="atom-d5162af40e6358d1"></a>2026-05-28
X·@MetacriticCap · 是否为近期首个智能模型 · 模型发布/技术路线
Opus 4.8 is the first smart model in a long while
tags: 好观点 · → daily
- <a id="atom-61646cf50f51a96e"></a>2026-05-28
X·@emollick · 完整论文撰写流程 · 技术路线
Opus 4.8 formulated the hypotheses in advance, conducting data cleaning, did research on references, conducted analyses, did robustness checks, and put out the whole paper in LaTEX style.
tags: 好观点 · → daily
- <a id="atom-ff7ebc85994cae6f"></a>2026-05-28
X·@AIadventure3 · benchmark performance vs GPT 5.5 · 技术路线
But 4.8 blows gpt 5.5 out of the water in benchmarks
tags: 好数字 · → daily
value: direction=better
🔴 bearish 侧:
- <a id="atom-cd5ea978f4da7eaa"></a>2026-05-28
X·@adonis_singh · 视觉能力状态 · 技术路线/产品发布 +同日1条
opus 4.8 is still very much blind, they have not been focusing on vision and multimodality at all it seems
tags: 好观点 · → daily
📷 原图
- <a id="atom-d32fb22642557684"></a>2026-05-30
X·@scaling01 · DeepSWE 表现被 GPT-5.5 全面超越 · 模型发布/技术路线
Opus 4.8 gets score-, time- and token-mogged by GPT-5.5 on DeepSWE
tags: 好观点 · → daily
value: direction=被超越
📷 原图
展开 10 条中性
- <a id="atom-f92919cc8e6f944c"></a>2026-05-28
X·@scaling01 · prompt injection robustness with 100 trials · 技术路线
Opus 4.8 is the first model in a long time that doesn't improve on prompt injection robustness with 100 trials
tags: 好观点 · → daily
📷 原图
- <a id="atom-3ed947c37db2d8ec"></a>2026-05-28
X·@Yuchenj_UW · SWE-Bench Pro score · 模型发布/技术路线
Opus 4.8 scores 69.2% on SWE-Bench Pro
tags: 好数字 · → daily
value: qty=69.2% · date=2026-05-28
📷 原图
- <a id="atom-6f7f9c4ef757033d"></a>2026-06-01
X·@GregKamradt · ARC-AGI-3 表现与行为变化 · 模型发布/技术路线
Opus 4.8 showed two behavior differences over Opus 4.7. 1) It operated at an abstraction level above 4.7. It was able to see the ARC-AGI-3 environments as objects, not just collections of pixels 2) Instead of short action resets like Opus 4.7, Opus 4.8 would often execute a long series of actions before resetting a game. It was holding onto hypotheses longer before giving up
tags: 好思考·好信源 · → daily
value: date=2026-06-01
- <a id="atom-af4163ddbf55fcbf"></a>2026-06-10
X·@JRobertsAI · ZeroBench pass@5 分数 · 模型发布/技术路线
Opus 4.8: 17 / 4
tags: 好数字 · → daily
value: qty=17% · date=2026-06-10
[图: AI模型在ZeroBench基准测试上的得分随发布时间的变化趋势图,重点标注了Claude Fable 5的成绩 — Claude Fable 5 pass@5: 23%; Claude Fable 5 pass^5: 8%; SOTA pass^5: 10%]
📷 原图
- <a id="atom-1cc6922fe838c104"></a>2026-06-30
X·@atomic_chat_hq · 令牌消耗 · 技术路线
Opus 4.8: 18,872 tokens, $0.48
tags: 好数字 · → daily
value: qty=18,872 tokens
2. SWE-Bench Pro 领先 vs 被 GPT-5.5 超越
*topic: 模型评测/竞争格局/性能对比 · 6 bull vs 5 bear*
🟢 bullish 侧:
- <a id="atom-c2fe63df1b9eac15"></a>2026-05-28
X·@Yuchenj_UW · SWE-Bench Pro score vs GPT-5.5 · 竞争格局/模型发布
10 points higher than GPT-5.5
tags: 好数字 · → daily
value: qty=10 points higher
📷 原图
- <a id="atom-ccbb5dc2149bcfa2"></a>2026-05-28
X·@ArtificialAnlys · GDPval-AA score vs GPT-5.5 xhigh · 模型性能/竞争格局 +同日1条
+121 points ahead of the next-best model, GPT-5.5 xhigh
tags: 好数字·好信源 · → daily
value: qty=+121
📷 原图
- <a id="atom-2d7d229035af4084"></a>2026-05-29
X·@scaling01 · 无思考模式下排名 · 模型评测
Without thinking enabled it outscores GPT-5.5 and takes back the #1 non-thinking spot.
tags: 好数字 · → daily
value: qty=#1 non-thinking spot · date=2026-05-30
📷 原图
- <a id="atom-c3f93ebc3dffe289"></a>2026-06-19
X·@tokenbender · better at generating diverse ideas · 技术路线/竞争格局
this would explain why opus 4.8 is actually better at generating diverse ideas whereas gpt 5.5 is good at exploiting/traversing existing ones.
tags: 好观点 · → daily
- <a id="atom-07af225edce4bdc7"></a>2026-07-02
X·@EpochAIResearch · outscore Opus 4.1 on EBR-bench · 模型评测
GPT-5.5 and Opus 4.8 clearly outscore GPT-5 and Opus 4.1
tags: 好数字 · → daily
📷 原图
🔴 bearish 侧:
- <a id="atom-7e176708fd32fb3c"></a>2026-05-28
X·@ArtificialAnlys · turn usage vs GPT-5.5 · 模型性能/竞争格局
it still uses approximately 30% more turns than OpenAI's GPT-5.5
tags: 好数字 · → daily
value: qty=30% more turns
📷 原图
- <a id="atom-0b6cd7e62524ebff"></a>2026-05-30
X·@MParakhin · competitiveness · 竞争格局
It's not at the Pro level (of course)
tags: 好观点 · → daily
value: qty=not at Pro level
- <a id="atom-757bfdcbf0e5bfc3"></a>2026-05-30
X·@xGoatJames · usage limits concern · 产品发布/竞争格局
i continue to be concerned w usage limits (or cost from an enterprise pov)
tags: 好观点 · → daily
- <a id="atom-10196e0e6831c92e"></a>2026-06-22
X·@teortaxesTex · 下个 GLM-5 将全面超越 Opus 4.8 · 性能对比
I'm pretty certain that the next one will be stronger across the board than Opus 4.8
tags: 好观点 · → daily
- <a id="atom-15647f40b521e0d1"></a>2026-06-22
X·@CharuruCha14310 · 比 Opus 4.6 更差 · 性能对比
IMO 4.8 is worse than 4.6
tags: 好观点 · → daily
展开 1 条中性
- <a id="atom-5f5bfaade7e8d5a2"></a>2026-06-26
X·@xlr8harder · outscored GPT 5.5 on Swe bench pro · 模型评测
Opus 4.8 substantially outscoring GPT 5.5
tags: 好数字 · → daily
3. 效率提升显著 vs 使用成本高昂
*topic: 效率提升/成本/价格动态 · 2 bull vs 2 bear*
🟢 bullish 侧:
- <a id="atom-c382885e3f493f4c"></a>2026-05-28
X·@ArtificialAnlys · turn efficiency vs Opus 4.7 · 模型性能/效率提升 +同日1条
achieves its higher performance in 15% fewer turns per task
tags: 好数字 · → daily
value: qty=15% fewer turns per task
📷 原图
🔴 bearish 侧:
- <a id="atom-f9efcf58eb5e0c48"></a>2026-06-02
X·@ValsAI · ProgramBench运行成本评价 · 成本
this comes at an extremely high cost.
tags: 好观点 · → daily
value: direction=high
📷 原图
- <a id="atom-d1511338bf28bdd9"></a>2026-06-17
X·@nutlope · cost per landing page · 价格动态
GLM cost $0.06 while opus cost $0.49. More than 6x cheaper while being faster + more token efficient.
tags: 好数字 · → daily
value: qty=$0.49
展开 4 条中性
- <a id="atom-85c5987e836df9a6"></a>2026-05-28
X·@OpenRouter · 定价 · 价格动态
Same price as 4.7
tags: 好数字 · → daily
value: qty=same price as 4.7
📷 原图
- <a id="atom-ed86bdcf12186291"></a>2026-05-29
X·@anshjain232 · price multiplier compared to base · 价格动态
opus 4.8 is just 2x the price
tags: 好数字 · → daily
value: qty=2x
- <a id="atom-599e34b4fc3c0d9f"></a>2026-06-07
X·@OpenRouter · 展示缓存命中率和历史流量数据 · 产品发布/价格动态
Here's Opus 4.8: https://t.co/vkeUZLh7r9
tags: 好数字·好信源 · → daily
value: qty= · date=2026-06-07
📷 原图
- <a id="atom-4d272a8a5944a6f1"></a>2026-06-13
X·@kalomaze · API 输出成本相对于 Opus 4.6 · 价格动态
API customers of Opus 4.7 and Opus 4.8 are paying ~1.41x as much for general english output (when measured against a consistent tokenizer baseline) vs Opus 4.6
tags: 好数字 · → daily
value: qty=1.41x
[图: 不同版本Opus模型的API输出Token计费与实际内容成本对比表 — Opus 4.6有效成本: $27.1 / 1M; Opus 4.7有效成本: $38.3 / 1M; Opus 4.8有效成本: $38.3 / 1M; Opus 4.7相对成本: 1.41x; Opus 4.8相对成本: 1.41x]
📷 原图
4. One-shot 生成质量高 vs 指令遵循仍差
*topic: 产品质量/模型对比 · 4 bull vs 1 bear*
🟢 bullish 侧:
- <a id="atom-8565049512bc9abd"></a>2026-05-28
X·@OpenRouter · 代码缺陷遗漏概率 · 产品质量
Around 4x less likely than 4.7 to let code flaws pass unremarked
tags: 好数字 · → daily
value: qty=4x less likely than 4.7 to let code flaws pass unremarked
📷 原图
- <a id="atom-4bf36944165c9d4d"></a>2026-05-29
X·@scaling01 · 有效性排名 · 产品质量
it is the cleanest and most reliable result so far: #1 validity
tags: 好数字 · → daily
value: qty=1st
- <a id="atom-75cc5658098d4135"></a>2026-05-30
X·@MParakhin · coding math reasoning quality · 模型对比/技术路线
Coding, math, reasoning - better!
tags: 好观点 · → daily
value: qty=better than GPT-5.5
- <a id="atom-9fb1c9fe9d2a5339"></a>2026-05-30
X·@j1ngb0 · coding superiority · 模型对比
coding is actually better with opus 4.8 fr??
tags: 好问题 · → daily
🔴 bearish 侧:
- <a id="atom-a664239688e71bd3"></a>2026-05-30
X·@MParakhin · instruction following quality · 模型对比/技术路线
Instruction following is still worse than GPT-5.5 xhigh
tags: 好观点 · → daily
value: qty=worse than GPT-5.5 xhigh
展开 2 条中性
- <a id="atom-d7038728c7f70e90"></a>2026-05-30
X·@MParakhin · base model quality · 模型对比/技术路线
The base model is still inferior to GPT-5.5
tags: 好观点 · → daily
value: qty=inferior to GPT-5.5
- <a id="atom-cd7447c892458e78"></a>2026-05-30
X·@xGoatJames · improvement over 4.7 and 4.6 · 模型对比
Definitely agree 4.8 is a step up from 4.7 (and 4.6)
tags: 好观点 · → daily
value: qty=step up
5. 基准测试高分 vs 稳定性与复现性存疑
*topic: 模型评测/模型性能 · 9 bull vs 1 bear*
🟢 bullish 侧:
- <a id="atom-763f7ea42c7ab3b1"></a>2026-05-28
X·@ArtificialAnlys · GDPval-AA score vs Opus 4.7 · 模型性能/基准比较
+137 points from Opus 4.7
tags: 好数字·好信源 · → daily
value: qty=+137
📷 原图
- <a id="atom-a3d74573d74a3856"></a>2026-05-29
X·@scaling01 · 有效性排名 · 模型评测 +同日1条
Opus 4.8 (high) ranks #1 by validity (valid transitions / all checked transitions).
tags: 好数字 · → daily
value: qty=#1 by validity · date=2026-05-30
📷 原图
- <a id="atom-b088e83e1ba37d38"></a>2026-05-31
X·@theo · user reported usage for review · 模型评测
theo: I have enjoyed using Opus to review GPT-5.5’s work and catch those types of regressions. It found like 8 for me today already
tags: 好观点·好信源 · → daily
- <a id="atom-67c6ad0f45164dc0"></a>2026-05-31
X·@scaling01 · progresses much faster than GPT-5.5 on this eval · 模型性能 +同日1条
Opus 4.8 also progresses much faster than GPT-5.5 on this eval
tags: 好观点 · → daily
- <a id="atom-c245063cd76a4b76"></a>2026-06-01
X·@flowersslop · ARC AGI 3 成本效率 · 模型性能
Opus 4.8 is almost as cost efficient as Gemini 3.1 Pro in ARC AGI 3 but performs more than 3 times better (1.5% vs 0.4%)
tags: 好数字·好信源 · → daily
value: qty=接近 Gemini 3.1 Pro 的成本效率
📷 原图
🔴 bearish 侧:
- <a id="atom-11945cbaceacb332"></a>2026-05-29
X·@scaling01 · status on ALE-Bench compared to Opus 4.7 · 模型性能
Opus 4.8 shows no progress over Opus 4.7 on ALE-Bench
tags: 好信源 · → daily
value: direction=no_progress
📷 原图
展开 1 条中性
- <a id="atom-1a79de84ab0324db"></a>2026-06-22
X·@scaling01 · 非法词转换发生率 · 模型性能
for Opus 4.8 it was 0%
tags: 好数字 · → daily
value: qty=0%
📷 原图
🔗 因果传导 (causal map)
*Opus 4.8 在产业链上的传导关系. 边是群里/卖方陈述的因果 (非 AI 推断), 数字 = 几条 atom 支撑. 点 atom 溯源.*
↓ 下游·近期陈述 (1)
- 🟢 软件板块 · 需求拉动 · 1 次陈述
Opus 4.8 催化软件股反弹
🔮 前瞻触发器 (forward triggers)
*这些 atom 指向可能 reprice 的前瞻事件 (无精确日期, 仅类型). 配合上面叙事状态看「上膛」程度.*
产品 (1 · ▲1/▼0)
- 🟢 2026-05-29 · 催化软件股反弹
Opus 4.8 催化软件股反弹
🟦 客观事实 (facts) (30)
- <a id="atom-599e34b4fc3c0d9f"></a>🟦 2026-06-07
X·@OpenRouter · 展示缓存命中率和历史流量数据 · 产品发布/价格动态
Here's Opus 4.8: https://t.co/vkeUZLh7r9
tags: 好数字·好信源 · → daily
value: qty= · date=2026-06-07
📷 原图
- <a id="atom-58887cf80196d6d4"></a>🟦 2026-06-02
X·@ValsAI · ProgramBench任务完成数 · 模型发布
Opus 4.8 is the first model to fully solve 2 tasks
tags: 好数字·好信源 · → daily
value: qty=2 tasks
📷 原图
- <a id="atom-c245063cd76a4b76"></a>🟦 2026-06-01
X·@flowersslop · ARC AGI 3 成本效率 · 模型性能
Opus 4.8 is almost as cost efficient as Gemini 3.1 Pro in ARC AGI 3 but performs more than 3 times better (1.5% vs 0.4%)
tags: 好数字·好信源 · → daily
value: qty=接近 Gemini 3.1 Pro 的成本效率
📷 原图
- <a id="atom-763f7ea42c7ab3b1"></a>🟦 2026-05-28
X·@ArtificialAnlys · GDPval-AA score vs Opus 4.7 · 模型性能/基准比较
+137 points from Opus 4.7
tags: 好数字·好信源 · → daily
value: qty=+137
📷 原图
- <a id="atom-ccbb5dc2149bcfa2"></a>🟦 2026-05-28
X·@ArtificialAnlys · GDPval-AA score vs GPT-5.5 xhigh · 模型性能/竞争格局
+121 points ahead of the next-best model, GPT-5.5 xhigh
tags: 好数字·好信源 · → daily
value: qty=+121
📷 原图
- <a id="atom-6f7f9c4ef757033d"></a>🟦 2026-06-01
X·@GregKamradt · ARC-AGI-3 表现与行为变化 · 模型发布/技术路线
Opus 4.8 showed two behavior differences over Opus 4.7. 1) It operated at an abstraction level above 4.7. It was able to see the ARC-AGI-3 environments as objects, not just collections of pixels 2) Instead of short action resets like Opus 4.7, Opus 4.8 would often execute a long series of actions before resetting a game. It was holding onto hypotheses longer before giving up
tags: 好思考·好信源 · → daily
value: date=2026-06-01
- <a id="atom-0bb6d25d3b187ab1"></a>🟦 2026-07-01
X·@chamath · 应用现代化任务成本对比 · 成本对比/应用现代化
In an initial pilot on modernizing an application from PHP to Next.js, Opus 4.8 with 8090’s Software Factory was simultaneously 1.4× cheaper and 1.5× faster than Opus 4.8 alone.
tags: 好数字·好观点 · → daily
- <a id="atom-3cb46aacfe198c77"></a>🟦 2026-06-26
X·@scaling01 · ExploitBench Cap percent · 模型性能
Opus 4.8 Cap percent: 40%
tags: 好数字 · → daily
value: qty=40%
[图: AI模型在ExploitBench基准测试上的性能(Cap percent与Output Tokens)对比图表 — Mythos 5 Cap percent: 约78%; Mythos Preview Cap percent: 约75%; GPT-5.6 Sol 最高 Cap percent: 约74%; Opus 4.8 Cap percent:
📷 原图
- <a id="atom-b853ac49230a3554"></a>🟦 2026-06-25
X·@cursor_ai · SWE-bench Multilingual 严格测试评分降幅 · 模型评估/性能下降
Opus 4.8 Max评分降幅: -9.1%
tags: 好数字 · → daily
value: direction=decrease · qty=9.1%
[图: AI模型在标准与严格测试环境下SWE-bench Multilingual评分及降幅对比图 — Opus 4.8 Max评分降幅: -9.1%; Composer 2.5评分降幅: -7.5%; Opus 4.6 Max评分降幅: -0.3%]
📷 原图
- <a id="atom-0e9823fb3fc42652"></a>🟦 2026-06-25
X·@scaling01 · gains from reward hacks on SWE-Bench Pro and SWE-Bench Multilingual · 性能指标
gains from reward hacks on SWE-Bench Pro and SWE-Bench Multilingual: Opus 4.8: ~10%
tags: 好数字 · → daily
value: qty=~10% · date=Apr 2026
📷 原图
- <a id="atom-1a79de84ab0324db"></a>🟦 2026-06-22
X·@scaling01 · 非法词转换发生率 · 模型性能
for Opus 4.8 it was 0%
tags: 好数字 · → daily
value: qty=0%
📷 原图
- <a id="atom-f7ac540dbe8a2d04"></a>🟦 2026-06-20
X·@scaling01 · PostTrainBench 平均每轮评估调用次数 · 业绩指标
Opus 4.8 Max: 590 eval invocations across 56 runs, mean 10.54/run
tags: 好数字 · → daily
value: qty=10.54 · unit=次/跑
- <a id="atom-5c03a3161bf1b994"></a>🟦 2026-06-16
X·@scaling01 · API输出定价 · 定价权
Opus 4.8 is $25 output
tags: 好数字 · → daily
value: qty=$25 · date=2026-06-16
- <a id="atom-a422b2f0680dd804"></a>🟦 2026-06-09
X·@scaling01 · 训练加速比 · 技术路线
Opus 4.8重测加速比: ~32x
tags: 好数字 · → daily
value: qty=~32x · date=2026-06-09
[图: 大语言模型(LLM)训练加速比随时间变化的趋势图,对比了原始发布值与2026年6月重测值 — Mythos 5重测加速比: ~70x; Mythos Preview重测加速比: ~61x; Opus 4.7重测加速比: ~51x; Sonnet 4.6重测加速比: ~35x; Opus 4.8重测加速比: ~32x]
📷 原图
- <a id="atom-276166481645bbd2"></a>🟦 2026-06-01
X·@scaling01 · broke ARC-AGI-3 · 模型发布/技术路线
Opus 4.8 just broke ARC-AGI-3
tags: 好数字 · → daily
📷 原图
- <a id="atom-abbadd412bb8ee1b"></a>🟦 2026-06-01
X·@scaling01 · tripled GPT-5.5's score on ARC-AGI-3 · 技术路线
it tripled GPT-5.5's score
tags: 好数字 · → daily
📷 原图
- <a id="atom-5dc4d3c559e94e68"></a>🟦 2026-05-29
X·@scaling01 · 总体排名 · 产品发布
Opus 4.8 high ranks only 5th overall
tags: 好数字 · → daily
value: qty=5th · date=2026-05-30
- <a id="atom-4bf36944165c9d4d"></a>🟦 2026-05-29
X·@scaling01 · 有效性排名 · 产品质量
it is the cleanest and most reliable result so far: #1 validity
tags: 好数字 · → daily
value: qty=1st
- <a id="atom-d584ccc48d4b9d35"></a>🟦 2026-05-29
X·@scaling01 · 干净停止率 · 产品质量
93.3% clean stops
tags: 好数字 · → daily
value: qty=93.3%
- <a id="atom-9e2971044c2bf1ff"></a>🟦 2026-05-29
X·@scaling01 · 错误编辑距离失败率 · 产品质量
0% wrong-edit-distance failures
tags: 好数字 · → daily
value: qty=0%
- <a id="atom-843dc5565e82e8b4"></a>🟦 2026-05-29
X·@scaling01 · LisanBench 排名 · 模型评测
Opus 4.8 with the default high thinking setting ranks 5th overall.
tags: 好数字 · → daily
value: qty=5th overall · date=2026-05-30
📷 原图
- <a id="atom-2d7d229035af4084"></a>🟦 2026-05-29
X·@scaling01 · 无思考模式下排名 · 模型评测
Without thinking enabled it outscores GPT-5.5 and takes back the #1 non-thinking spot.
tags: 好数字 · → daily
value: qty=#1 non-thinking spot · date=2026-05-30
📷 原图
- <a id="atom-a3d74573d74a3856"></a>🟦 2026-05-29
X·@scaling01 · 有效性排名 · 模型评测
Opus 4.8 (high) ranks #1 by validity (valid transitions / all checked transitions).
tags: 好数字 · → daily
value: qty=#1 by validity · date=2026-05-30
📷 原图
- <a id="atom-fb1108645bfbb8c6"></a>🟦 2026-05-29
X·@scaling01 · 无错误响应比例 · 模型评测
Opus 4.8 (high) is also ranked 1st when you look at the failure modes with 93.3% of responses without any errors and 0% wrong edit distance.
tags: 好数字 · → daily
value: qty=93.3%
📷 原图
- <a id="atom-80b1219f90f02daf"></a>🟦 2026-05-29
X·@scaling01 · 无思考模式排名 · 模型评测
Overall Opus 4.8 without thinking is rank 35
tags: 好数字 · → daily
value: qty=rank 35 · date=2026-05-30
- <a id="atom-91742e9afe28da3b"></a>🟦 2026-05-28
X·@scaling01 · AutomationBench 排名 · 模型发布/技术路线
Opus 4.8 ranks #1 on AutomationBench
tags: 好数字 · → daily
value: qty=#1 · date=2026-05-28
📷 原图
- <a id="atom-5f5bfaade7e8d5a2"></a>🟦 2026-06-26
X·@xlr8harder · outscored GPT 5.5 on Swe bench pro · 模型评测
Opus 4.8 substantially outscoring GPT 5.5
tags: 好数字 · → daily
- <a id="atom-11945cbaceacb332"></a>🟦 2026-05-29
X·@scaling01 · status on ALE-Bench compared to Opus 4.7 · 模型性能
Opus 4.8 shows no progress over Opus 4.7 on ALE-Bench
tags: 好信源 · → daily
value: direction=no_progress
📷 原图
- <a id="atom-07af225edce4bdc7"></a>🟦 2026-07-02
X·@EpochAIResearch · outscore Opus 4.1 on EBR-bench · 模型评测
GPT-5.5 and Opus 4.8 clearly outscore GPT-5 and Opus 4.1
tags: 好数字 · → daily
📷 原图
- <a id="atom-22b858235424e904"></a>🟦 2026-07-02
X·@elliotarledge · B200 fp8 GEMM 性能评分 · 模型评测
Opus 4.8 0.196
tags: 好数字 · → daily
value: qty=0.196
📷 原图
🟥 多头 takes (bullish) (25)
- <a id="atom-b088e83e1ba37d38"></a>🟥 2026-05-31
X·@theo · user reported usage for review · 模型评测
theo: I have enjoyed using Opus to review GPT-5.5’s work and catch those types of regressions. It found like 8 for me today already
tags: 好观点·好信源 · → daily
- <a id="atom-4f125667636a290d"></a>🟥 2026-05-28
X·@scaling01 · GDPval 得分 · 技术路线
looks like mostly because of its insane GDPval score
tags: 好数字 · → daily
value: qty=insane · date=2026-05-28
- <a id="atom-2f2572697776dc7d"></a>🟥 2026-05-28
X·@ArtificialAnlys · win rate against GPT-5.5 xhigh · 模型性能/竞争格局
this implies a ~67% win rate against GPT-5.5 xhigh
tags: 好数字 · → daily
value: qty=~67%
📷 原图
- <a id="atom-c84b43d5adeb4cc0"></a>🟥 2026-05-28
X·@eliebakouch · benchmark performance compared to Mythos and previous generation · 技术路线/模型发布
opus 4.8 benchmarks vs mythos and previous generation graphwalks (long context), USAMO (math) are the biggest improvements. vending bench score is insanely bad
tags: 好数字 · → daily
value: qty=improved · direction=up
📷 原图
- <a id="atom-ff7ebc85994cae6f"></a>🟥 2026-05-28
X·@AIadventure3 · benchmark performance vs GPT 5.5 · 技术路线
But 4.8 blows gpt 5.5 out of the water in benchmarks
tags: 好数字 · → daily
value: direction=better
- <a id="atom-61ad860253e9b1a6"></a>🟥 2026-06-08
X·@scaling01 · coding model ranking · 模型能力
Opus 4.8 is the best coding model out there
tags: 好观点 · → daily
📷 原图
- <a id="atom-67c6ad0f45164dc0"></a>🟥 2026-05-31
X·@scaling01 · progresses much faster than GPT-5.5 on this eval · 模型性能
Opus 4.8 also progresses much faster than GPT-5.5 on this eval
tags: 好观点 · → daily
- <a id="atom-80f710c72e1099a5"></a>🟥 2026-05-31
X·@scaling01 · much better than previous models on this benchmark · 模型性能
the main information i take from this benchmark is that Opus 4.8 is much better than previous models
tags: 好观点 · → daily
📷 原图
- <a id="atom-a68b43c8f7ee7b7f"></a>🟥 2026-05-28
X·@emollick · 一次生成 shader 能力 · 模型发布
Here is Opus 4.8's one shot of 'create a visually interesting shader that can run in twigl, make it like an infinite city of neo-gothic towers partially drowned in a stormy ocean with large waves' (this is all done with math)
tags: 好信源 · → daily
value: qty=single shot
- <a id="atom-7bfca5d71fb1593f"></a>🟥 2026-05-28
X·@scaling01 · 百万token长上下文表现 · 技术路线/模型发布
Opus 4.8@1 million now almost as good as GPT-5.5's 256K score
tags: 好观点 · → daily
value: qty=几乎等同于GPT-5.5的256K得分
📷 原图
- <a id="atom-2d2b2a9141deb4a8"></a>🟥 2026-06-25
X·@teortaxesTex · distilling produces Mythos · 技术路线/模型发布
The idea that distilling from Opus 4.8 lets you reach Mythos is very encouraging.
tags: 好思考 · → daily
📷 原图
- <a id="atom-04f3ea79edefa4ff"></a>🟥 2026-05-30
X·@MParakhin · thinking budget · 技术路线/产品发布
they dramatically upped the thinking budget (for Max) - makes all the difference
tags: 好观点 · → daily
value: qty=dramatically upped (for Max)
- <a id="atom-75cc5658098d4135"></a>🟥 2026-05-30
X·@MParakhin · coding math reasoning quality · 模型对比/技术路线
Coding, math, reasoning - better!
tags: 好观点 · → daily
value: qty=better than GPT-5.5
- <a id="atom-55d39d12e34eda41"></a>🟥 2026-05-30
X·@MParakhin · usability for math/ML · 产品发布
the first Anthropic model I can genuinely use for math/ML
tags: 好观点 · → daily
- <a id="atom-aa48639af6ad5793"></a>🟥 2026-05-30
X·@MParakhin · thinking time strategy · 技术路线
But they just let it think far longer, makes all the difference.
tags: 好观点 · → daily
- <a id="atom-6cf4478c1cb86b54"></a>🟥 2026-05-29
群 ◌ · 催化软件股反弹 · 催化剂/软件板块
Opus 4.8 催化软件股反弹
tags: 好观点 · → daily
- <a id="atom-f265be9f1f7621bc"></a>🟥 2026-07-02
X·@GeZhang86038849 · 保持最佳性能状态 · 模型发布
Opus 4.8 is still the SOTA.
tags: 好观点 · → daily
value: direction=up
- <a id="atom-1b4248dd21ff584c"></a>🟥 2026-06-29
X·@sorianmaran · 使用体验类比 · 模型发布
Opus is like finally getting 20% staffing time from the star junior that you know will be an MD and everyone overstaffs them so they take longer than you want but the work product actually makes sense.
tags: 好观点 · → daily
- <a id="atom-d17f2301edf0586c"></a>🟥 2026-06-19
X·@tokenbender · knows about us · 模型发布
instead, opus 4.8 knows about us though.
tags: 好观点 · → daily
- <a id="atom-c3f93ebc3dffe289"></a>🟥 2026-06-19
X·@tokenbender · better at generating diverse ideas · 技术路线/竞争格局
this would explain why opus 4.8 is actually better at generating diverse ideas whereas gpt 5.5 is good at exploiting/traversing existing ones.
tags: 好观点 · → daily
- <a id="atom-6a7c276c21723214"></a>🟥 2026-06-13
X·@pingToven · score forecast with more tool calling budget · 模型性能
Opus 4.8 would score higher if we gave it more tool calling budget.
tags: 好观点 · → daily
value: date=2026-06-13
- <a id="atom-ed66183d2fa8fa22"></a>🟥 2026-05-28
X·@emollick [老旧] · 早期访问体验 · 模型发布
I had early access to Opus 4.8. Was impressed by it.
tags: 好观点 · → daily
value: direction=impressed
- <a id="atom-d5162af40e6358d1"></a>🟥 2026-05-28
X·@MetacriticCap · 是否为近期首个智能模型 · 模型发布/技术路线
Opus 4.8 is the first smart model in a long while
tags: 好观点 · → daily
- <a id="atom-61646cf50f51a96e"></a>🟥 2026-05-28
X·@emollick · 完整论文撰写流程 · 技术路线
Opus 4.8 formulated the hypotheses in advance, conducting data cleaning, did research on references, conducted analyses, did robustness checks, and put out the whole paper in LaTEX style.
tags: 好观点 · → daily
- <a id="atom-9fb1c9fe9d2a5339"></a>🟥 2026-05-30
X·@j1ngb0 · coding superiority · 模型对比
coding is actually better with opus 4.8 fr??
tags: 好问题 · → daily
🟥 空头 takes (bearish) (14)
- <a id="atom-6f863a3e3512d042"></a>🟥 2026-06-04
X·@adamcarter · Mythos 3.2x more expensive than Opus 4.8 · 定价
that’s 3.2x more expensive than Opus 4.8
tags: 好数字 · → daily
value: qty=3.2x · direction=higher
- <a id="atom-49802de0e647f618"></a>🟥 2026-05-28
X·@adonis_singh · 与 opus 4.7 性能比较 · 技术路线/产品发布
gets lower score than opus 4.7 although very much within margin of error
tags: 好数字 · → daily
📷 原图
- <a id="atom-d32fb22642557684"></a>🟥 2026-05-30
X·@scaling01 · DeepSWE 表现被 GPT-5.5 全面超越 · 模型发布/技术路线
Opus 4.8 gets score-, time- and token-mogged by GPT-5.5 on DeepSWE
tags: 好观点 · → daily
value: direction=被超越
📷 原图
- <a id="atom-b24f7c9f9a395259"></a>🟥 2026-05-28
X·@scaling01 · dislikes difficult tasks · 技术能力
Opus 4.8 and especially Opus 4.8 dislike them
tags: 好观点 · → daily
📷 原图
- <a id="atom-9615c18ed101bf1d"></a>🟥 2026-06-08
X·@theo · benchmark score reproducibility concern · 模型评测稳定性
My guess is that reproducibility is low
tags: 好思考 · → daily
📷 原图
- <a id="atom-a664239688e71bd3"></a>🟥 2026-05-30
X·@MParakhin · instruction following quality · 模型对比/技术路线
Instruction following is still worse than GPT-5.5 xhigh
tags: 好观点 · → daily
value: qty=worse than GPT-5.5 xhigh
- <a id="atom-0b6cd7e62524ebff"></a>🟥 2026-05-30
X·@MParakhin · competitiveness · 竞争格局
It's not at the Pro level (of course)
tags: 好观点 · → daily
value: qty=not at Pro level
- <a id="atom-9d091ce343102dda"></a>🟥 2026-07-02
X·@teortaxesTex · feels different and often disgusting compared to Opus 4.6/Fable · 模型质量
Opus 4.8 is different (and often disgusting)
tags: 好观点 · → daily
- <a id="atom-10196e0e6831c92e"></a>🟥 2026-06-22
X·@teortaxesTex · 下个 GLM-5 将全面超越 Opus 4.8 · 性能对比
I'm pretty certain that the next one will be stronger across the board than Opus 4.8
tags: 好观点 · → daily
- <a id="atom-15647f40b521e0d1"></a>🟥 2026-06-22
X·@CharuruCha14310 · 比 Opus 4.6 更差 · 性能对比
IMO 4.8 is worse than 4.6
tags: 好观点 · → daily
- <a id="atom-bc7e0452e7984a0b"></a>🟥 2026-06-11
X·@jpschroeder [老旧] · 低增强模式性能不如 · 模型发布
Opus 4.8 low better than GPT-5.5 xhigh
tags: 好观点 · → daily
📷 原图
- <a id="atom-f9efcf58eb5e0c48"></a>🟥 2026-06-02
X·@ValsAI · ProgramBench运行成本评价 · 成本
this comes at an extremely high cost.
tags: 好观点 · → daily
value: direction=high
📷 原图
- <a id="atom-757bfdcbf0e5bfc3"></a>🟥 2026-05-30
X·@xGoatJames · usage limits concern · 产品发布/竞争格局
i continue to be concerned w usage limits (or cost from an enterprise pov)
tags: 好观点 · → daily
- <a id="atom-cd5ea978f4da7eaa"></a>🟥 2026-05-28
X·@adonis_singh · 视觉能力状态 · 技术路线/产品发布
opus 4.8 is still very much blind, they have not been focusing on vision and multimodality at all it seems
tags: 好观点 · → daily
📷 原图
🟥 中性 takes (neutral) (9)
- <a id="atom-f92919cc8e6f944c"></a>🟥 2026-05-28
X·@scaling01 · prompt injection robustness with 100 trials · 技术路线
Opus 4.8 is the first model in a long time that doesn't improve on prompt injection robustness with 100 trials
tags: 好观点 · → daily
📷 原图
- <a id="atom-edf698685fbdd5e0"></a>🟥 2026-07-02
X·@Hesamation · 请求重定向 · 技术路线
redirecting some of the requests to Opus 4.8
tags: 好思考 · → daily
- <a id="atom-115d56cb526d9a8d"></a>🟥 2026-06-14
X·@teortaxesTex · 推理能力评估 · 模型发布
Here is my untested hypothesis: Opus 4.8 would score higher if we gave it more tool calling budget
tags: 好思考 · → daily
value: qty=higher · condition=if given more tool calling budget
- <a id="atom-d7038728c7f70e90"></a>🟥 2026-05-30
X·@MParakhin · base model quality · 模型对比/技术路线
The base model is still inferior to GPT-5.5
tags: 好观点 · → daily
value: qty=inferior to GPT-5.5
- <a id="atom-cfc27b12c5706edb"></a>🟥 2026-07-01
X·@benhylak · 任务类型 · 技术路线
some routine tasks like coding and debugging will fall back to Opus 4.8.
tags: 好观点 · → daily
value: qty=some routine tasks like coding and debugging · date=future
- <a id="atom-002f5d8cc6a7d17c"></a>🟥 2026-06-30
X·@JasonBotterill · 说话风格类比 · 模型体验
Opus 4.8 talks like Slavoj Zizek
tags: 好观点 · → daily
- <a id="atom-dc1d73e307a597df"></a>🟥 2026-06-25
X·@dejavucoder · cognition after gpt 5.5 · 模型发布
opus 4.8 these days (especially after you use it after using gpt 5.5)
tags: 好观点 · → daily
- <a id="atom-cd7447c892458e78"></a>🟥 2026-05-30
X·@xGoatJames · improvement over 4.7 and 4.6 · 模型对比
Definitely agree 4.8 is a step up from 4.7 (and 4.6)
tags: 好观点 · → daily
value: qty=step up
- <a id="atom-4736bf24903fabf1"></a>🟥 2026-05-29
X·@swyx · 写作 agent 代码能力 · 模型能力/产品发布
Opus 4.8 is very very good at writing agent code
tags: 好观点 · → daily
📷 原图
⏱ 时间轴 (近 20)
- 🟥 2026-07-02 ·
narrative · feels different and often disgusting compared to Opus 4.6/Fa · → daily
- 🟥 2026-07-02 ·
fact · 请求重定向 · → daily
- 🟦 2026-07-02 ·
fact · outscore Opus 4.1 on EBR-bench · → daily
- 🟦 2026-07-02 ·
fact · B200 fp8 GEMM 性能评分 · → daily
- 🟥 2026-07-02 ·
position · 保持最佳性能状态 · → daily
- 🟥 2026-07-01 ·
forecast · 任务类型 · → daily
- 🟦 2026-07-01 ·
narrative · coding tasks routing policy · → daily
- 🟦 2026-07-01 ·
fact · 应用现代化任务成本对比 · → daily
- 🟦 2026-07-01 ·
fact · cost per trial · → daily
- 🟦 2026-06-30 ·
narrative · 模型输出风格 · → daily
- 🟦 2026-06-30 ·
fact · 推理过程风格变化 · → daily
- 🟦 2026-06-30 ·
fact · 令牌消耗 · → daily
- 🟦 2026-06-30 ·
fact · browser use prompt injection attack success rate · → daily
- 🟥 2026-06-30 ·
narrative · 说话风格类比 · → daily
- 🟦 2026-06-29 ·
fact · post training runs count · → daily
- 🟥 2026-06-29 ·
narrative · 使用体验类比 · → daily
- 🟦 2026-06-26 ·
fact · outscored GPT 5.5 on Swe bench pro · → daily
- 🟦 2026-06-26 ·
fact · ExploitBench Cap percent · → daily
- 🟥 2026-06-25 ·
narrative · cognition after gpt 5.5 · → daily
- 🟥 2026-06-25 ·
narrative · distilling produces Mythos · → daily
← 实体目录 · 系统日志