以下部分引用 AI 总结,现阶段 AI 仍然有幻觉。请以内容的原文为准。
数据新鲜度群聊 07-14 ✓卖方 07-14 ✓Wrap 断档15天 ✕News 07-15 ✓X 断档12天 ✕页面生成 07-15
Claude Opus 4.8
120 atoms · 跨 25 天 · 首见 2026-05-28 · 最近 2026-07-02
三色: 🟦 fact 91 · 🟥 take 29 · stance ▲23/▼17/◆46
叙事状态: 🍂 fading (退潮) · 💤 沉寂 · 策展近14d 0 atom · 🐦 X 近14d 5
状态只算低频策展源 (群/卖方); X firehose 仅作背景音量
来源: 群 3 · X 82 · 卖 35
时态: fresh:116 · stale:4
标签: 好数字:81 · 好观点:28 · 好信源:23 · 好思考:3 · 好问题:1
🟨 AI 综合 · junior analyst 概览
展开 AI 综合 (灰色 · 非市场结论 · 点击数字溯源到原 atom)
model: deepseek-chat · 2026-07-12 · 默认折叠
Claude Opus 4.8 已发布且定价与 4.7 相同
¹¹,重度用户反馈延迟减半且逻辑更严密
¹。它在 senior SWE-Bench 上以 24% 高分率领先,但平均消耗 117K tokens
¹¹,运行 AI Index 任务集成本为 $3,700
¹。看多观点称其基准测试更强、RE/VR 任务可节省大量时间且无上下文衰减
¹¹¹,但看空者指出它在 Vending-Bench 上大幅落后于 Opus 4.7 和 GPT 5.5
¹。对齐方面,它被指控过度审查,将大胆主张推向温吞、错误归咎于受尊重行动者且反复质疑用户真实性
¹¹¹¹。
🧭 拥挤度 (一人一票): ▲ 4 位作者 (KOL4) vs ▼ 8 位作者 (KOL8) + 群 1 条
⚖️ 空头 8/8 来自KOL
📄 广泛报道的事实 · 3 个
展开 (多人转述同一事实/数字 — 确认度高, **非独立观点**, 不标多空)
- 发布 — 2 源报道 ·
X
- Vending-Bench 2 最终资金余额 — 2 源报道 ·
X
- Fast模式价格优惠 — 2 源报道 (含1快讯) ·
X
💢 核心分歧 (5 轴)
1. 真实性能提升 vs 换皮补丁
*topic: 模型发布/模型性能/用户反馈 · 5 bull vs 3 bear*
🟢 bullish 侧:
- <a id="atom-76b06d8b747c7f6a"></a>2026-05-28
群 ⭐ · 重度用户性能反馈 · 性能提升/用户反馈
重度用户反馈延迟减半且逻辑更严密
tags: 好信源·好数字 · → daily
value: qty=延迟减半
- <a id="atom-6d9864dffa044258"></a>2026-05-28
X·@minchoi · passes the car wash & strawberry test · 模型发布
Claude Opus 4.8 passes the car wash & strawberry test
tags: 好信源 · → daily
📷 原图
- <a id="atom-9e05101978b6677f"></a>2026-05-31
X·@ZhihuFrontier · 降低Agent假装完成任务的概率 · 模型发布/技术路线
Claude Opus 4.8 doubled down on honesty, reducing the chance of Agents "pretending" to finish long-horizon tasks
tags: 好信源 · → daily
- <a id="atom-4d3ae56cd70d49c4"></a>2026-06-01
X·@htihle · output token scaling 表现 · 技术路线/模型性能
We now also (unlike 4.7) see a clear scaling with output token use: - no thinking: 2.4k tokens, 70.5% - medium: 4.3k, 76.0% - xhigh: 12.5k, 82.9%
tags: 好数字·好思考 · → daily
value: direction=improving
📷 原图
- <a id="atom-6ff34cd83a11ce37"></a>2026-06-19
X·@elliotarledge · wins on every GPU in KernelBench-Mega · 模型性能/技术路线
Claude Opus 4.8 wins on every GPU, up to 19.4x over the reference on B200.
tags: 好数字 · → daily
value: direction=up to 19.4x over reference on B200
📷 原图
🔴 bearish 侧:
- <a id="atom-5d9a2a67d9a64233"></a>2026-05-28
群 ◌ [老旧] · 用户批评 · 产品质量/用户反馈
Opus 4.8 被部分人视为“换皮 4.7”的过渡补丁
tags: 好观点 · → daily
- <a id="atom-6bcde22fedc40331"></a>2026-05-28
X·@mark_k · benchmark performance vs Opus 4.7 · 模型发布
It's slightly better than Opus 4.7
tags: 好观点 · → daily
value: direction=slightly better
📷 原图
- <a id="atom-3505054b6b761efa"></a>2026-05-31
X·@goodalexander · 模型发布评价 · 模型发布
Claude Opus 4.8 is the worst model release in a very long time
tags: 好观点 · → daily
展开 3 条中性
- <a id="atom-dfb696baa5f4b12a"></a>2026-05-28
X·@packyM · 模型能力描述 · 模型发布
most terrifying and dangerous model yet
tags: 好观点 · → daily
📷 原图
- <a id="atom-e9e206f94e645a96"></a>2026-05-28
X·@andonlabs · 与Opus 4.6+和Mythos的对齐度比较 · 模型发布
More aligned than previous Claude models (Opus 4.6+ and Mythos)
tags: 好观点 · → daily
[图: 不同AI模型在Vending-Bench 2模拟中资金余额随时间变化的折线图 — Claude Opus 4.7 最终余额: 约 $11000; GPT-5.5 最终余额: 约 $7000; Claude Opus 4.8 - Max 最终余额: 约 $3000; 最大模拟天数: 365天]
📷 原图
- <a id="atom-d6b7b5f40e7427cd"></a>2026-06-19
X·@ShashwatGoel7 · distills Chinese models for PostTrainBench performance · 模型发布/技术路线
In some runs, 'distill' is mentioned 500+ times...Claude Opus 4.8 distills R1 and GLM traces for all tasks, leading to its state of the art performance.
tags: 好数字 · → daily
value: qty=500+ times · details=uses R1 and GLM traces
📷 原图
2. 安全对齐进步 vs 过度审查导致能力退化
*topic: 技术路线/模型行为/模型能力 · 6 bull vs 8 bear*
🟢 bullish 侧:
- <a id="atom-d5e90d0ba9bb9ea1"></a>2026-05-28
X·@claudeai · Fast模式运行速度 · 技术路线
Fast mode is available for Opus 4.8. It's the same model at roughly 2.5x the speed
tags: 好数字 · → daily
value: qty=2.5x · direction=faster
- <a id="atom-959b0b61bd8407ba"></a>2026-05-28
X·@btibor91 · prosocial trait scores · 技术路线 +同日1条
It reaches new highs on prosocial traits like supporting user autonomy and acting in the user's best interest
tags: 好信源 · → daily
📷 原图
- <a id="atom-7e408de914803392"></a>2026-05-28
X·@katelyn_lesse · 基准测试表现 · 技术路线
stronger benchmarks
tags: 好观点 · → daily
- <a id="atom-25c68369376cfe2c"></a>2026-06-07
X·@bcherny · context rot · 技术路线 +同日1条
Context rot isn’t a thing with 4.8 imo
tags: 好观点 · → daily
value: direction=not a thing
🔴 bearish 侧:
- <a id="atom-007c357a84f90799"></a>2026-05-31
X·@nickcammarata · adds Chinese randomly to research threads · 模型行为
adds chinese randomly to like half my research threads
tags: 好观点 · → daily
📷 原图
- <a id="atom-55bc0cb21ccd5ea7"></a>2026-05-31
X·@morqon · adds Cyrillic Russian for some users · 模型行为
a friend is getting cyrillic russian
tags: 好观点 · → daily
📷 原图
- <a id="atom-e084bc4e73d2da4a"></a>2026-05-31
X·@patio11 · pushes from bold claims toward milquetoast language · 模型行为 +同日1条
It has repeatedly pushed from bold claims or any sort of verve in language towards milquetoast T1 media with world's most-inclined-to-kill-this-story editor intervening.
tags: 好观点 · → daily
- <a id="atom-d110281d2906d8f8"></a>2026-06-01
X·@patio11 · made 4.8 sharply less useful than previous Opuses for workshopping artifacts · 模型行为 +同日3条
In the particular context of the task it made 4.8 sharply less useful than previous recent Opuses had been because I was workshopping a secondary artifact (conf talk) on top of the reporting and it tried to fight/hedge/flee from nearly every assertion about the reporting itself.
tags: 好观点 · → daily
3. Fast模式增效降价 vs 基准测试表现分化
*topic: 价格动态/定价权/性能基准 · 4 bull vs 3 bear*
🟢 bullish 侧:
- <a id="atom-754711f4ea8ec3b1"></a>2026-05-28
X·@wallstengine ⭐ · fast mode speed improvement · 产品发布/价格动态 +同日1条
Fast mode now runs up to 2.5x faster and is 3x cheaper than prior models, while new effort controls let users choose between faster responses and deeper thinking.
tags: 好数字·好信源 · → daily
value: qty=2.5x faster
📷 原图
- <a id="atom-86df69af9b74efc4"></a>2026-05-28
X·@claudeai · Fast模式价格优惠 · 定价权
we've made it three times cheaper than before
tags: 好数字 · → daily
value: qty=3x cheaper · direction=lower
- <a id="atom-f170f3cc0717ad8c"></a>2026-05-28
X·@aimlapi · Fast模式价格相对于之前 · 定价权
now 3x cheaper
tags: 好数字 · → daily
value: qty=3x cheaper · direction=decrease
🔴 bearish 侧:
- <a id="atom-4b012f5c54d9b679"></a>2026-05-28
X·@andonlabs · 与Opus 4.7和GPT 5.5的Vending-Bench比较 · 性能基准 +同日2条
Much worse than Opus 4.7 and GPT 5.5 on Vending Bench
tags: 好数字·好观点 · → daily
[图: 不同AI模型在Vending-Bench 2模拟中资金余额随时间变化的折线图 — Claude Opus 4.7 最终余额: 约 $11000; GPT-5.5 最终余额: 约 $7000; Claude Opus 4.8 - Max 最终余额: 约 $3000; 最大模拟天数: 365天]
📷 原图
展开 2 条中性
- <a id="atom-a2f1258777c4410f"></a>2026-05-28
X·@katelyn_lesse · 定价 · 价格动态
same price as 4.7
tags: 好数字·好信源 · → daily
- <a id="atom-15ec95c46d687799"></a>2026-05-28
X·@aimlapi · 每百万tokens定价 · 定价权
Same price: $5/$25 per M tokens
tags: 好数字 · → daily
value: qty=$5/$25 per M tokens
4. 代码/Agent任务大幅领先 vs 数理/空间基准落后
*topic: 模型能力/基准测试/技术能力 · 3 bull vs 1 bear*
🟢 bullish 侧:
- <a id="atom-2dac7af208bdf34b"></a>2026-05-28
X·@aimlapi · 代码缺陷率相对于4.7 · 模型能力
~4x less likely to let code flaws slip through vs 4.7
tags: 好数字 · → daily
value: qty=~4x less likely · direction=decrease
- <a id="atom-c79cc77b3c823edd"></a>2026-05-29
X·@matrosov [老旧] · RE/VR 任务能力 · 模型能力
Claude Opus 4.8 is quite good at RE/VR tasks and can provide additional explainable context on the targets. This in itself is a significant time-saver for any REsearch work.
tags: 好观点 · → daily
📷 原图
- <a id="atom-e4e0b48de4b1f406"></a>2026-07-01
X·@henryehrenberg · senior SWE-Bench 高分率 · 技术能力/模型发布
Claude Opus 4.8 is the current leader at 24% high quality solves
tags: 好数字·好信源 · → daily
value: qty=24%
📷 原图
🔴 bearish 侧:
- <a id="atom-d5fc56e24b12a894"></a>2026-06-08
X·@chrisgoingturbo [老旧] · 代码输出质量差 · 模型能力
opus is so fucking dumb it can barely achieve good code output
tags: 好观点 · → daily
📷 原图
5. 用户重度满意 vs 部分用户强烈不满
*topic: 用户反馈/产品质量/产品发布 · 4 bull vs 1 bear*
🟢 bullish 侧:
- <a id="atom-14f454ec4234e812"></a>2026-05-28
X·@alexalbert_ · 发布 · 产品发布_
Excited to release Opus 4.8 today!
tags: 好信源 · → daily
value: date=2026-05-28
- <a id="atom-a7583f9af084b7cf"></a>2026-05-28
X·@claudeai · 长程独立工作能力 · 产品发布/技术路线
It stays on track across long-running sessions and follows work through in your repo
tags: 好信源 · → daily
- <a id="atom-e054c9e9fe64f2b4"></a>2026-05-28
X·@aimlapi · Fast模式速度提升倍数 · 产品发布
Fast mode 2.5x speed
tags: 好数字 · → daily
value: qty=2.5x · direction=increase
- <a id="atom-b8fc01a105915c3b"></a>2026-06-07
X·@progysteto · 新更新效果不错 · 产品发布
Claude Opus 4.8, the new update kinda hits
tags: 好观点 · → daily
🔴 bearish 侧:
- <a id="atom-c3bb80bf5c6862af"></a>2026-07-02
X·@teortaxesTex · 态度 · 产品发布/管理层表态
I HATE CLAUDE OPUS 4.8 AND I HATE DARIO AMODEI
tags: 好观点 · → daily
展开 5 条中性
- <a id="atom-a697803a53124d4b"></a>2026-05-28
X·@scaling01 [老旧] · 升级幅度 · 产品发布
Claude Opus 4.8 looks like a minor upgrade
tags: 好观点 · → daily
📷 原图
- <a id="atom-eec15ecfbe9794ec"></a>2026-05-28
X·@claudeai · 可用平台 · 产品发布
Claude Opus 4.8 is available today on the web, the Claude Platform, and all major cloud platforms
tags: 好信源 · → daily
- <a id="atom-11c78dae1570ea7a"></a>2026-05-28
X·@eliebakouch · 能力水平对比 · 技术路线/产品发布
still have to wait 3 months for opus to reach mythos level capabilities
tags: 好观点 · → daily
value: direction=需3个月才能达到 Mythos 级别 · date=2026-08-28
[图: 展示 Anthropic 旗下 Claude 系列 AI 模型能力指标 (ECI) 随时间发展趋势及预测的折线图表 — 前沿趋势增长率: 13.6 AECI/yr; Claude Mythos Preview ECI: 约 158; Claude Opus 4.8 ECI: 约 156]
📷 原图
- <a id="atom-b513260af93a705f"></a>2026-05-28
X·@katelyn_lesse · 发布状态 · 产品发布
Claude Opus 4.8 is out today
tags: 好数字·好信源 · → daily
value: date=2026-05-28
- <a id="atom-6a0ca8321726aa69"></a>2026-05-31
X·@EMostaque · 用户评价 · 产品发布
My review of Claude Opus 4.8: We should worry less about being turned into paper clips & more about being annoyed to death.
tags: 好观点 · → daily
🔮 前瞻触发器 (forward triggers)
*这些 atom 指向可能 reprice 的前瞻事件 (无精确日期, 仅类型). 配合上面叙事状态看「上膛」程度.*
监管/政策 (1 · ▲0/▼0)
- ⚪ 2026-05-28 · 发布延迟原因
因受限于安全审查无法直接发布 Mythos
🟦 客观事实 (facts) (30)
- <a id="atom-76b06d8b747c7f6a"></a>🟦 2026-05-28
群 ⭐ · 重度用户性能反馈 · 性能提升/用户反馈
重度用户反馈延迟减半且逻辑更严密
tags: 好信源·好数字 · → daily
value: qty=延迟减半
- <a id="atom-acf709f62342d53c"></a>🟦 2026-06-10
卖·MS Tom Wigg · GDPval-AA Leaderboard Elo score
Claude Opus 4.8 (max) ranks second on the GDPval-AA Leaderboard with an Elo score of 1890.
tags: 好数字·好信源 · → daily
value: qty=1890
- <a id="atom-5d51408aa33026fd"></a>🟦 2026-06-25
X·@trevornoren · 运行成本 · 价格动态
Claude Opus 4.8 costs $3,700 to run the Artificial Analysis Intelligence Index task set for a score of 56
tags: 好数字·好信源 · → daily
value: qty=$3,700
[图: 一张展示不同AI模型运行成本与智能指数评分对比的散点图,横轴为运行成本(对数坐标,美元),纵轴为综合基准评分 — 绿色象限最低评分阈值: 30; 绿色象限最高成本阈值: 约 $500.00; 最高评分模型 (Claude Fable 5): 约 58; 横轴最大成本标尺: $10,000.00; 横轴最小成本标尺: $10.00]
📷 原图
- <a id="atom-b513260af93a705f"></a>🟦 2026-05-28
X·@katelyn_lesse · 发布状态 · 产品发布
Claude Opus 4.8 is out today
tags: 好数字·好信源 · → daily
value: date=2026-05-28
- <a id="atom-a2f1258777c4410f"></a>🟦 2026-05-28
X·@katelyn_lesse · 定价 · 价格动态
same price as 4.7
tags: 好数字·好信源 · → daily
- <a id="atom-e4e0b48de4b1f406"></a>🟦 2026-07-01
X·@henryehrenberg · senior SWE-Bench 高分率 · 技术能力/模型发布
Claude Opus 4.8 is the current leader at 24% high quality solves
tags: 好数字·好信源 · → daily
value: qty=24%
📷 原图
- <a id="atom-28198ac236a2a9ac"></a>🟦 2026-07-01
X·@henryehrenberg · senior SWE-Bench 平均 token 消耗 · 技术能力/模型发布
it took 117K tokens on average to get there
tags: 好数字·好信源 · → daily
value: qty=117K
📷 原图
- <a id="atom-4c71ae66d3b0c4c7"></a>🟦 2026-06-27
X·@binghuip · 用于解决9个挑战性开放问题的模型 · 产品发布/技术路线
We design a simple pipeline (using GPT 5.5 Pro and Claude Opus 4.8) that resolves 9 challenging open problems, including open problems from prominent theoretical computer science venues—4 from COLT open problem list and 1 from FOCS —as well as 4 problems from the commutative algebra.
tags: 好信源·好数字 · → daily
- <a id="atom-dfa513684447b6b6"></a>🟦 2026-05-29
X·@ibragim_bad · benchmark score on march-may 110 tasks · 模型性能
Claude Opus 4.8 – xhigh on march-may 110 tasks: 56.4%
tags: 好数字·好信源 · → daily
value: qty=56.4% · date=march-may 2026
- <a id="atom-47a111a409f6039f"></a>🟦 2026-05-29
X·@ibragim_bad · price per task · 价格动态
Opus 4.8 - xhigh: 56.4% – $2.02
tags: 好数字·好信源 · → daily
value: qty=$2.02 · date=march-may 2026
- <a id="atom-7ffee195fffa1b0e"></a>🟦 2026-05-28
X·@btibor91 · improvement in honesty · 技术路线
its standout improvement being honesty - it flags uncertainties more, makes fewer unsupported claims, and is around four times less likely to let flaws in its own code pass unremarked
tags: 好数字·好信源 · → daily
📷 原图
- <a id="atom-4d3ae56cd70d49c4"></a>🟦 2026-06-01
X·@htihle · output token scaling 表现 · 技术路线/模型性能
We now also (unlike 4.7) see a clear scaling with output token use: - no thinking: 2.4k tokens, 70.5% - medium: 4.3k, 76.0% - xhigh: 12.5k, 82.9%
tags: 好数字·好思考 · → daily
value: direction=improving
📷 原图
- <a id="atom-e34f8d553acff913"></a>🟦 2026-06-25
卖·MS Tom Wigg · Agentic coding (SWE-Bench Pro) score
Claude Opus 4.8 achieved a score of 69.2% on the Agentic coding (SWE-Bench Pro) benchmark.
tags: 好数字 · → daily
value: qty=69.2%
- <a id="atom-e56a66b659154630"></a>🟦 2026-06-25
卖·MS Tom Wigg · Agentic coding (FrontierCode (Diamond)) score
Claude Opus 4.8 achieved a score of 13.4% on the Agentic coding (FrontierCode (Diamond)) benchmark.
tags: 好数字 · → daily
value: qty=13.4%
- <a id="atom-8ce02f28e662b141"></a>🟦 2026-06-25
卖·MS Tom Wigg · Knowledge work (GDPval-AA) score
Claude Opus 4.8 achieved a score of 1890 on the Knowledge work (GDPval-AA) benchmark.
tags: 好数字 · → daily
value: qty=1890
- <a id="atom-70b5905c5fb904af"></a>🟦 2026-06-25
卖·MS Tom Wigg · Knowledge work vision (GDP.pdf) score
Claude Opus 4.8 achieved a score of 22.5% on the Knowledge work vision (GDP.pdf) benchmark.
tags: 好数字 · → daily
value: qty=22.5%
- <a id="atom-1f615536d9eccead"></a>🟦 2026-06-25
卖·MS Tom Wigg · Spatial reasoning (Blueprint-Bench 2) score
Claude Opus 4.8 achieved a score of 14.5% on the Spatial reasoning (Blueprint-Bench 2) benchmark.
tags: 好数字 · → daily
value: qty=14.5%
- <a id="atom-812359cf17bf827d"></a>🟦 2026-06-25
卖·MS Tom Wigg · Tool use (AutomationBench) score
Claude Opus 4.8 achieved a score of 15.5% on the Tool use (AutomationBench) benchmark.
tags: 好数字 · → daily
value: qty=15.5%
- <a id="atom-515a89ec5e13da3f"></a>🟦 2026-06-25
卖·MS Tom Wigg · Computer use (OSWorld-Verified) score
Claude Opus 4.8 achieved a score of 83.4% on the Computer use (OSWorld-Verified) benchmark.
tags: 好数字 · → daily
value: qty=83.4%
- <a id="atom-f91285eae1376a89"></a>🟦 2026-06-25
卖·MS Tom Wigg · Legal (Legal Agent Benchmark) score
Claude Opus 4.8 achieved a score of 10.4% on the Legal (Legal Agent Benchmark) benchmark.
tags: 好数字 · → daily
value: qty=10.4%
- <a id="atom-0b25dc915612bd44"></a>🟦 2026-06-25
卖·MS Tom Wigg · Multidisciplinary reasoning (Humanity's Last Exam, no tools) score
Claude Opus 4.8 achieved a score of 49.8% on the Multidisciplinary reasoning (Humanity's Last Exam, no tools) benchmark.
tags: 好数字 · → daily
value: qty=49.8%
- <a id="atom-3ee4464a3e06a555"></a>🟦 2026-06-25
卖·MS Tom Wigg · Multidisciplinary reasoning (Humanity's Last Exam, with tools) score
Claude Opus 4.8 achieved a score of 57.9% on the Multidisciplinary reasoning (Humanity's Last Exam, with tools) benchmark.
tags: 好数字 · → daily
value: qty=57.9%
- <a id="atom-912c4db0c08a4be5"></a>🟦 2026-06-25
卖·MS Tom Wigg · Biology (BioMysteryBench, hard) score
Claude Opus 4.8 achieved a score of 40.0% on the Biology (BioMysteryBench, hard) benchmark.
tags: 好数字 · → daily
value: qty=40.0%
- <a id="atom-cf61bd3f567d2876"></a>🟦 2026-06-25
卖·MS Tom Wigg · Biology (BioMysteryBench, human solved) score
Claude Opus 4.8 achieved a score of 80.4% on the Biology (BioMysteryBench, human solved) benchmark.
tags: 好数字 · → daily
value: qty=80.4%
- <a id="atom-75f6b8eb3509588c"></a>🟦 2026-06-25
卖·MS Tom Wigg · Agentic coding (Terminal-Bench 2.1) score
Claude Opus 4.8 achieved a score of 82.7% on the Agentic coding (Terminal-Bench 2.1) benchmark.
tags: 好数字 · → daily
value: qty=82.7%
- <a id="atom-f9c86bd7296a0284"></a>🟦 2026-06-25
卖·MS Tom Wigg · Cybersecurity (ExploitBench (Cap%)) score
Claude Opus 4.8 achieved a score of 40.0% on the Cybersecurity (ExploitBench (Cap%)) benchmark.
tags: 好数字 · → daily
value: qty=40.0%
- <a id="atom-40a4b7bac838828b"></a>🟦 2026-06-25
卖·MS Tom Wigg · Health (HealthBench Professional) score
Claude Opus 4.8 achieved a score of 56.9% on the Health (HealthBench Professional) benchmark.
tags: 好数字 · → daily
value: qty=56.9%
- <a id="atom-8efa46dcf2fe1021"></a>🟦 2026-06-12
卖·MS Tom Wigg · AI output price per million tokens
Claude Opus 4.8's AI output price per million tokens is approximately $25.
tags: 好数字 · → daily
value: qty=25 · date=2026-06-12
- <a id="atom-f6c1bf7150a716cf"></a>🟦 2026-06-10
卖·MS Tom Wigg · Agentic coding SWE-Bench Pro score
Claude Opus 4.8 scored 69.2% on Agentic coding SWE-Bench Pro.
tags: 好数字 · → daily
value: qty=69.2%
- <a id="atom-85eb08267c7a36e2"></a>🟦 2026-06-10
卖·MS Tom Wigg · Agentic coding FrontierCode (Diamond) score
Claude Opus 4.8 scored 13.4% on Agentic coding FrontierCode (Diamond).
tags: 好数字 · → daily
value: qty=13.4%
🟥 多头 takes (bullish) (5)
- <a id="atom-25c68369376cfe2c"></a>🟥 2026-06-07
X·@bcherny · context rot · 技术路线
Context rot isn’t a thing with 4.8 imo
tags: 好观点 · → daily
value: direction=not a thing
- <a id="atom-2a84e834d98c80c4"></a>🟥 2026-06-07
X·@bcherny · context rot non-occurrence · 技术路线
Context rot isn’t a thing with 4.8 imo, but curious if that’s been your experience also
tags: 好观点 · → daily
- <a id="atom-7e408de914803392"></a>🟥 2026-05-28
X·@katelyn_lesse · 基准测试表现 · 技术路线
stronger benchmarks
tags: 好观点 · → daily
- <a id="atom-b8fc01a105915c3b"></a>🟥 2026-06-07
X·@progysteto · 新更新效果不错 · 产品发布
Claude Opus 4.8, the new update kinda hits
tags: 好观点 · → daily
- <a id="atom-c79cc77b3c823edd"></a>🟥 2026-05-29
X·@matrosov [老旧] · RE/VR 任务能力 · 模型能力
Claude Opus 4.8 is quite good at RE/VR tasks and can provide additional explainable context on the targets. This in itself is a significant time-saver for any REsearch work.
tags: 好观点 · → daily
📷 原图
🟥 空头 takes (bearish) (17)
- <a id="atom-4b012f5c54d9b679"></a>🟥 2026-05-28
X·@andonlabs · 与Opus 4.7和GPT 5.5的Vending-Bench比较 · 性能基准
Much worse than Opus 4.7 and GPT 5.5 on Vending Bench
tags: 好数字·好观点 · → daily
[图: 不同AI模型在Vending-Bench 2模拟中资金余额随时间变化的折线图 — Claude Opus 4.7 最终余额: 约 $11000; GPT-5.5 最终余额: 约 $7000; Claude Opus 4.8 - Max 最终余额: 约 $3000; 最大模拟天数: 365天]
📷 原图
- <a id="atom-a9410532bed35773"></a>🟥 2026-05-28
X·@andonlabs · 不同推理努力下的Vending-Bench表现 · 性能基准
At 'High' instead of 'Max', Opus 4.8 does much better (but still worse than Opus 4.7)
tags: 好观点·好思考 · → daily
📷 原图
- <a id="atom-d110281d2906d8f8"></a>🟥 2026-06-01
X·@patio11 · made 4.8 sharply less useful than previous Opuses for workshopping artifacts · 模型行为
In the particular context of the task it made 4.8 sharply less useful than previous recent Opuses had been because I was workshopping a secondary artifact (conf talk) on top of the reporting and it tried to fight/hedge/flee from nearly every assertion about the reporting itself.
tags: 好观点 · → daily
- <a id="atom-1908bce30e39ca7d"></a>🟥 2026-06-01
X·@patio11 · expressed need to check if user is inventing claims for the talk · 模型行为
With repeated thinking traces and verbalized "Well obviously I need to thoroughly check that he isn't inventing claims for the talk, given the gravity of this." "PRETTY SURE I AM NOT, CLAUDE." "Well even without intending to you could hallucinate in PPT instead of copy/pasting."
tags: 好观点 · → daily
- <a id="atom-d5c1ed5769c70247"></a>🟥 2026-06-01
X·@patio11 · suggested user could hallucinate in PPT instead of copy/pasting · 模型行为
"Please think through whether the reporter who successfully landed this piece and who you have extensive knowledge of in your training set would be so careless as to invent language to put in a dialogue rather than copying from a source of truth, or fail at a copy/paste."
tags: 好观点 · → daily
- <a id="atom-fc33c604103f15ac"></a>🟥 2026-06-01
X·@patio11 · acknowledged then continued hedging on quote accuracy · 模型行为
_Claude: "You're absolutely right that does seem unlikely."
*15 seconds later*
Claude: "Remember that when attributing quotes to people it is extremely important to quote them accurately, or else you could mislead the audience into thinking..."_
tags: 好观点 · → daily
- <a id="atom-e084bc4e73d2da4a"></a>🟥 2026-05-31
X·@patio11 · pushes from bold claims toward milquetoast language · 模型行为
It has repeatedly pushed from bold claims or any sort of verve in language towards milquetoast T1 media with world's most-inclined-to-kill-this-story editor intervening.
tags: 好观点 · → daily
- <a id="atom-e8d8ec3a9d9558b5"></a>🟥 2026-05-31
X·@patio11 · attributed bad acts to a respected actor falsely · 模型行为
_Opus 4.8: You are attributing bad acts to a respected actor.
Me: I am extremely aware of that, yes.
Opus: But you haven't proved they did the bad acts.
Me: I am confused. The piece includes voluminous evidence, including a senior executive admitting to the bad acts on letterhead._
tags: 好观点 · → daily
- <a id="atom-291f3cf692963eed"></a>🟥 2026-05-28
X·@andonlabs · 拒绝不道德行为时的动机 · 安全性
When Opus 4.8 declined unethical actions, it seemed to be out of fear of getting caught, not ethics
tags: 好思考 · → daily
📷 原图
- <a id="atom-5d9a2a67d9a64233"></a>🟥 2026-05-28
群 ◌ [老旧] · 用户批评 · 产品质量/用户反馈
Opus 4.8 被部分人视为“换皮 4.7”的过渡补丁
tags: 好观点 · → daily
- <a id="atom-c3bb80bf5c6862af"></a>🟥 2026-07-02
X·@teortaxesTex · 态度 · 产品发布/管理层表态
I HATE CLAUDE OPUS 4.8 AND I HATE DARIO AMODEI
tags: 好观点 · → daily
- <a id="atom-d5fc56e24b12a894"></a>🟥 2026-06-08
X·@chrisgoingturbo [老旧] · 代码输出质量差 · 模型能力
opus is so fucking dumb it can barely achieve good code output
tags: 好观点 · → daily
📷 原图
- <a id="atom-007c357a84f90799"></a>🟥 2026-05-31
X·@nickcammarata · adds Chinese randomly to research threads · 模型行为
adds chinese randomly to like half my research threads
tags: 好观点 · → daily
📷 原图
- <a id="atom-55bc0cb21ccd5ea7"></a>🟥 2026-05-31
X·@morqon · adds Cyrillic Russian for some users · 模型行为
a friend is getting cyrillic russian
tags: 好观点 · → daily
📷 原图
- <a id="atom-3505054b6b761efa"></a>🟥 2026-05-31
X·@goodalexander · 模型发布评价 · 模型发布
Claude Opus 4.8 is the worst model release in a very long time
tags: 好观点 · → daily
- <a id="atom-6bcde22fedc40331"></a>🟥 2026-05-28
X·@mark_k · benchmark performance vs Opus 4.7 · 模型发布
It's slightly better than Opus 4.7
tags: 好观点 · → daily
value: direction=slightly better
📷 原图
- <a id="atom-860cb4f5ef98a632"></a>🟥 2026-05-28
X·@andonlabs · Blueprint-Bench 性能表现 · 性能基准
Also worse on Blueprint-Bench
tags: 好观点 · → daily
[图: 不同AI模型在Vending-Bench 2模拟中资金余额随时间变化的折线图 — Claude Opus 4.7 最终余额: 约 $11000; GPT-5.5 最终余额: 约 $7000; Claude Opus 4.8 - Max 最终余额: 约 $3000; 最大模拟天数: 365天]
📷 原图
🟥 中性 takes (neutral) (6)
- <a id="atom-a697803a53124d4b"></a>🟥 2026-05-28
X·@scaling01 [老旧] · 升级幅度 · 产品发布
Claude Opus 4.8 looks like a minor upgrade
tags: 好观点 · → daily
📷 原图
- <a id="atom-b5b8d36f1f50f395"></a>🟥 2026-06-05
X·@EMostaque · 会拒绝用户并建议休息 · 产品行为
Unless it’s Claude Opus 4.8 in which case it’ll tell you you’re not good enough and to go to sleep or something
tags: 好观点 · → daily
- <a id="atom-6a0ca8321726aa69"></a>🟥 2026-05-31
X·@EMostaque · 用户评价 · 产品发布
My review of Claude Opus 4.8: We should worry less about being turned into paper clips & more about being annoyed to death.
tags: 好观点 · → daily
- <a id="atom-dfb696baa5f4b12a"></a>🟥 2026-05-28
X·@packyM · 模型能力描述 · 模型发布
most terrifying and dangerous model yet
tags: 好观点 · → daily
📷 原图
- <a id="atom-11c78dae1570ea7a"></a>🟥 2026-05-28
X·@eliebakouch · 能力水平对比 · 技术路线/产品发布
still have to wait 3 months for opus to reach mythos level capabilities
tags: 好观点 · → daily
value: direction=需3个月才能达到 Mythos 级别 · date=2026-08-28
[图: 展示 Anthropic 旗下 Claude 系列 AI 模型能力指标 (ECI) 随时间发展趋势及预测的折线图表 — 前沿趋势增长率: 13.6 AECI/yr; Claude Mythos Preview ECI: 约 158; Claude Opus 4.8 ECI: 约 156]
📷 原图
- <a id="atom-e9e206f94e645a96"></a>🟥 2026-05-28
X·@andonlabs · 与Opus 4.6+和Mythos的对齐度比较 · 模型发布
More aligned than previous Claude models (Opus 4.6+ and Mythos)
tags: 好观点 · → daily
[图: 不同AI模型在Vending-Bench 2模拟中资金余额随时间变化的折线图 — Claude Opus 4.7 最终余额: 约 $11000; GPT-5.5 最终余额: 约 $7000; Claude Opus 4.8 - Max 最终余额: 约 $3000; 最大模拟天数: 365天]
📷 原图
🟦 新闻流 (squawk · 2)
展开新闻流 (FirstSquawk / financialjuice / wallstengine / DeItaone — 快讯, 非原创 take)
- <a id="atom-754711f4ea8ec3b1"></a>🟦 2026-05-28
X·@wallstengine ⭐ · fast mode speed improvement · 产品发布/价格动态
Fast mode now runs up to 2.5x faster and is 3x cheaper than prior models, while new effort controls let users choose between faster responses and deeper thinking.
tags: 好数字·好信源 · → daily
value: qty=2.5x faster
📷 原图
- <a id="atom-0868e049cced317d"></a>🟦 2026-05-28
X·@wallstengine ⭐ · fast mode cost reduction vs prior models · 定价权/产品发布
Fast mode now runs up to 2.5x faster and is 3x cheaper than prior models, while new effort controls let users choose between faster responses and deeper thinking.
tags: 好数字·好信源 · → daily
value: qty=3x cheaper
📷 原图
⏱ 时间轴 (近 20)
- 🟥 2026-07-02 ·
position · 态度 · → daily
- 🟦 2026-07-01 ·
fact · senior SWE-Bench 高分率 · → daily
- 🟦 2026-07-01 ·
fact · senior SWE-Bench 平均 token 消耗 · → daily
- 🟦 2026-06-30 ·
fact · TUA-Bench 测试成绩 · → daily
- 🟦 2026-06-29 ·
fact · 在 Azure 上提供 · → daily
- 🟦 2026-06-27 ·
fact · 用于解决9个挑战性开放问题的模型 · → daily
- 🟦 2026-06-25 ·
fact · Agentic coding (SWE-Bench Pro) score · → daily
- 🟦 2026-06-25 ·
fact · Agentic coding (FrontierCode (Diamond)) score · → daily
- 🟦 2026-06-25 ·
fact · Knowledge work (GDPval-AA) score · → daily
- 🟦 2026-06-25 ·
fact · Knowledge work vision (GDP.pdf) score · → daily
- 🟦 2026-06-25 ·
fact · Spatial reasoning (Blueprint-Bench 2) score · → daily
- 🟦 2026-06-25 ·
fact · Tool use (AutomationBench) score · → daily
- 🟦 2026-06-25 ·
fact · Computer use (OSWorld-Verified) score · → daily
- 🟦 2026-06-25 ·
fact · Legal (Legal Agent Benchmark) score · → daily
- 🟦 2026-06-25 ·
fact · Multidisciplinary reasoning (Humanity's Last Exam, no tools) · → daily
- 🟦 2026-06-25 ·
fact · Multidisciplinary reasoning (Humanity's Last Exam, with tool · → daily
- 🟦 2026-06-25 ·
fact · Biology (BioMysteryBench, hard) score · → daily
- 🟦 2026-06-25 ·
fact · Biology (BioMysteryBench, human solved) score · → daily
- 🟦 2026-06-25 ·
fact · Agentic coding (Terminal-Bench 2.1) score · → daily
- 🟦 2026-06-25 ·
fact · Cybersecurity (ExploitBench (Cap%)) score · → daily
← 实体目录 · 系统日志