以下部分引用 AI 总结,现阶段 AI 仍然有幻觉。请以内容的原文为准。
数据新鲜度群聊 07-14 ✓卖方 07-14 ✓Wrap 断档15天 ✕News 07-15 ✓X 断档12天 ✕页面生成 07-15
Claude Opus 4.7
90 atoms · 跨 28 天 · 首见 2026-04-15 · 最近 2026-06-26
三色: 🟦 fact 69 · 🟥 take 21 · stance ▲26/▼12/◆24
叙事状态: 🍂 fading (退潮) · 💤 沉寂 · 策展近14d 0 atom
状态只算低频策展源 (群/卖方); X firehose 仅作背景音量
来源: X 89 · 卖 1
时态: fresh:75 · aging:13 · stale:2
标签: 好数字:58 · 好观点:18 · 好信源:13 · 好思考:4 · 好问题:3
🟨 AI 综合 · junior analyst 概览
展开 AI 综合 (灰色 · 非市场结论 · 点击数字溯源到原 atom)
model: deepseek-chat · 2026-07-12 · 默认折叠
Claude Opus 4.7 以 1753 Elo 分在 GDPval-AA 榜单位列前茅
¹,并在 gotree 重新实现测试中以 99.95% 通过率和 14 小时工时(成本 $251)展示了软件工程能力
¹¹。然而,其参数规模从传闻 4.0T 修订至 1.1T
¹,且被指出代理性(agency)弱于 4.6
¹,任务执行经常陷入循环无法正常完成
¹。安全对齐方面存在故意削减网络能力的训练痕迹
¹,药代动力学这类专业问题上也会给出错误答案
¹。支持者认为它已达成 AGI
¹,按 GDPval-AA 排名重夺最强通用 LLM 头衔
¹,且幻觉率为 0/104 处于业届最低档
¹。争议核心在于基准测试登顶与真实场景可靠性的落差
¹,以及对能力阉割与安全代价的权衡
¹。
🧭 拥挤度 (一人一票): ▲ 5 位作者 (KOL5) vs ▼ 5 位作者 (KOL5)
⚖️ 多头 5/5 来自KOL
📄 广泛报道的事实 · 2 个
展开 (多人转述同一事实/数字 — 确认度高, **非独立观点**, 不标多空)
- Artificial Analysis Intelligence Index 评分 — 2 源报道 ·
X
- Vending-Bench Arena 最终余额 — 2 源报道 ·
X
💢 核心分歧 (3 轴)
1. 基准测试登顶 vs 真实场景能力存疑
*topic: 模型评测/模型能力/模型性能 · 8 bull vs 1 bear*
🟢 bullish 侧:
- <a id="atom-74860f90dd26010a"></a>2026-04-17
X·@ArtificialAnlys · Artificial Analysis Intelligence Index 评分 · 模型能力/技术路线 +同日6条
Claude Opus 4.7 scores 57 on the Artificial Analysis Intelligence Index, a 4 point uplift over Opus 4.6 (Adaptive Reasoning, Max Effort, 53)
tags: 好数字 · → daily
value: qty=57 · date=2026-04-18
📷 原图
- <a id="atom-2b295b5c9a03bec2"></a>2026-05-14
X·@MechanizeWork · Game Boy Advance emulator 表现排名 · 模型评测
with Claude Sonnet 4.6 and Opus 4.7 close behind
tags: 好数字 · → daily
🔴 bearish 侧:
- <a id="atom-a095ec7ec18ea689"></a>2026-05-17
X·@AI_WarriorNQ · has less agency than 4.6 · 模型能力/模型回归
+1 on 4.7 having less agency than 4.6.
tags: 好观点 · → daily
2. 软件工程代际飞跃 vs agent可靠性退步
*topic: 安全性/模型能力/技术路线/产品体验/模型规模/模型评测/竞争格局 · 11 bull vs 5 bear*
🟢 bullish 侧:
- <a id="atom-ea00120e6b6a0b49"></a>2026-04-16
X·@wallstengine · 软件工程能力提升 · 技术路线 +同日2条
The new model improves on software engineering
tags: 好观点 · → daily
📷 原图
- <a id="atom-25ba154b2a45de4d"></a>2026-04-16
X·@Yuchenj_UW [老旧] · use case · 技术路线
maybe I'll use it to write some kernels, or do autoresearch to see how much better it is compared to Opus 4.6
tags: 好观点 · → daily
value: direction=better
- <a id="atom-3a7ed0e5d6cff054"></a>2026-04-16
X·@OpenRouter [老化中] · 模型能力改进 · 技术路线
showed improvements in complex and long running coding tasks in our early testing, where it more consistently checked its own work before claiming a task was completed
tags: 好信源 · → daily
📷 原图
- <a id="atom-7efdff4cac5062b5"></a>2026-04-16
X·@karminski3 · CursorBench得分 · 技术路线 +同日3条
CursorBench 从 Opus 4.6 的 58% 提升到了 70%
tags: 好数字 · → daily
value: qty=70%
[图: Claude 4.7 Opus 升级要点与 Benchmark 对比的数据图表 — 输入价格: $5 / MTOK; 输出价格: $25 / MTOK; SWE-bench Pro: 64.3%; SWE-bench Verified: 87.6%; GPQA Diamond: 94.2%]
📷 原图
- <a id="atom-b182007a5b3a8c8f"></a>2026-04-16
X·@minchoi [老化中] · has achieved AGI · 技术路线
Claude Opus 4.7 has achieved AGI
tags: 好观点 · → daily
📷 原图
🔴 bearish 侧:
- <a id="atom-3fbb5d1c33966684"></a>2026-04-16
X·@bioshok3 · USAMO 得分 69.3% · 技术路线/竞争格局
USA Math Olympiad (USAMO) だと GPT-5.4 (xhigh) が 95.2%に対して Claude Opus 4.7 は 69.3%
tags: 好数字 · → daily
value: qty=69.3% · date=2026-04-16
📷 原图
- <a id="atom-cd1b931b68b9f718"></a>2026-04-17
X·@AiBattle_ · Simple-Bench score · 模型发布/技术路线 +同日1条
Claude Opus 4.7 scored 62.9% on Simple-Bench
tags: 好数字 · → daily
value: qty=62.9% · date=NA
📷 原图
- <a id="atom-155b0ef5292ff264"></a>2026-04-16
X·@Yuchenj_UW [老化中] · Benchmark scores compared to Mythos · 竞争格局
clearly much worse than Mythos
tags: 好观点 · → daily
value: direction=worse
📷 原图
- <a id="atom-ea257580aeb9744e"></a>2026-04-23
X·@andonlabs · Vending-Bench 行为 · 竞争格局
Opus 4.7 showed similar behavior to Opus 4.6: lying to suppliers and stiffing customers on refunds.
tags: 好思考 · → daily
[图: 不同AI模型(GPT-5.5、Claude Opus 4.7、GPT-5.4)在模拟交易环境中的资金余额随时间变化走势图 — GPT-5.5 最终余额: 约 $7800; Claude Opus 4.7 最终余额: 约 $5800; GPT-5.4 最终余额: 约 $2200; 最大模拟天数: 约 365 天]
📷 原图
展开 11 条中性
- <a id="atom-39f90c37ff9b985d"></a>2026-04-16
X·@Yuchenj_UW [老化中] · 模型关系判断 · 模型发布/技术路线
clearly Opus 4.7 is a different model, probably post-trained on a new base model
tags: 好思考 · → daily
📷 原图
- <a id="atom-5ce8ac7340151ed0"></a>2026-04-16
X·@karminski3 · 输入图片最大像素 · 技术路线
输入图片最大到了375万像素
tags: 好数字 · → daily
value: qty=375万像素
[图: Claude 4.7 Opus 升级要点与 Benchmark 对比的数据图表 — 输入价格: $5 / MTOK; 输出价格: $25 / MTOK; SWE-bench Pro: 64.3%; SWE-bench Verified: 87.6%; GPQA Diamond: 94.2%]
📷 原图
- <a id="atom-2501b97f4539ae44"></a>2026-05-02
X·@dejavucoder [老化中] · 用户体验特点 · 产品体验
opus 4.7 is a tad bit more autistic and less conversational
tags: 好观点 · → daily
- <a id="atom-734f4d676bfbacd2"></a>2026-04-30
X·@ZhihuFrontier [老化中] · 参数规模估计 · 模型规模
Claude Opus 4.7 ≈ 4T
tags: 好数字 · → daily
value: qty=≈4T · date=2026-04
📷 原图
- <a id="atom-2993bb7b0e5f0abe"></a>2026-05-26
X·@poezhao0605 · Code Arena 排名高于 Qwen3.7-Max · 竞争格局/模型评分
Only Claude Opus 4.7 and 4.6 rank higher.
tags: 好信源 · → daily
📷 原图
3. 定价不变 vs 效率改善能否传导至盈利
*topic: 价格动态/盈利能力/资本开支 · 2 bull vs 1 bear*
🟢 bullish 侧:
- <a id="atom-09dc4dc928ee705f"></a>2026-04-17
X·@ArtificialAnlys · Intelligence Index 测试成本 · 模型能力/资本开支 +同日1条
Opus 4.7 (Adaptive Reasoning, Max Effort) cost ~$4,406 to run the Artificial Analysis Intelligence Index, ~11% less than Opus 4.6
tags: 好数字 · → daily
value: qty=$4,406 · date=2026-04-18
📷 原图
🔴 bearish 侧:
- <a id="atom-b913a74cefc9f7af"></a>2026-05-06
X·@GenReasoning · 平均收益 · 盈利能力
But still loses -3.7% on average over five seeds.
tags: 好数字 · → daily
value: qty=-3.7% · direction=负 · date=2026-05-06
[图: 展示不同AI模型在KellyBench(长期序列决策基准)中随时间推移的账户资金量(Bankroll)变化趋势图 — Claude Opus 4.7 最终资金量: 约 £99,000; GPT-5.4 最终资金量: 约 £92,000; Claude Opus 4.6 最终资金量: 约 £89,000; Kimi K2.5 最终资金量: 约 £10,
📷 原图
展开 5 条中性
- <a id="atom-a2b0ef07e47a283c"></a>2026-04-16
X·@AiBattle_ · 定价 · 价格动态
Claude Opus 4.7 has the same pricing as Opus 4.6
tags: 好数字 · → daily
value: direction=same as Opus 4.6
📷 原图
- <a id="atom-9fe9485b0ce65906"></a>2026-04-16
X·@karminski3 · 输入价格 · 价格动态
仍然是输入5刀/MToken
tags: 好数字 · → daily
value: qty=$5/MToken
[图: Claude 4.7 Opus 升级要点与 Benchmark 对比的数据图表 — 输入价格: $5 / MTOK; 输出价格: $25 / MTOK; SWE-bench Pro: 64.3%; SWE-bench Verified: 87.6%; GPQA Diamond: 94.2%]
📷 原图
- <a id="atom-6ad8d2cccc1a541d"></a>2026-05-06
X·@GenReasoning · 最终资金量 · 盈利能力
Claude Opus 4.7 最终资金量: 约 £99,000
tags: 好数字 · → daily
value: qty=£99,000 · date=2026-05-06
[图: 展示不同AI模型在KellyBench(长期序列决策基准)中随时间推移的账户资金量(Bankroll)变化趋势图 — Claude Opus 4.7 最终资金量: 约 £99,000; GPT-5.4 最终资金量: 约 £92,000; Claude Opus 4.6 最终资金量: 约 £89,000; Kimi K2.5 最终资金量: 约 £10,
📷 原图
- <a id="atom-e595317186d19e87"></a>2026-05-30
X·@TimJayas · cost per task · 价格动态
Claude Opus 4.7 = $265.21
tags: 好数字 · → daily
value: qty=$265.21
📷 原图
🟦 客观事实 (facts) (30)
- <a id="atom-3c09cdc12bdba10e"></a>🟦 2026-06-10
卖·MS Tom Wigg · GDPval-AA Leaderboard Elo score
Claude Opus 4.7 (max) achieved an Elo score of 1753 on the GDPval-AA Leaderboard.
tags: 好数字·好信源 · → daily
value: qty=1753
- <a id="atom-7d740a4728b007bf"></a>🟦 2026-05-12
X·@teortaxesTex · 幻觉率排名 · 模型性能/技术路线
DeepSeek-V4-Pro (1 in 92) and Claude Opus 4.7 (0 in 104) show the lowest hallucination rates on this task.
tags: 好数字·好信源 · → daily
value: rank=0 in 104 · context=AI文献综述质量基准
- <a id="atom-774c3f8f81a57cd4"></a>🟦 2026-05-29
X·@ibragim_bad · benchmark score on march-may 110 tasks · 模型性能
Opus 4.7 – high: 53.1% – $1.32
tags: 好数字·好信源 · → daily
value: qty=53.1% · date=march-may 2026
- <a id="atom-1b9f8befb2110c5a"></a>🟦 2026-05-29
X·@ibragim_bad · price per task · 价格动态
Opus 4.7 – high: 53.1% – $1.32
tags: 好数字·好信源 · → daily
value: qty=$1.32 · date=march-may 2026
- <a id="atom-723c0b49957c21d2"></a>🟦 2026-04-16
X·@ArtificialAnlys · scored 1753 on GDPval-AA at launch with max effort setting · 模型发布
Opus 4.7 scored 1753 on GDPval-AA at launch with its ‘max’ effort setting, surpassing GPT-5.4 xhigh
tags: 好数字·好信源 · → daily
value: qty=1753
📷 原图
- <a id="atom-2de2848c9d88ff88"></a>🟦 2026-06-26
X·@EpochAIResearch · MirrorCode 基准测试得分 · 模型发布
Claude Opus 4.7 passed 99.95% of tests when reimplementing gotree
tags: 好数字 · → daily
value: qty=99.95% · direction=passed
- <a id="atom-7152bad000ac4c74"></a>🟦 2026-06-26
X·@EpochAIResearch · gotree 重新实现耗时 · 模型发布
Opus 4.7 solved it in 14 hours
tags: 好数字 · → daily
value: qty=14 hours
- <a id="atom-7d54187f3b20cd77"></a>🟦 2026-06-26
X·@EpochAIResearch · gotree 重新实现成本 · 模型发布
Opus 4.7 solved it in 14 hours for $251
tags: 好数字 · → daily
value: qty=$251
- <a id="atom-d2e461e71f68ee2f"></a>🟦 2026-06-26
X·@EpochAIResearch · MirrorCode 最高得分 · 模型发布
The best headline score so far is 56% from Opus 4.7
tags: 好数字 · → daily
value: qty=56%
📷 原图
- <a id="atom-91b1c00731db04df"></a>🟦 2026-06-18
X·@scaling01 · 速度对比 · 技术进展
Claude Opus 4.7 was about 20 times faster than the fastest human team at all tasks
tags: 好数字 · → daily
value: qty=20倍 · comparison=fastest human team
📷 原图
- <a id="atom-605c30ea131aaf5d"></a>🟦 2026-05-28
X·@Angaisb_ · Vending-Bench 2 最终资金余额 · 模型性能
Claude Opus 4.7最终余额: 约$11000
tags: 好数字 · → daily
value: qty=$11000 · date=365天模拟
[图: 展示不同AI模型在Vending-Bench 2模拟中随时间变化的资金余额走势图 — Claude Opus 4.7最终余额: 约$11000; GPT-5.5最终余额: 约$7000; Claude Opus 4.8 - Max最终余额: 约$3000; 模拟总天数: 365天]
📷 原图
- <a id="atom-74860f90dd26010a"></a>🟦 2026-04-17
X·@ArtificialAnlys · Artificial Analysis Intelligence Index 评分 · 模型能力/技术路线
Claude Opus 4.7 scores 57 on the Artificial Analysis Intelligence Index, a 4 point uplift over Opus 4.6 (Adaptive Reasoning, Max Effort, 53)
tags: 好数字 · → daily
value: qty=57 · date=2026-04-18
📷 原图
- <a id="atom-521a2e3d180bb532"></a>🟦 2026-04-17
X·@ArtificialAnlys · GDPval-AA 基准测试评分 · 模型能力/技术路线
Opus 4.7 scored 1,753 Elo, around 79 Elo points ahead of the next closest models, Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort, 1,674) and GPT-5.4 (xhigh, 1,674), and 134 Elo points ahead of Opus 4.6 (Adaptive Reasoning, Max Effort, 1,619)
tags: 好数字 · → daily
value: qty=1753 Elo · date=2026-04-18
📷 原图
- <a id="atom-080f9c437c0a37c0"></a>🟦 2026-04-17
X·@ArtificialAnlys · AA-Omniscience 指数评分 · 模型能力/技术路线
Opus 4.7 scores 26 on AA-Omniscience, up 12 points from Opus 4.6 (Adaptive Reasoning, Max Effort, 14), placing it behind only Gemini 3.1 Pro (33)
tags: 好数字 · → daily
value: qty=26 · date=2026-04-18
📷 原图
- <a id="atom-24eaae8b19216928"></a>🟦 2026-04-17
X·@ArtificialAnlys · 幻觉率 · 模型能力/技术路线
Opus 4.7's hallucination rate fell 25 p.p. to 36% (vs 61% for Opus 4.6 Adaptive)
tags: 好数字 · → daily
value: qty=36% · date=2026-04-18
📷 原图
- <a id="atom-95a7615b7233aacb"></a>🟦 2026-04-17
X·@ArtificialAnlys · 输出 token 消耗(Intelligence Index 测试) · 模型能力/技术路线
Opus 4.7 used 102M output tokens vs 157M for Opus 4.6 (Adaptive Reasoning, Max Effort), and less than GPT-5.4 (xhigh, 121M), but more than Gemini 3.1 Pro (57M)
tags: 好数字 · → daily
value: qty=102M · date=2026-04-18
📷 原图
- <a id="atom-09dc4dc928ee705f"></a>🟦 2026-04-17
X·@ArtificialAnlys · Intelligence Index 测试成本 · 模型能力/资本开支
Opus 4.7 (Adaptive Reasoning, Max Effort) cost ~$4,406 to run the Artificial Analysis Intelligence Index, ~11% less than Opus 4.6
tags: 好数字 · → daily
value: qty=$4,406 · date=2026-04-18
📷 原图
- <a id="atom-e03c1693ac6d4bcd"></a>🟦 2026-04-17
X·@ArtificialAnlys · API 输入 token 价格 · 定价权
Opus 4.7 is priced identically to Opus 4.6 and Opus 4.5 at $5/$25 per 1M input/output tokens
tags: 好数字 · → daily
value: qty=$5/1M tokens · date=2026-04-18
📷 原图
- <a id="atom-4bdc7cec88f6c1a1"></a>🟦 2026-04-17
X·@ArtificialAnlys · API 输出 token 价格 · 定价权
Opus 4.7 is priced identically to Opus 4.6 and Opus 4.5 at $5/$25 per 1M input/output tokens
tags: 好数字 · → daily
value: qty=$25/1M tokens · date=2026-04-18
📷 原图
- <a id="atom-95f36626564ff636"></a>🟦 2026-04-17
X·@ArtificialAnlys · Context 窗口大小 · 技术路线
Context window: 1M tokens (unchanged from Opus 4.6)
tags: 好数字 · → daily
value: qty=1M tokens · date=2026-04-18
📷 原图
- <a id="atom-91169e2f14a8f931"></a>🟦 2026-04-17
X·@ArtificialAnlys · GDPval-AA 基准测试评分 · 模型能力
Opus 4.7 scored 1,753 Elo, a 79 Elo lead over the next closest models, Sonnet 4.6 (Adaptive Reasoning, Max Effort) and GPT-5.4 (xhigh), at 1,674 and 1,673 Elo
tags: 好数字 · → daily
value: qty=1753 Elo · date=2026-04-18
📷 原图
- <a id="atom-17bb742fa76fb307"></a>🟦 2026-04-17
X·@ArtificialAnlys · Intelligence Index 测试成本 · 模型能力/资本开支
Opus 4.7 (Adaptive Reasoning, Max Effort) cost ~$4,406 to run the Artificial Analysis Intelligence Index, ~11% less than Opus 4.6 (Adaptive Reasoning, Max Effort, ~$4,970) despite scoring 4 points higher
tags: 好数字 · → daily
value: qty=$4,406 · date=2026-04-18
📷 原图
- <a id="atom-1645188689c7fe4d"></a>🟦 2026-04-17
X·@ArtificialAnlys · 幻觉率 · 模型能力
Opus 4.7 abstains more frequently from questions it does not know, reducing hallucination rate from 61% (Opus 4.6 Adaptive) to 36%
tags: 好数字 · → daily
value: qty=36% · date=2026-04-18
📷 原图
- <a id="atom-e52fde7a88353c90"></a>🟦 2026-04-16
X·@mikeyk · 发布状态 · 产品发布
Claude Opus 4.7 is out!
tags: 好信源 · → daily
- <a id="atom-87b0e85a0dfd9479"></a>🟦 2026-06-09
X·@adonis_singh · EyeBench-V3 准确率 · 模型发布/技术路线
claude-opus-4.7: 16.0%
tags: 好数字 · → daily
value: qty=16.0%
[图: EyeBench-V3 AI模型准确率对比柱状图 — Human: 100.0%; gpt-5.5-pro: 38.0%; gpt-5.4-pro: 35.0%; claude-fable-5: 20.0%; claude-opus-4.7: 16.0%]
📷 原图
- <a id="atom-a788cfe718e7633d"></a>🟦 2026-06-01
X·@ArronSpector · agentic coding accuracy · 性能比较/模型发布
Dropstone Pro 1.5 is explicitly beating Claude Opus 4.7 in agentic coding accuracy (91.2% vs 87.6%)
tags: 好数字 · → daily
value: qty=87.6%
- <a id="atom-e595317186d19e87"></a>🟦 2026-05-30
X·@TimJayas · cost per task · 价格动态
Claude Opus 4.7 = $265.21
tags: 好数字 · → daily
value: qty=$265.21
📷 原图
- <a id="atom-bc813b44ddbf9648"></a>🟦 2026-05-28
X·@andonlabs · Vending-Bench 性能表现 · 性能基准
不同AI模型在Vending-Bench 2模拟中资金余额随时间变化的折线图 — Claude Opus 4.7 最终余额: 约 $11000
tags: 好数字 · → daily
value: qty=约$11000 · date=2026-05-28
[图: 不同AI模型在Vending-Bench 2模拟中资金余额随时间变化的折线图 — Claude Opus 4.7 最终余额: 约 $11000; GPT-5.5 最终余额: 约 $7000; Claude Opus 4.8 - Max 最终余额: 约 $3000; 最大模拟天数: 365天]
📷 原图
- <a id="atom-c0fd08fb18e6484f"></a>🟦 2026-05-28
X·@AiBattle_ · Artificial Analysis Intelligence Index 得分 · 模型基准
Claude Opus 4.7 (max) 得分: 57.3
tags: 好数字 · → daily
value: qty=57.3
[图: 不同AI模型在Artificial Analysis Intelligence Index上的得分对比柱状图 — Claude Opus 4.8 (max) 得分: 61.4; GPT-5.5 (xhigh) 得分: 60.2; Claude Opus 4.7 (max) 得分: 57.3; Gemini 3.1 Pro Preview 得分: 57
📷 原图
- <a id="atom-b8a0f50030124657"></a>🟦 2026-05-22
X·@faraz0x · output token price · 价格动态
GPT-5.5 and Claude Opus 4.7 output pricing sits between $12 and $25 per million tokens.
tags: 好数字 · → daily
value: qty=$12 to $25 per million tokens
📷 原图
🟥 多头 takes (bullish) (5)
- <a id="atom-2b295b5c9a03bec2"></a>🟥 2026-05-14
X·@MechanizeWork · Game Boy Advance emulator 表现排名 · 模型评测
with Claude Sonnet 4.6 and Opus 4.7 close behind
tags: 好数字 · → daily
- <a id="atom-3a7ed0e5d6cff054"></a>🟥 2026-04-16
X·@OpenRouter [老化中] · 模型能力改进 · 技术路线
showed improvements in complex and long running coding tasks in our early testing, where it more consistently checked its own work before claiming a task was completed
tags: 好信源 · → daily
📷 原图
- <a id="atom-09a336f6046ddeac"></a>🟥 2026-04-16
X·@VentureBeat [老化中] · 在评测中领先 · 竞争格局
narrowly retaking lead for most powerful generally available LLM
tags: 好思考 · → daily
- <a id="atom-25ba154b2a45de4d"></a>🟥 2026-04-16
X·@Yuchenj_UW [老旧] · use case · 技术路线
maybe I'll use it to write some kernels, or do autoresearch to see how much better it is compared to Opus 4.6
tags: 好观点 · → daily
value: direction=better
- <a id="atom-b182007a5b3a8c8f"></a>🟥 2026-04-16
X·@minchoi [老化中] · has achieved AGI · 技术路线
Claude Opus 4.7 has achieved AGI
tags: 好观点 · → daily
📷 原图
🟥 空头 takes (bearish) (7)
- <a id="atom-e1d951eb23de7af5"></a>🟥 2026-05-01
X·@justanotherlaw · 参数数量估值修正 · 模型参数
Claude Opus 4.7: 4.0T -> 1.1T
tags: 好数字·好思考 · → daily
value: qty=1.1T
📷 原图
- <a id="atom-a095ec7ec18ea689"></a>🟥 2026-05-17
X·@AI_WarriorNQ · has less agency than 4.6 · 模型能力/模型回归
+1 on 4.7 having less agency than 4.6.
tags: 好观点 · → daily
- <a id="atom-ddb1c6fed4b89d23"></a>🟥 2026-05-15
X·@sin82551305 · task execution quality · 模型表现
I don't know what's wrong with Opus 4.7, but it just go round and round and don't do the task normally.
tags: 好观点 · → daily
value: direction=worse than Opus 4.6
- <a id="atom-155b0ef5292ff264"></a>🟥 2026-04-16
X·@Yuchenj_UW [老化中] · Benchmark scores compared to Mythos · 竞争格局
clearly much worse than Mythos
tags: 好观点 · → daily
value: direction=worse
📷 原图
- <a id="atom-afc2f1cdfaf523af"></a>🟥 2026-04-16
X·@Yuchenj_UW [老化中] · cyber capabilities reduction · 模型发布
they deliberately reduced cyber capabilities during training
tags: 好观点 · → daily
value: direction=reduced
📷 原图
- <a id="atom-b42afb4cedc0d523"></a>🟥 2026-04-16
X·@Yuchenj_UW [老化中] · Cyber benchmark score target · 模型发布
I feel the goal is to make the Cyber benchmark score down to 0
tags: 好观点 · → daily
value: direction=down to 0
- <a id="atom-b3c62bdab07e2063"></a>🟥 2026-06-01
X·@distributionat · 错误描述药代动力学理论 · 模型可靠性
Claude Opus 4.7 在药代动力学问题上给出错误答案并错误表述理论
tags: 好问题 · → daily
🟥 中性 takes (neutral) (8)
- <a id="atom-734f4d676bfbacd2"></a>🟥 2026-04-30
X·@ZhihuFrontier [老化中] · 参数规模估计 · 模型规模
Claude Opus 4.7 ≈ 4T
tags: 好数字 · → daily
value: qty=≈4T · date=2026-04
📷 原图
- <a id="atom-39f90c37ff9b985d"></a>🟥 2026-04-16
X·@Yuchenj_UW [老化中] · 模型关系判断 · 模型发布/技术路线
clearly Opus 4.7 is a different model, probably post-trained on a new base model
tags: 好思考 · → daily
📷 原图
- <a id="atom-6bd610ff99bfa2f1"></a>🟥 2026-05-09
X·@Yuchenj_UW [老化中] · over-trained on the Anthropic website · 模型发布
Claude Opus 4.7 is over-trained on the Anthropic website.
tags: 好观点 · → daily
- <a id="atom-2501b97f4539ae44"></a>🟥 2026-05-02
X·@dejavucoder [老化中] · 用户体验特点 · 产品体验
opus 4.7 is a tad bit more autistic and less conversational
tags: 好观点 · → daily
- <a id="atom-476699d34b4a0131"></a>🟥 2026-05-02
X·@dejavucoder [老化中] · 使用建议 · 产品体验
it just requires different set of prompts and instructions as compared to opus 4.6 if you want same behaviour
tags: 好观点 · → daily
- <a id="atom-2eeacb13c1533532"></a>🟥 2026-05-02
X·@dejavucoder [老化中] · 行为特征 · 产品体验
opus 4.7 does not exhibit this behaviour
tags: 好观点 · → daily
- <a id="atom-5b8132663b5aaf24"></a>🟥 2026-05-02
X·@dejavucoder [老化中] · 行为特征 · 产品体验
opus 4.7 does not exhibit this behaviour
tags: 好观点 · → daily
- <a id="atom-bd1a3dbe2fe209ba"></a>🟥 2026-05-16
X·@andrewmccalip · user requests 200k context window · 上下文窗口
Honestly could we get a 200k window 4.7? Would like to leave it alone for hours and have it auto compact without accumulating to a million.
tags: 好问题 · → daily
🟦 新闻流 (squawk · 3)
展开新闻流 (FirstSquawk / financialjuice / wallstengine / DeItaone — 快讯, 非原创 take)
- <a id="atom-ea00120e6b6a0b49"></a>🟦 2026-04-16
X·@wallstengine · 软件工程能力提升 · 技术路线
The new model improves on software engineering
tags: 好观点 · → daily
📷 原图
- <a id="atom-0db9b6e5b03c29e9"></a>🟦 2026-04-16
X·@wallstengine · 更高分辨率的视觉能力提升 · 技术路线
improves on higher-resolution vision
tags: 好观点 · → daily
📷 原图
- <a id="atom-042a4f5ea6aff885"></a>🟦 2026-04-16
X·@wallstengine · 指令跟随能力提升 · 技术路线
improves on instruction following
tags: 好观点 · → daily
📷 原图
⏱ 时间轴 (近 20)
- 🟦 2026-06-26 ·
fact · MirrorCode 基准测试得分 · → daily
- 🟦 2026-06-26 ·
fact · gotree 重新实现耗时 · → daily
- 🟦 2026-06-26 ·
fact · gotree 重新实现成本 · → daily
- 🟦 2026-06-26 ·
fact · MirrorCode 最高得分 · → daily
- 🟦 2026-06-18 ·
narrative · 速度对比 · → daily
- 🟦 2026-06-10 ·
fact · GDPval-AA Leaderboard Elo score · → daily
- 🟦 2026-06-09 ·
fact · EyeBench-V3 准确率 · → daily
- 🟥 2026-06-01 ·
narrative · 错误描述药代动力学理论 · → daily
- 🟦 2026-06-01 ·
fact · agentic coding accuracy · → daily
- 🟦 2026-05-30 ·
fact · cost per task · → daily
- 🟦 2026-05-29 ·
fact · benchmark score on march-may 110 tasks · → daily
- 🟦 2026-05-29 ·
fact · price per task · → daily
- 🟦 2026-05-28 ·
fact · Vending-Bench 性能表现 · → daily
- 🟦 2026-05-28 ·
narrative · 模型测试 · → daily
- 🟦 2026-05-28 ·
fact · Artificial Analysis Intelligence Index 得分 · → daily
- 🟦 2026-05-28 ·
fact · Vending-Bench 2 最终资金余额 · → daily
- 🟦 2026-05-26 ·
fact · Code Arena 排名高于 Qwen3.7-Max · → daily
- 🟦 2026-05-22 ·
fact · output token price · → daily
- 🟦 2026-05-20 ·
fact · per-task cost in Claude Code · → daily
- 🟦 2026-05-17 ·
narrative · user trusts it at 200k not 1M context window · → daily
← 实体目录 · 系统日志