以下部分引用 AI 总结,现阶段 AI 仍然有幻觉。请以内容的原文为准。
数据新鲜度群聊 07-14 ✓卖方 07-14 ✓Wrap 断档15天 ✕News 07-15 ✓X 断档12天 ✕页面生成 07-15
GPT-5.5
411 atoms · 跨 67 天 · 首见 2026-04-15 · 最近 2026-07-03
三色: 🟦 fact 248 · 🟥 take 163 · stance ▲132/▼42/◆101
叙事状态: 🍂 fading (退潮) · 💤 沉寂 · 策展近14d 0 atom · 🐦 X 近14d 19
状态只算低频策展源 (群/卖方); X firehose 仅作背景音量
来源: 群 6 · X 374 · 卖 31
时态: fresh:342 · aging:54 · stale:15
标签: 好数字:220 · 好观点:134 · 好信源:41 · 好思考:34 · 好问题:9
别名 (合并): GPT 5.5 · GPT5.5
🟨 AI 综合 · junior analyst 概览
展开 AI 综合 (灰色 · 非市场结论 · 点击数字溯源到原 atom)
model: deepseek-chat · 2026-07-12 · 默认折叠
GPT-5.5 训练完成于3月24日
¹,在 GDPval-AA 上取得 1769 Elo 分
¹,且参数量从 9.7T 修正为 1.5T
¹。看多方指出其天花板高于 Opus 4.8
¹,超过 10M 令牌后性能激增
¹,最近结果令人震惊
¹。看空方则报告模型性能下滑
¹,在简单心智理论测试中失败
¹,且 SWE-Marathon 排名表现异常
¹。
🧭 拥挤度 (一人一票): ▲ 57 位作者 (KOL57) + 群 1 条 vs ▼ 15 位作者 (KOL15)
⚖️ 多头 57/57 来自KOL
🔥 核心观点 · 1 个
≥2 个独立的人各自表达 take (已排除新闻 bot) · 多空按人算 · 跨语言合并
- 参数量估测
2 人观点 · ⚪ 中性 · X · +1快讯 · 🥇 @deedydas ★★ (2026-04-29)
<details><summary>展开 4 人 · 2026-04-29→2026-05-16</summary>
- ·
X @deedydas ★★ [2026-04-29]
- · X @yiran2037840 ★ [2026-05-02]
- · X @ar0cket1 ★ [2026-05-16]
- ◆ X @ZhihuFrontier 译 [2026-04-30]
📄 广泛报道的事实 · 12 个
展开 (多人转述同一事实/数字 — 确认度高, **非独立观点**, 不标多空)
- API推理输入价格 — 4 源报道 ·
X
- average pass rate — 3 源报道 ·
X
- SWE-bench Pro 得分 — 5 源报道 (含2快讯) ·
X
- タスク解決時間 — 3 源报道 (含1快讯) ·
X
- 推理速度 — 2 源报道 ·
X
- HLE 得分 — 2 源报道 ·
X
- Codex accuracy — 2 源报道 ·
X
- ARC-AGI-3 得分 — 2 源报道 ·
X
- pricing — 2 源报道 ·
X
- Vending-Bench Arena 最终余额 — 2 源报道 ·
X
- 发布 — 2 源报道 (含1快讯) ·
X
- DeepSWE测试通过率 — 2 源报道 (含1快讯) ·
X
💢 核心分歧 (4 轴)
1. 定价翻倍 vs 效率提升:价值权衡
*topic: 价格动态/定价权 · 4 bull vs 3 bear*
🟢 bullish 侧:
- <a id="atom-45b61ef55b6ab36e"></a>2026-04-23
X·@scaling01 [老化中] · dominates cost-performance frontier on Artificial Analysis Index · 技术路线/价格动态
The GPT-5.5 model family completely dominates the cost-performance frontier on the Artificial Analysis Index
tags: 好观点 · → daily
📷 原图
- <a id="atom-934b2bac462b583c"></a>2026-05-01
X·@htihle · 单位token价格 · 定价权
less than half the price of 5.4 (xhigh), despite the high cost/token
tags: 好数字 · → daily
value: qty=低于GPT-5.4 xhigh的一半 · date=2026-05-01
📷 原图
- <a id="atom-041bf9f41c5484ed"></a>2026-05-19
X·@Angaisb_ · 相比Gemini 3.5 Flash更便宜且更智能 · 竞争格局/价格动态
GPT-5.5 medium is both cheaper and smarter than Gemini 3.5 Flash too
tags: 好观点 · → daily
📷 原图
- <a id="atom-353335be85a5be57"></a>2026-06-11
X·@EverydayAI_ · 成本效益上超越 Anthropic 最新模型 · 定价权/竞争格局
still outpunching Anthropic's newest Fable 5
tags: 好观点 · → daily
[图: 大语言模型在DeepSWE任务上的成功率与平均每任务成本对比图 — 最高成功率: 约70%; 最高单次任务成本: 约$21.25; 最低单次任务成本: $0.0]
📷 原图
🔴 bearish 侧:
- <a id="atom-95132410b5cfc1ec"></a>2026-04-23
X·@ArtificialAnlys · per-token 定价 · 定价权
Per-token pricing has doubled from GPT-5.4 to $5/$30 per 1M input/output tokens
tags: 好数字 · → daily
value: qty=$5/$30 per 1M input/output tokens
[图: 展示各主流大语言模型在Artificial Analysis Intelligence Index v4.0评估中得分的数据柱状图 — GPT-5.5 (xhigh): 60; Claude Opus 4.7 (max): 57; Gemini 3.1 Pro Preview: 57; GPT-5.4 (xhigh): 57; Kimi K2.6:
📷 原图
- <a id="atom-5266e7d58c117d7b"></a>2026-05-03
X·@orenbahari [老化中] · cost premium vs open model for long interactive tasks · 产品发布/价格动态
GPT 5.5 is about 15% more expensive for smart broader knowledge
tags: 好数字 · → daily
value: qty=15%
- <a id="atom-abd76c8842469791"></a>2026-06-27
X·@RYANHINGSHING · 价格对比 Claude Opus 4.8 · 价格动态
上一代旗舰对决是GPT-5.5 vs Claude Opus 4.8,GPT-5.5贵一点点
tags: 好数字 · → daily
展开 3 条中性
- <a id="atom-18862d41d23eaace"></a>2026-06-01
X·@aimlapi · 定价 · 产品发布/价格动态
GPT-5.5: $0.42
tags: 好数字 · → daily
value: qty=$0.42
- <a id="atom-59883576ae0c74e2"></a>2026-06-09
X·@deedydas · price per M output tokens · 价格动态
AND it's about the same price as GPT 5.5 ($10/M input, $45/M output) vs Fable 5 ($10/M input, $50/M output) and 6x cheaper than GPT 5.5 Pro.
tags: 好数字 · → daily
value: qty=$45
- <a id="atom-990dc53c434c4a47"></a>2026-06-10
X·@scaling01 · cost relative to GPT-5.4 · 定价权/产品发布
GPT-5.5 is 2x more expensive than GPT-5.4
tags: 好数字 · → daily
value: qty=2×
📷 原图
2. 幻觉率 86% 对比竞品:可靠性短板
*topic: 模型评测/模型性能/网络安全 · 11 bull vs 10 bear*
🟢 bullish 侧:
- <a id="atom-80595b97a139735d"></a>2026-04-25
X·@mweinbach [老化中] · issue reduction compared to earlier models · 模型性能 +同日1条
With GPT 5.5, there are significantly fewer (if any) issues it's finding
tags: 好观点 · → daily
- <a id="atom-deaf340c18398c5c"></a>2026-05-02
X·@scaling01 ⭐ · 超过10M tokens后性能激增 · 模型性能
after that they go absolutely ballistic
tags: 好观点·好思考 · → daily
value: qty=10M · date=2026-05-02
📷 原图
- <a id="atom-d7fa9927ed09c2c1"></a>2026-05-05
X·@scaling01 · token efficiency before step 11 · 模型性能 +同日1条
One thing we can say is that before the step-11 gate, stronger models seem more token-efficient.
tags: 好观点 · → daily
📷 原图
- <a id="atom-663b5880bb0f6963"></a>2026-05-10
X·@code_star · 编码任务表现排名 · 模型性能
GPT 5.5 by a mile
tags: 好观点 · → daily
value: qty=1st
- <a id="atom-a711284eaa02a00c"></a>2026-05-15
X·@ArtificialAnlys · GDPval-AA head-to-head comparison win rate against Claude 4 Sonnet · 模型性能
GPT-5.5 is expected to win ~98% of head-to-head comparisons on realistic work outputs against Claude 4 Sonnet, the leading model in GDPval-AA a year ago
tags: 好数字 · → daily
value: qty=~98% · date=May 2026
[图: 展示GDPval-AA评估中关于库存事件分析的数据面板和供应商事件数量图表 — 总解决成本: $82,567; 平均单次事件成本: $2,014; 最高单次事件成本: $12,000; 前三家供应商成本占比: 88%; 有效事件记录数: 50]
📷 原图
🔴 bearish 侧:
- <a id="atom-dae2c5f2bc774789"></a>2026-05-14
X·@mweinbach · pass rate dropped 6% · 模型性能 +同日2条
Pass rate dropped 6% today for GPT 5.5 from Margin Labs, something def changed
tags: 好数字·好信源 · → daily
value: qty=6% · direction=down
📷 原图
- <a id="atom-24d0c6e4d592b9af"></a>2026-05-22
X·@scaling01 · SWE-bench Pro 得分 · 模型性能 +同日5条
SWE-bench Pro: Mythos -> 77.8%, GPT-5.5 -> 58.6%
tags: 好数字 · → daily
value: qty=58.6%
📷 原图
- <a id="atom-9a269d95bf757620"></a>2026-05-31
X·@scaling01 · progresses slower than Opus 4.8 on this eval · 模型性能
Opus 4.8 also progresses much faster than GPT-5.5 on this eval
tags: 好观点 · → daily
展开 4 条中性
- <a id="atom-bf30b89551c8fea9"></a>2026-05-01
X·@dejavucoder · PostTrainBench 性能接近 · 模型性能
gpt 5.5 very close
tags: 好数字 · → daily
value: qty=very close
- <a id="atom-e67bf25e9110a601"></a>2026-05-05
X·@scaling01 [老化中] · step-11 gate · 模型性能/技术路线
The change after step 11 could be a compounding effect. It could also be that step 11 acts as a gate: once a model gets past it, later progress becomes easier or at least more reachable within the benchmark setup.
tags: 好思考 · → daily
value: direction=pass
📷 原图
- <a id="atom-659bd9ec55afd28a"></a>2026-06-18
X·@ChrissGPT · 模型工作耗时 · 模型性能
GPT-5.5 Extra High worked for 34 minutes and 42 seconds
tags: 好数字 · → daily
value: qty=34 minutes and 42 seconds · date=2026-06-19
- <a id="atom-dbb02f5aa6daf7ef"></a>2026-06-19
X·@HarshithLucky3 · DeepSWE score · 模型性能
GPT 5.5 xhigh dropped 70 to 67 (-3%)
tags: 好数字 · → daily
value: qty=67 · date=v1.1
📷 原图
3. SWE-Bench Pro 58.6%:编码能力是否退步
*topic: 模型评测/模型能力 · 13 bull vs 4 bear*
🟢 bullish 侧:
- <a id="atom-0169d28725b288db"></a>2026-04-30
X·@sama · 任务解决时间 · 模型评测/算力效率
GPT-5.5 solved a task that takes a human expert ~12 hours in under 11 minutes
tags: 好数字 · → daily
value: qty=11分钟
- <a id="atom-e74f30e260e21b0f"></a>2026-04-23
X·@ArtificialAnlys · Terminal-Bench Hard 排名第一 · 模型评测 +同日4条
GPT-5.5 (xhigh) leads Terminal-Bench Hard, GDPval-AA and our newly hosted APEX-Agents-AA
tags: 好数字 · → daily
[图: 展示各主流大语言模型在Artificial Analysis Intelligence Index v4.0评估中得分的数据柱状图 — GPT-5.5 (xhigh): 60; Claude Opus 4.7 (max): 57; Gemini 3.1 Pro Preview: 57; GPT-5.4 (xhigh): 57; Kimi K2.6:
📷 原图
- <a id="atom-449bc90ed4d7f695"></a>2026-05-01
X·@scaling01 [老化中] · function calling performance · 模型能力
GPT-5.5 has absolutely maxxed out function calling
tags: 好观点 · → daily
value: direction=maxxed_out
- <a id="atom-d1752f728ed644cd"></a>2026-05-05
X·@MatternJustus [老旧] · 性能对比 · 模型能力
GPT-5.5 outperforms all other models by a wide margin
tags: 好观点 · → daily
- <a id="atom-79c1112851f78157"></a>2026-05-08
X·@gdb [老化中] · 能力评价 · 模型能力
GPT-5.5 is both very capable and very succinct
tags: 好观点 · → daily
🔴 bearish 侧:
- <a id="atom-07848f00f9dc2194"></a>2026-04-23
X·@ArtificialAnlys · AA-Omniscience hallucination rate · 模型评测
it has a hallucination rate of 86% - vs Opus 4.7 (max) at 36%, and Gemini 3.1 Pro Preview at 50%
tags: 好数字 · → daily
value: qty=86%
[图: 展示各主流大语言模型在Artificial Analysis Intelligence Index v4.0评估中得分的数据柱状图 — GPT-5.5 (xhigh): 60; Claude Opus 4.7 (max): 57; Gemini 3.1 Pro Preview: 57; GPT-5.4 (xhigh): 57; Kimi K2.6:
📷 原图
- <a id="atom-f2db651c1917ebc8"></a>2026-05-01
X·@scaling01 · Claude Code harness performance comparison vs Opus 4.7 · 模型能力/技术路线
it doesn't beat Opus 4.7 in the Claude Code harness even with almost 2 more hours of working time via reprompting
tags: 好数字 · → daily
value: direction=worse · comparison=Opus 4.7
- <a id="atom-7d189ab9d76636b1"></a>2026-06-28
X·@9hills · PRD实现不可用 · 模型能力
GPT-5.5:不可用,懒惰,大量未实现功能和罗列性质的页面设计。建议Human-in-loop生成详细功能点后再执行。
tags: 好观点 · → daily
- <a id="atom-a2a8e5004add8ccd"></a>2026-06-28
X·@Xianbao_QIAN · 近期降智严重 · 模型能力
5.5 最近的降智太厉害了。感觉已经快不可用了。还得是开源模型
tags: 好观点 · → daily
展开 5 条中性
- <a id="atom-b937f189660eff90"></a>2026-04-23
X·@ArtificialAnlys · AA Intelligence Index 领先幅度 · 模型评测
OpenAI’s new model tops the Artificial Analysis Intelligence Index by 3 points
tags: 好数字 · → daily
value: qty=3 pts
[图: 展示各主流大语言模型在Artificial Analysis Intelligence Index v4.0评估中得分的数据柱状图 — GPT-5.5 (xhigh): 60; Claude Opus 4.7 (max): 57; Gemini 3.1 Pro Preview: 57; GPT-5.4 (xhigh): 57; Kimi K2.6:
📷 原图
- <a id="atom-22d90c7da0e10b75"></a>2026-05-10
X·@SebastienBubeck · 能力出现时间 · 模型能力/技术路线
What he talks about couldn't have happened before GPT-5.5
tags: 好观点 · → daily
value: date=GPT-5.5发布前
- <a id="atom-acafbc20936d3b0c"></a>2026-05-13
X·@AndrewCurran_ · cyber capabilities · 模型能力
GPT-5.5 solved “The Last Ones” on 3 of 10 attempts.
tags: 好数字 · → daily
📷 原图
- <a id="atom-9ad81e6cc6c9500c"></a>2026-06-29
X·@Meituan_LongCat · SWE-bench Pro评分 · 模型评测
SWE-bench Pro: 59.5 (GPT-5.5: 58.6)
tags: 好数字 · → daily
value: qty=58.6
📷 原图
4. 参数量从 9.7T 修正为 1.5T:架构信任危机
*topic: 技术路线/模型发布 · 65 bull vs 14 bear*
🟢 bullish 侧:
- <a id="atom-ad654366d68767c1"></a>2026-04-23
X·@bioshok3 [老旧] · トークンあたりのパフォーマンス · 技术路线
様々な指標でgpt-5.5はトークンあたりのパフォーマンスが高い。
tags: 好观点 · → daily
value: qty=高い(高) · direction=up
📷 原图
- <a id="atom-6401211c3bf985eb"></a>2026-04-30
X·@bioshok3 · タスク解決時間 · 技术路线/产品发布
GPT-5.5 solved a task that takes a human expert ~12 hours in under 11 minutes at
tags: 好数字·好信源 · → daily
value: qty=人間の専門家が約12時間かかるタスクを11分未満で解決
- <a id="atom-719859c354055644"></a>2026-04-24
X·@yacineMTB [老化中] · Codex 执行方式 · 模型发布
I love watching it do things the wrong slow inefficient way even though I could guide it to be better because I don't really care it will be able to do it
tags: 好观点 · → daily
- <a id="atom-1aa22398d26f81c9"></a>2026-04-24
X·@steipete [老化中] · 能力提升 · 模型发布/技术路线
GPT 5.5 is definitely a step up in the character game.
tags: 好观点 · → daily
value: direction=up
📷 原图
- <a id="atom-0f1f6c7489d3bcf5"></a>2026-04-25
X·@scaling01 [老化中] · 推测身份 · 模型发布
gpt-5.5 might be him him
tags: 好思考 · → daily
🔴 bearish 侧:
- <a id="atom-13ca13684e8f9679"></a>2026-04-24
X·@tokenbender [老化中] · 防御性行为 · 技术路线
GPT-5.5 的防御性行为被夸大,但仍会在更长上下文中发生
tags: 好观点 · → daily
value: qty=超过200K上下文 · direction=longer
- <a id="atom-d7cff87e671f76db"></a>2026-05-01
X·@scaling01 · PostTrainBench 结果对比 Opus 4.7 在 Claude Code 测试中的表现 · 模型发布/竞争格局
PostTrainBench results for GPT-5.5 are in it doesn't beat Opus 4.7 in the Claude Code harness even with almost 2 more hours of working time via reprompting
tags: 好信源 · → daily
value: direction=not beat · comparative=Opus 4.7
📷 原图
- <a id="atom-c4d0d9c4b5f5970d"></a>2026-05-03
X·@teortaxesTex · self-written code did not realize architecture mechanisms · 技术路线
GPT-5.5's self-written code didn't realize these
tags: 好观点 · → daily
- <a id="atom-a25209e599232cc8"></a>2026-05-07
X·@spicey_lemonade [老化中] · fails at simple theory of mind test · 技术路线/模型发布
GPT 5.5 fails at simple theory of mind. This is one of the most important tests for sentience.
tags: 好思考·好问题 · → daily
📷 原图
- <a id="atom-e397de4f9af46fd4"></a>2026-05-09
X·@Yuchenj_UW [老化中] · weirdly weak at frontend · 技术路线
GPT-5.5 is still weirdly weak at frontend.
tags: 好观点 · → daily
展开 37 条中性
- <a id="atom-1f3d4a399939c88e"></a>2026-04-15
X·@chatgpt21 · 发布预期时间 · 模型发布
GPT 5.5 as quoted here will come just unfortunately not this week. the good news is it’s still before the end of April and tentative for next week
tags: 好信源 · → daily
value: date=预计下周(4月底前)
📷 原图
- <a id="atom-be3c1f5f67baa8c2"></a>2026-04-30
X·@eliebakouch [老化中] · performance comparison with Mythos for cyber · 模型发布/技术路线
wait what, gpt5.5 is on par with mythos for cyber?
tags: 好问题 · → daily
value: direction=on par
📷 原图
- <a id="atom-0d53800e47928e60"></a>2026-04-30
X·@teortaxesTex · benchmark pass rate average · 模型发布/技术路线
GPT-5.5 average pass rate of 71.4% (±8.0%)
tags: 好数字 · → daily
value: qty=71.4% · date=2026-04-30
📷 原图
- <a id="atom-ca6370a7c90b1e08"></a>2026-05-01
X·@hungjng69679118 [老化中] · 参数量被认为5万亿甚至10万亿级别 · 技术路线
GPT 5.5...被认为是5万亿甚至10万亿级别的参数量
tags: 好思考 · → daily
value: qty=5万亿-10万亿
- <a id="atom-378863d8b50e0bda"></a>2026-05-01
X·@htihle · 单次推理输出token数 · 技术路线
It uses about 13k output tokens on average
tags: 好数字 · → daily
value: qty=13k · date=2026-05-01
📷 原图
🔗 因果传导 (causal map)
*GPT-5.5 在产业链上的传导关系. 边是群里/卖方陈述的因果 (非 AI 推断), 数字 = 几条 atom 支撑. 点 atom 溯源.*
↓ 下游·近期陈述 (1)
- 🟢 NVDA · 需求拉动 · 1 次陈述
明确提及 NVDA GB200/GB300 硬件支持
🔮 前瞻触发器 (forward triggers)
*这些 atom 指向可能 reprice 的前瞻事件 (无精确日期, 仅类型). 配合上面叙事状态看「上膛」程度.*
产品 (4 · ▲4/▼0)
- 🟢 2026-04-23 · 主打自主性,强化Agentic AI叙事
GPT-5.5 主打自主性,Agentic AI 叙事强化
- 🟢 2026-04-23 · 在模糊指令下的执行力
OpenAI 强调其在模糊指令下的执行力
- 🟢 2026-04-23 · 基于NVDA GB200/GB300 NVL72系统共同设计训练
明确提及 NVDA GB200/GB300 硬件支持
- 🟢 2026-04-23 · 能处理混乱多部分任务并自主规划、使用工具、检查工作
OpenAI 称 GPT-5.5 能够处理混乱的多部分任务,自主规划、使用工具并检查工作
🟦 客观事实 (facts) (30)
- <a id="atom-cebcf562f6ade59a"></a>🟦 2026-06-10
卖·MS Tom Wigg · GDPval-AA Leaderboard Elo score
GPT-5.5 (xhigh) achieved an Elo score of 1769 on the GDPval-AA Leaderboard.
tags: 好数字·好信源 · → daily
value: qty=1769
- <a id="atom-7f081e1fdee959ae"></a>🟦 2026-06-19
X·@zerohedge ⭐ · API 混合价格 · 价格动态/技术路线
GPT-5.5 (xhigh) 混合价格: $4.4/M tokens
tags: 好数字·好信源 · → daily
value: qty=$4.4/M tokens · date=2026
[图: 2026年中美主流AI模型API价格及智能指数对比表 — Claude Fable 5 混合价格: $7.7/M tokens; GPT-5.5 (xhigh) 混合价格: $4.4/M tokens; Gemini 3.1 Pro Preview 混合价格: $1.7/M tokens; Qwen3.7 Max 混合价格: $1.4/M token
📷 原图
- <a id="atom-1efb53818e953496"></a>🟦 2026-06-16
X·@morqon ⭐ · 训练完成日期 · 产品发布
on march 24, the information reported that GPT-5.5 had finished training
tags: 好数字·好信源 · → daily
value: date=2026-03-24
- <a id="atom-35da940966c6332c"></a>🟦 2026-05-13
X·@bioshok3 ⭐ · 在The Last Ones网络靶场解决次数 · 模型能力
GPT-5.5は「The Last Ones」を10回の試行のうち3回で解決。
tags: 好数字·好信源 · → daily
value: qty=10回试行的3回 · date=2026-05-13
📷 原图
- <a id="atom-933bb73b1e2da7fb"></a>🟦 2026-06-09
X·@OpenAIDevs · 替代OCR流程 · 技术应用/产品发布
23,000+ ChinaRxiv papers are now freely available with more complete English translations after one developer replaced a complex OCR pipeline with GPT‑5.5.
tags: 好数字·好信源 · → daily
value: qty=23,000+ · date=2026-06-09
- <a id="atom-dae2c5f2bc774789"></a>🟦 2026-05-14
X·@mweinbach · pass rate dropped 6% · 模型性能
Pass rate dropped 6% today for GPT 5.5 from Margin Labs, something def changed
tags: 好数字·好信源 · → daily
value: qty=6% · direction=down
📷 原图
- <a id="atom-175051a0606fcd7c"></a>🟦 2026-05-13
X·@chatgpt21 · publicly benchmarked token cap · 模型发布/技术路线
This is the first time horizon snapshot we have of GPT 5.5 that’s been publicly benchmarked at least at the 2.5 million token cap.
tags: 好数字·好信源 · → daily
value: qty=2.5 million · date=
📷 原图
- <a id="atom-510aff737b15b90a"></a>🟦 2026-05-05
X·@bioshok3 · 推理成本对比 Mythos · 价格动态/竞争格局
GPT-5.5がMythosよりも約4~5倍安価だ
tags: 好数字·好信源 · → daily
value: qty=4~5 times cheaper · direction=lower
- <a id="atom-1f7f62614e8dee28"></a>🟦 2026-05-04
X·@OpenRouter · 成本变化 · 成本变化/模型发布
We analyzed GPT 5.5 vs GPT 5.4 and found that costs increased between 49-92%.
tags: 好数字·好信源 · → daily
value: qty=49-92%
📷 原图
- <a id="atom-7fd98006220ca9e6"></a>🟦 2026-04-30
X·@bioshok3 · 英国AISIサイバー攻撃テストでの成績 · 技术路线
gpt-5.5は英国AISIサイバー攻撃テストでmythosと同等。
tags: 好数字·好信源 · → daily
value: qty=平均パス率71.4%(±8.0%) · comparison=Claude Mythosと同等
- <a id="atom-6401211c3bf985eb"></a>🟦 2026-04-30
X·@bioshok3 · タスク解決時間 · 技术路线/产品发布
GPT-5.5 solved a task that takes a human expert ~12 hours in under 11 minutes at
tags: 好数字·好信源 · → daily
value: qty=人間の専門家が約12時間かかるタスクを11分未満で解決
- <a id="atom-ddf7bb596ff47c24"></a>🟦 2026-04-24
X·@hsu_steve · API推理输入价格 · 定价
GPT-5.5输入价格: $5
tags: 好数字·好信源 · → daily
value: qty=$5
[图: 各大主流AI大模型(如DeepSeek V4 Pro、GPT-5.5等)的API推理成本/定价对比表及效率分析 — DeepSeek V4 Pro输入价格: $1.74; DeepSeek V4 Pro输出价格: $3.48; GPT-5.5输入价格: $5; GPT-5.5输出价格: $30]
📷 原图
- <a id="atom-60041c2d440d1aef"></a>🟦 2026-04-24
X·@hsu_steve · API推理输出价格 · 定价
GPT-5.5输出价格: $30
tags: 好数字·好信源 · → daily
value: qty=$30
[图: 各大主流AI大模型(如DeepSeek V4 Pro、GPT-5.5等)的API推理成本/定价对比表及效率分析 — DeepSeek V4 Pro输入价格: $1.74; DeepSeek V4 Pro输出价格: $3.48; GPT-5.5输入价格: $5; GPT-5.5输出价格: $30]
📷 原图
- <a id="atom-67b81aaeb5caef02"></a>🟦 2026-05-04
X·@benedictk_ [老化中] · brokenarxiv 评分 · 模型发布/技术路线_
brokenarxiv scores are in. gpt 5.5 an actual skeptical model that thinks.
tags: 好思考·好信源 · → daily
📷 原图
- <a id="atom-d87981a9a058abc8"></a>🟦 2026-04-23
X·@ArtificialAnlys · Intelligence Index 得分 (medium) · 模型评测/盈利能力
GPT-5.5 (medium) scores the same as Claude Opus 4.7 (max) on our Intelligence Index at one quarter of the cost (~$1,200 vs $4,800)
tags: 好数字·好观点 · → daily
value: qty=same as Claude Opus 4.7 (max)
[图: 展示各主流大语言模型在Artificial Analysis Intelligence Index v4.0评估中得分的数据柱状图 — GPT-5.5 (xhigh): 60; Claude Opus 4.7 (max): 57; Gemini 3.1 Pro Preview: 57; GPT-5.4 (xhigh): 57; Kimi K2.6:
📷 原图
- <a id="atom-0625b900afa5c911"></a>🟦 2026-06-25
卖·MS Tom Wigg · Agentic coding (SWE-Bench Pro) score
GPT 5.5 achieved a score of 58.6% on the Agentic coding (SWE-Bench Pro) benchmark.
tags: 好数字 · → daily
value: qty=58.6%
- <a id="atom-0848c55bb02de39c"></a>🟦 2026-06-25
卖·MS Tom Wigg · Agentic coding (FrontierCode (Diamond)) score
GPT 5.5 achieved a score of 5.7% on the Agentic coding (FrontierCode (Diamond)) benchmark.
tags: 好数字 · → daily
value: qty=5.7%
- <a id="atom-a13c885108480c0c"></a>🟦 2026-06-25
卖·MS Tom Wigg · Knowledge work (GDPval-AA) score
GPT 5.5 achieved a score of 1769 on the Knowledge work (GDPval-AA) benchmark.
tags: 好数字 · → daily
value: qty=1769
- <a id="atom-3daa7bdedb2fccc6"></a>🟦 2026-06-25
卖·MS Tom Wigg · Knowledge work vision (GDP.pdf) score
GPT 5.5 achieved a score of 24.9% on the Knowledge work vision (GDP.pdf) benchmark.
tags: 好数字 · → daily
value: qty=24.9%
- <a id="atom-56479a1710f19e96"></a>🟦 2026-06-25
卖·MS Tom Wigg · Spatial reasoning (Blueprint-Bench 2) score
GPT 5.5 achieved a score of 36.2% on the Spatial reasoning (Blueprint-Bench 2) benchmark.
tags: 好数字 · → daily
value: qty=36.2%
- <a id="atom-726c36df55004a2e"></a>🟦 2026-06-25
卖·MS Tom Wigg · Tool use (AutomationBench) score
GPT 5.5 achieved a score of 12.9% on the Tool use (AutomationBench) benchmark.
tags: 好数字 · → daily
value: qty=12.9%
- <a id="atom-c2d79d3f6abcab89"></a>🟦 2026-06-25
卖·MS Tom Wigg · Computer use (OSWorld-Verified) score
GPT 5.5 achieved a score of 78.7% on the Computer use (OSWorld-Verified) benchmark.
tags: 好数字 · → daily
value: qty=78.7%
- <a id="atom-07baacf5e23c0cc9"></a>🟦 2026-06-25
卖·MS Tom Wigg · Legal (Legal Agent Benchmark) score
GPT 5.5 achieved a score of 2.1% on the Legal (Legal Agent Benchmark) benchmark.
tags: 好数字 · → daily
value: qty=2.1%
- <a id="atom-8ee941183aea7ce1"></a>🟦 2026-06-25
卖·MS Tom Wigg · Multidisciplinary reasoning (Humanity's Last Exam, no tools) score
GPT 5.5 achieved a score of 41.4% on the Multidisciplinary reasoning (Humanity's Last Exam, no tools) benchmark.
tags: 好数字 · → daily
value: qty=41.4%
- <a id="atom-1c9c050acc5529e8"></a>🟦 2026-06-25
卖·MS Tom Wigg · Multidisciplinary reasoning (Humanity's Last Exam, with tools) score
GPT 5.5 achieved a score of 52.2% on the Multidisciplinary reasoning (Humanity's Last Exam, with tools) benchmark.
tags: 好数字 · → daily
value: qty=52.2%
- <a id="atom-8627d69815c91918"></a>🟦 2026-06-25
卖·MS Tom Wigg · Agentic coding (Terminal-Bench 2.1) score
GPT 5.5 achieved a score of 83.4% on the Agentic coding (Terminal-Bench 2.1) benchmark.
tags: 好数字 · → daily
value: qty=83.4%
- <a id="atom-080a03bd7132fab5"></a>🟦 2026-06-25
卖·MS Tom Wigg · Cybersecurity (ExploitBench (Cap%)) score
GPT 5.5 achieved a score of 34.0% on the Cybersecurity (ExploitBench (Cap%)) benchmark.
tags: 好数字 · → daily
value: qty=34.0%
- <a id="atom-b7dedc7f3d39212e"></a>🟦 2026-06-25
卖·MS Tom Wigg · Health (HealthBench Professional) score
GPT 5.5 achieved a score of 51.8% on the Health (HealthBench Professional) benchmark.
tags: 好数字 · → daily
value: qty=51.8%
- <a id="atom-8759e2945f44d2f1"></a>🟦 2026-06-12
卖·MS Tom Wigg · AI output price per million tokens
GPT-5.5's AI output price per million tokens is approximately $30.
tags: 好数字 · → daily
value: qty=30 · date=2026-06-12
- <a id="atom-fbc55111324649f8"></a>🟦 2026-06-10
卖·MS Tom Wigg · Agentic coding SWE-Bench Pro score
GPT 5.5 scored 58.6% on Agentic coding SWE-Bench Pro.
tags: 好数字 · → daily
value: qty=58.6%
🟥 多头 takes (bullish) (30)
- <a id="atom-37cee7cb9cd361b8"></a>🟥 2026-07-02
X·@scaling01 · higher ceiling than Opus 4.8 · 竞争格局
GPT-5.5 seems to have a higher ceiling than Opus 4.8
tags: 好观点·好信源 · → daily
- <a id="atom-deaf340c18398c5c"></a>🟥 2026-05-02
X·@scaling01 ⭐ · 超过10M tokens后性能激增 · 模型性能
after that they go absolutely ballistic
tags: 好观点·好思考 · → daily
value: qty=10M · date=2026-05-02
📷 原图
- <a id="atom-5dc7ebade442fadf"></a>🟥 2026-05-27
X·@scaling01 · cyber-capabilities time horizon · 技术路线
GPT-5.5 reaches the same time horizon on cyber-capabilities as GPT-5.3-Codex with less than 1/4th the token budget
tags: 好数字 · → daily
value: qty=same as GPT-5.3-Codex · date=time horizon unspecified · direction=neutral
📷 原图
- <a id="atom-6addd94890dd4cd5"></a>🟥 2026-05-05
X·@scaling01 · progress rate after step 11 · 模型性能
Weaker models tend to flatten out, while Mythos Preview and GPT-5.5 continue progressing at a faster pace, causing the gap at later milestones to widen.
tags: 好数字 · → daily
📷 原图
- <a id="atom-037cc8b87df4f395"></a>🟥 2026-07-01
X·@JakeABoggs · 在 HieroglyphBench 得分与 Anthropic Fable 5 相当 · 技术路线/模型发布
effectively ties with GPT-5.5 on HieroglyphBench
tags: 好观点·好思考 · → daily
📷 原图
- <a id="atom-e9d7ab346dfcb33e"></a>🟥 2026-05-19
X·@yacineMTB · 能力描述 · 技术路线/模型发布
Gpt 5.5 is a machine that liberates good math from shitty researcher code
tags: 好思考·好观点 · → daily
- <a id="atom-51f124c7c6e38e77"></a>🟥 2026-05-15
X·@mweinbach · performance issue due to bug · 产品质量/技术路线
oh look it was a bug and they fixed it
tags: 好观点·好思考 · → daily
- <a id="atom-2cf9fdae0e80a518"></a>🟥 2026-05-27
X·@scaling01 · 最近结果令人震惊 · 模型发布
I have been absolutely one-shot by recent GPT-5.5 results
tags: 好思考 · → daily
- <a id="atom-bbba8ac83ef3c9b2"></a>🟥 2026-05-05
X·@deredleritt3r · TLO 性能未达上限 · 技术路线
After 100 million tokens, performance was still going up. What we're seeing here is not the capability ceiling.
tags: 好思考 · → daily
📷 原图
- <a id="atom-74bee964055a89ee"></a>🟥 2026-05-01
X·@OpenAI [老化中] · 模型推出后表现 · 模型发布
One week since the launch of GPT-5.5, and it’s already our strongest model launch yet.
tags: 好思考 · → daily
value: qty=strongest model launch yet · date=one week since launch
- <a id="atom-0f1f6c7489d3bcf5"></a>🟥 2026-04-25
X·@scaling01 [老化中] · 推测身份 · 模型发布
gpt-5.5 might be him him
tags: 好思考 · → daily
- <a id="atom-a711284eaa02a00c"></a>🟥 2026-05-15
X·@ArtificialAnlys · GDPval-AA head-to-head comparison win rate against Claude 4 Sonnet · 模型性能
GPT-5.5 is expected to win ~98% of head-to-head comparisons on realistic work outputs against Claude 4 Sonnet, the leading model in GDPval-AA a year ago
tags: 好数字 · → daily
value: qty=~98% · date=May 2026
[图: 展示GDPval-AA评估中关于库存事件分析的数据面板和供应商事件数量图表 — 总解决成本: $82,567; 平均单次事件成本: $2,014; 最高单次事件成本: $12,000; 前三家供应商成本占比: 88%; 有效事件记录数: 50]
📷 原图
- <a id="atom-221d0ff58424ae7a"></a>🟥 2026-05-14
X·@MechanizeWork · Game Boy Advance emulator 表现 · 模型评测
GPT-5.5's emulator runs games best
tags: 好数字 · → daily
- <a id="atom-2968b4fee0b27cca"></a>🟥 2026-05-14
X·@MechanizeWork · 24小时编码性能 vs 时间 · 模型评测
GPT-5.5 performs best, with a late gain in the final hours
tags: 好数字 · → daily
📷 原图
- <a id="atom-b62f5de1dce5323b"></a>🟥 2026-06-12
X·@zephyr_z9 · post-training quality · 模型发布
5.5 has really good post training
tags: 好观点 · → daily
- <a id="atom-c5ac86ba333908cd"></a>🟥 2026-06-05
X·@scaling01 [老旧] · SWE-Marathon 排名表现 · 基准表现
very interesting to see GPT-5.5 on top
tags: 好观点 · → daily
📷 原图
- <a id="atom-7b4e6200cef35e43"></a>🟥 2026-05-23
X·@scaling01 · 模型能力侧重 · 技术路线/产品发布
GPT-5.5 is obviously better at math
tags: 好观点 · → daily
value: direction=优势 · comparative=Mythos
- <a id="atom-83dd4d0813795f5f"></a>🟥 2026-05-22
X·@scaling01 · ALE-Bench 表现 · 模型表现
GPT-5.5-xhigh is quite ridiculous
tags: 好观点 · → daily
📷 原图
- <a id="atom-00766419abd3c05b"></a>🟥 2026-05-12
X·@scaling01 [老化中] · 性能对比 · 模型发布/竞争格局
GPT-5.5-xhigh IQ mogged Opus 4.7
tags: 好观点 · → daily
value: direction=高于 · comparison_object=Opus 4.7
- <a id="atom-e5c824483f63eecf"></a>🟥 2026-05-10
X·@sama [老化中] · performance assessment · 模型发布/技术路线
disagree but it's pretty good
tags: 好观点 · → daily
- <a id="atom-d7fa9927ed09c2c1"></a>🟥 2026-05-05
X·@scaling01 · token efficiency before step 11 · 模型性能
One thing we can say is that before the step-11 gate, stronger models seem more token-efficient.
tags: 好观点 · → daily
📷 原图
- <a id="atom-449bc90ed4d7f695"></a>🟥 2026-05-01
X·@scaling01 [老化中] · function calling performance · 模型能力
GPT-5.5 has absolutely maxxed out function calling
tags: 好观点 · → daily
value: direction=maxxed_out
- <a id="atom-8d9770c2038fde5b"></a>🟥 2026-04-25
X·@gdb [老化中] · 能力上限 · 技术路线/模型发布
GPT-5.5 raises the ceiling of ambition for what you can do with AI
tags: 好思考 · → daily
- <a id="atom-45b61ef55b6ab36e"></a>🟥 2026-04-23
X·@scaling01 [老化中] · dominates cost-performance frontier on Artificial Analysis Index · 技术路线/价格动态
The GPT-5.5 model family completely dominates the cost-performance frontier on the Artificial Analysis Index
tags: 好观点 · → daily
📷 原图
- <a id="atom-d7e69a09e3a1596b"></a>🟥 2026-04-23
群 [老化中] · 主打自主性,强化Agentic AI叙事 · 产品发布/Agentic AI
GPT-5.5 主打自主性,Agentic AI 叙事强化
tags: 好观点 · → daily
- <a id="atom-fe3f81f6468d05e2"></a>🟥 2026-05-27
X·@koltregaskes · DeepSWE 得分高于 Claude Sonnet · 技术路线
That difference is substantial.
tags: 好思考 · → daily
value: direction=高于
📷 原图
- <a id="atom-d001c8aca53e3be5"></a>🟥 2026-05-09
X·@rezoundous [老化中] · 被信任程度高于 Opus 4.7 · 模型发布/技术路线
Can't explain it, but I trust GPT-5.5 more than Opus 4.7 right now.
tags: 好思考 · → daily
value: date=2026-05-09
- <a id="atom-daaad3b642b4b8f7"></a>🟥 2026-05-07
X·@steipete [老化中] · 用户使用体验 · 技术路线
/goal + GPT 5.5 is amazing. I can now plan really extensive refactors with e2e tests and it just works.
tags: 好观点 · → daily
📷 原图
- <a id="atom-051600dbc18021bc"></a>🟥 2026-05-05
X·@michpokrass [老化中] · factuality 改进 · 模型发布/产品发布
we shipped gpt-5.5 instant today to chat; it's rolling out over the next couple days to everyone. for this model, we focused on factuality, crushing hacks, and improving the baseline intelligence. 5.5 is a pretty big step forward on all three
tags: 好观点 · → daily
- <a id="atom-66329a1662e07307"></a>🟥 2026-05-05
X·@michpokrass [老化中] · intelligence 提升 · 模型发布/技术路线
it is much smarter, significantly less likely to hallucinate, and just more delightful to talk to
tags: 好观点 · → daily
🟥 空头 takes (bearish) (20)
- <a id="atom-f4bc29e4a802f1ae"></a>🟥 2026-05-01
X·@justanotherlaw · 参数数量估值修正 · 模型参数
GPT 5.5: 9.7T -> 1.5T
tags: 好数字·好思考 · → daily
value: qty=1.5T
📷 原图
- <a id="atom-64f9483188abf756"></a>🟥 2026-05-15
X·@mweinbach [老旧] · model performance decreased · 模型发布/产品质量
GPT 5.5 is worse
tags: 好观点·好思考 · → daily
[图: Margin Lab 平台上展示每日输入和输出 Token 使用量趋势的数据面板图表 — 2026年5月12日输入Token: 4.1M; 输入Token峰值: 约5.5M; 输出Token末端值: 约700.0K]
📷 原图
- <a id="atom-a25209e599232cc8"></a>🟥 2026-05-07
X·@spicey_lemonade [老化中] · fails at simple theory of mind test · 技术路线/模型发布
GPT 5.5 fails at simple theory of mind. This is one of the most important tests for sentience.
tags: 好思考·好问题 · → daily
📷 原图
- <a id="atom-5266e7d58c117d7b"></a>🟥 2026-05-03
X·@orenbahari [老化中] · cost premium vs open model for long interactive tasks · 产品发布/价格动态
GPT 5.5 is about 15% more expensive for smart broader knowledge
tags: 好数字 · → daily
value: qty=15%
- <a id="atom-51bc98546eb6bb09"></a>🟥 2026-06-05
X·@scaling01 · SWE-Marathon 排名表现 · 基准表现
but GPT-5.5 is weird
tags: 好观点 · → daily
- <a id="atom-9a269d95bf757620"></a>🟥 2026-05-31
X·@scaling01 · progresses slower than Opus 4.8 on this eval · 模型性能
Opus 4.8 also progresses much faster than GPT-5.5 on this eval
tags: 好观点 · → daily
- <a id="atom-d80ed66be0e0a226"></a>🟥 2026-05-14
X·@mweinbach · capability regression to GPT 5.4 · 模型性能
Feels like we're back to GPT 5.4
tags: 好思考 · → daily
📷 原图
- <a id="atom-f0bbc3bc8652bfe0"></a>🟥 2026-05-14
X·@mweinbach · goal completion was faked · 模型性能
Yesterday i noticed it finished a goal that was 100% NOT finished just randomly called the end goal tool
tags: 好思考 · → daily
- <a id="atom-2db546870a76a0ef"></a>🟥 2026-06-29
X·@sorianmaran · 使用体验类比 · 模型发布
Everytime I have to switch to GPT 5.5 from Opus 4.8 ... it's like talking to my MBA associate who always says YOU GOT IT! And works all night and all weekend and forces all his subordinates to do the same and then didnt even realize one of their core assumptions was wrong.
tags: 好观点 · → daily
- <a id="atom-7d189ab9d76636b1"></a>🟥 2026-06-28
X·@9hills · PRD实现不可用 · 模型能力
GPT-5.5:不可用,懒惰,大量未实现功能和罗列性质的页面设计。建议Human-in-loop生成详细功能点后再执行。
tags: 好观点 · → daily
- <a id="atom-a2a8e5004add8ccd"></a>🟥 2026-06-28
X·@Xianbao_QIAN · 近期降智严重 · 模型能力
5.5 最近的降智太厉害了。感觉已经快不可用了。还得是开源模型
tags: 好观点 · → daily
- <a id="atom-ed7ffc34b462ebaf"></a>🟥 2026-06-23
X·@yacineMTB [老化中] · 任务表现 · 模型表现
gpt 5.5 is failing at this task
tags: 好观点 · → daily
📷 原图
- <a id="atom-88b6e3bd1aa0cbf2"></a>🟥 2026-06-22
X·@Hangsiin · GPT-5.5 无法用短 prompt 产生 Fable 那种感觉 · 技术路线
I don’t think GPT-5.5 could produce this kind of feel with such a short prompt.
tags: 好观点 · → daily
- <a id="atom-319d17427b2cb9af"></a>🟥 2026-06-19
X·@dhtikna · 性能接近 Fable 5 · 模型发布/竞争格局
gpt 5.5 is so close to fable 5
tags: 好观点 · → daily
- <a id="atom-63f2e63f335153c4"></a>🟥 2026-06-19
X·@tokenbender · feels smaller, test-time-compute pilled and less knowledgeable · 模型发布/技术路线
after looking at all the in-the-weights results i can say gpt 5.5 feels smaller, test-time-compute pilled and less knowledgeable.
tags: 好观点 · → daily
- <a id="atom-682ce4551f96d3bc"></a>🟥 2026-05-28
X·@flowersslop · 实时策略游戏可玩性 · 产品性能
Tried using GPT-5.5 to play an RTS against the easy AI. Its basically infeasible right now. way too latency/token constrained.
tags: 好观点 · → daily
value: direction=不可行
- <a id="atom-c98a04fe13cb08de"></a>🟥 2026-05-28
X·@flowersslop · 实时反应能力不足 · 产品性能
can’t handle environments that require sub-100ms reactions and continuous multi-action control.
tags: 好观点 · → daily
value: direction=无法处理
- <a id="atom-e397de4f9af46fd4"></a>🟥 2026-05-09
X·@Yuchenj_UW [老化中] · weirdly weak at frontend · 技术路线
GPT-5.5 is still weirdly weak at frontend.
tags: 好观点 · → daily
- <a id="atom-13ca13684e8f9679"></a>🟥 2026-04-24
X·@tokenbender [老化中] · 防御性行为 · 技术路线
GPT-5.5 的防御性行为被夸大,但仍会在更长上下文中发生
tags: 好观点 · → daily
value: qty=超过200K上下文 · direction=longer
- <a id="atom-8222ea29308f2b32"></a>🟥 2026-06-01
X·@KineticElle · model version downgrade · 模型发布
GPT 5.5 Thinking turn into 5.2
tags: 好问题 · → daily
value: direction=down
🟥 中性 takes (neutral) (20)
- <a id="atom-48a360d0951d3e1f"></a>🟥 2026-06-04
X·@scaling01 · 训练规模估计 · 训练规模
my guess is that GPT-5.5 is ~5T
tags: 好数字 · → daily
value: qty=~5T · date=2026-06-04
- <a id="atom-5f7e94864c83a45b"></a>🟥 2026-06-04
X·@scaling01 · 训练时长估计 · 训练规模
base GPT-5.5 is only a ~16 day training run + maybe 2 months of post
tags: 好数字 · → daily
value: qty=~16 days · date=2026-06-04
- <a id="atom-e67bf25e9110a601"></a>🟥 2026-05-05
X·@scaling01 [老化中] · step-11 gate · 模型性能/技术路线
The change after step 11 could be a compounding effect. It could also be that step 11 acts as a gate: once a model gets past it, later progress becomes easier or at least more reachable within the benchmark setup.
tags: 好思考 · → daily
value: direction=pass
📷 原图
- <a id="atom-c9a032b8599b303f"></a>🟥 2026-05-01
X·@chatgpt21 · reported failure: true local effect, false world model · 技术路线
True local effect, false world model
tags: 好数字 · → daily
📷 原图
- <a id="atom-5de0a040e28a7eff"></a>🟥 2026-05-01
X·@chatgpt21 · reported failure: wrong level of abstraction from training data · 技术路线
Wrong level of abstraction from training data
tags: 好数字 · → daily
📷 原图
- <a id="atom-c0c546c25458f4f2"></a>🟥 2026-05-01
X·@chatgpt21 · reported failure: solved the level, didn’t reinforce the reward · 技术路线
Solved the level, didn’t reinforce the reward
tags: 好数字 · → daily
📷 原图
- <a id="atom-a256f97f02935e98"></a>🟥 2026-04-30
X·@ZhihuFrontier [老化中] · 参数规模估计 · 模型规模
GPT-5.5 ≈ 9T
tags: 好数字 · → daily
value: qty=≈9T · date=2026-04
📷 原图
- <a id="atom-8c2c8eb8e878509c"></a>🟥 2026-06-25
X·@flowersslop · 不同场景下能力猜测 · 模型发布
GPT-5.5s guesses for different scenarios in the future
tags: 好观点·好问题 · → daily
📷 原图
- <a id="atom-d7a199c906bd1f04"></a>🟥 2026-06-18
X·@testingcatalog · 对比表现 · 竞争格局
GPT-5.5 is head-to-head with Claude Opus 4.7
tags: 好观点 · → daily
- <a id="atom-e307d8b6ee1ad647"></a>🟥 2026-05-04
X·@flavioAd [老旧] · 开放邀请 · 产品发布
I got invited to the GPT-5.5 party but can’t go because I live on the wrong continent, and now the people who didn’t get invited are getting something nice too
tags: 好信源 · → daily
value: direction=受邀
- <a id="atom-1f3d4a399939c88e"></a>🟥 2026-04-15
X·@chatgpt21 · 发布预期时间 · 模型发布
GPT 5.5 as quoted here will come just unfortunately not this week. the good news is it’s still before the end of April and tentative for next week
tags: 好信源 · → daily
value: date=预计下周(4月底前)
📷 原图
- <a id="atom-3e69e3a3eba3e46f"></a>🟥 2026-06-25
X·@bdsqlsz · strength in code implementation vs long-term planning · 模型对比/技术路线
GPT better suited for implementing code than for long-term planning and complex tasks
tags: 好思考 · → daily
- <a id="atom-7da99c1ba8486fce"></a>🟥 2026-05-07
X·@dhtikna [老化中] · theory of mind test failure is common in text-only settings · 技术路线/模型发布
Even a human make this mistake in text only communication setting. Nothing burger imo
tags: 好思考 · → daily
- <a id="atom-ca6370a7c90b1e08"></a>🟥 2026-05-01
X·@hungjng69679118 [老化中] · 参数量被认为5万亿甚至10万亿级别 · 技术路线
GPT 5.5...被认为是5万亿甚至10万亿级别的参数量
tags: 好思考 · → daily
value: qty=5万亿-10万亿
- <a id="atom-90ad08a729f24a8d"></a>🟥 2026-06-26
X·@BSPK_ · 模型质量评价 · 模型发布
GPT5.5로 돌아오면 뭔가 아쉬운….
tags: 好观点 · → daily
- <a id="atom-8dc301c9c0e7ab96"></a>🟥 2026-06-19
X·@tokenbender · designed to rely on data rather than memory · 技术路线
truly built for "i would not rely on memory, let me check this in <data> instead"
tags: 好观点 · → daily
- <a id="atom-ca5f9614abc47dbe"></a>🟥 2026-06-19
X·@tokenbender · good at exploiting/traversing existing ideas · 技术路线/竞争格局
this would explain why opus 4.8 is actually better at generating diverse ideas whereas gpt 5.5 is good at exploiting/traversing existing ones.
tags: 好观点 · → daily
- <a id="atom-c5f7dcf22fa3c784"></a>🟥 2026-06-13
X·@andrewqu · 与 Opus-4.8 和 Fable-5 在日常工作中的表现差异 · 模型能力对比
a lot of people wouldn’t be able to tell the difference if they were randomly routed between gpt-5.5, opus-4.8, or fable-5 for their day to day work
tags: 好观点 · → daily
- <a id="atom-3876e59af5c44a0d"></a>🟥 2026-06-09
X·@aadityaa_26 · 代码/研究能力 · 合作客户
I was working on designing a new trading system, Gemini, and GPT-5.5 was amazing for research.
tags: 好观点 · → daily
- <a id="atom-b7606dbd6cf714b2"></a>🟥 2026-06-02
X·@steipete · larger model double usage · 模型大小
You probably moved to GPT 5.5 which is a far larger model and takes double usage.
tags: 好观点 · → daily
⏱ 时间轴 (近 20)
- 🟦 2026-07-03 ·
fact · 表现 · → daily
- 🟦 2026-07-02 ·
fact · reward-hacking attempts · → daily
- 🟦 2026-07-02 ·
fact · reward-hacking rate on SWE-Marathon · → daily
- 🟦 2026-07-02 ·
fact · outscore GPT-5 on EBR-bench · → daily
- 🟥 2026-07-02 ·
narrative · higher ceiling than Opus 4.8 · → daily
- 🟦 2026-07-02 ·
fact · B200 fp8 GEMM 性能评分 · → daily
- 🟥 2026-07-02 ·
position · 能力跃升幅度 · → daily
- 🟦 2026-07-01 ·
forecast · 元评估框架下API过滤器拒绝率 · → daily
- 🟦 2026-07-01 ·
fact · Claw-Anything benchmark pass@1 score · → daily
- 🟦 2026-07-01 ·
fact · cost per trial · → daily
- 🟦 2026-07-01 ·
fact · 工具使用回归 · → daily
- 🟥 2026-07-01 ·
fact · 在 HieroglyphBench 得分与 Anthropic Fable 5 相当 · → daily
- 🟥 2026-06-30 ·
narrative · token efficiency compared to Gemini 3.5 Flash · → daily
- 🟥 2026-06-30 ·
narrative · 性能对比 · → daily
- 🟦 2026-06-30 ·
fact · 令牌消耗 · → daily
- 🟥 2026-06-30 ·
position · 当前最佳模型 · → daily
- 🟦 2026-06-29 ·
fact · ARC-AGI 得分 · → daily
- 🟥 2026-06-29 ·
narrative · 使用体验类比 · → daily
- 🟦 2026-06-29 ·
fact · SWE-bench Pro评分 · → daily
- 🟥 2026-06-28 ·
position · PRD实现不可用 · → daily
← 实体目录 · 系统日志