以下部分引用 AI 总结,现阶段 AI 仍然有幻觉。请以内容的原文为准。
数据新鲜度群聊 07-14 ✓卖方 07-14 ✓Wrap 断档15天 ✕News 07-15 ✓X 断档12天 ✕页面生成 07-15
GPT-5.5 Pro
45 atoms · 跨 22 天 · 首见 2026-04-25 · 最近 2026-07-02
三色: 🟦 fact 34 · 🟥 take 11 · stance ▲22/▼6/◆6
叙事状态: 🍂 fading (退潮) · 💤 沉寂 · 策展近14d 0 atom · 🐦 X 近14d 3
状态只算低频策展源 (群/卖方); X firehose 仅作背景音量
来源: X 45
时态: fresh:41 · aging:4
标签: 好数字:27 · 好观点:11 · 好信源:6 · 好思考:4
🟨 AI 综合 · junior analyst 概览
展开 AI 综合 (灰色 · 非市场结论 · 点击数字溯源到原 atom)
model: deepseek-chat · 2026-07-12 · 默认折叠
GPT-5.5 Pro 在数学与科学推理上展现出显著优势,包括在开放数学问题上远超其他模型
¹,以及能够发现学术论文中的重大错误并提供建设性反馈
¹。它目前是 Epoch Capabilities Index 的领先者
¹,在 FrontierMath Tier 4 和 CritPt 基准测试中均获得最高分
¹¹。然而,在日常实用性方面存在分歧:其输出曾被形容为“糟糕”的博客草稿
¹,且被认为对于会议摘要等简单任务过于大材小用
¹。更关键的是,它在可靠医疗使用方面仍被认为不足
¹,并且在 ECI 得分上仍略低于某个竞品
¹。因此,模型在高阶创新任务与基础应用场景之间存在明显的表现落差。
🧭 拥挤度 (一人一票): ▲ 4 位作者 (KOL4) vs ▼ 4 位作者 (KOL4)
📄 广泛报道的事实 · 1 个
展开 (多人转述同一事实/数字 — 确认度高, **非独立观点**, 不标多空)
- VoxelBench 排名第 1 — 2 源报道 (含2快讯) ·
X
💢 核心分歧 (1 轴)
1. 成本与速度优势 vs 记忆与可靠性隐患
*topic: 技术路线/产品发布/推理效率/模型能力/学术研究 · 18 bull vs 3 bear*
🟢 bullish 侧:
- <a id="atom-c478720b8c59b746"></a>2026-04-28
X·@bioshok3 · ECI score · 模型发布/技术路线 +同日2条
GPT-5.5 Proがエポック能力指数(ECI)で159という新たな高得点を達成
tags: 好数字 · → daily
value: qty=159 · date=2026-04-29
- <a id="atom-237e9ee0dec9b423"></a>2026-04-29
X·@emollick [老化中] · task performance comparison · 产品发布/竞争格局
But here is GPT-5.5 Pro for comparison. Sadly, it took this assignment very seriously and ethically and thus was no fun at all.
tags: 好思考 · → daily
📷 原图
- <a id="atom-3086ae61ba205f72"></a>2026-04-30
X·@ArtificialAnlys · CritPt评估成本与token使用降低60% · 技术路线
60% lower cost and token use in our frontier science eval, CritPt
tags: 好数字 · → daily
value: qty=60% · date=2026-04-30
📷 原图
- <a id="atom-0dc1fa2ea70bfb3e"></a>2026-05-03
X·@marmaduke091 · VoxelBench 排名第 1 · 模型发布/技术路线
GPT-5.5 Pro takes the #1 spot on VoxelBench!
tags: 好数字 · → daily
value: qty=#1
📷 原图
- <a id="atom-a6c630fa6a285786"></a>2026-05-08
X·@prz_chojecki [老化中] · open math problem solving capability · 技术路线/模型发布
From purely math perspective GPT-5.5 Pro is way better at approaching open math problems than any other model (released or not)
tags: 好观点·好信源 · → daily
🔴 bearish 侧:
- <a id="atom-3f36bd7a15d1fc23"></a>2026-05-24
X·@aramh · 推理耗时变化 · 产品发布/技术路线 +同日1条
a response used to take 40 minutes at minimum and now it never takes more than 4 minutes
tags: 好数字 · → daily
value: direction=decrease
- <a id="atom-6e973916237d7d7c"></a>2026-06-07
X·@sundeep · 能力范围评价 · 技术路线/产品发布
meeting summaries and calendar updates don’t require GPT-5.5 Pro.
tags: 好观点 · → daily
value: direction=not required for meeting summaries and calendar updates
📷 原图
展开 4 条中性
- <a id="atom-76327a54dd00ed98"></a>2026-06-12
X·@quantum_aram · unit distance 求解能力 · 技术路线
GPT5.5 Pro solves unit distance already.
tags: 好数字 · → daily
- <a id="atom-80323ba92b49197c"></a>2026-06-12
X·@SebastienBubeck · unit distance 不是一次求解 · 技术路线
not one-shot though; but soon
tags: 好观点 · → daily
value: direction=not one-shot
- <a id="atom-4679843ab6cbe498"></a>2026-06-20
X·@emollick ⭐ · 能够分析数据并扩展论文论证 · AI能力/学术研究
It found new data, analyzed it, created reproducible files, extended the key argument...
tags: 好数字·好信源 · → daily
📷 原图
- <a id="atom-2fa7cb7d68f55a56"></a>2026-06-27
X·@yishan · radiology image interpretation benchmark score vs prior best · 模型能力/基准测试
79/100 vs 69/100
tags: 好数字 · → daily
value: qty=79/100 vs 69/100
🟦 客观事实 (facts) (30)
- <a id="atom-4679843ab6cbe498"></a>🟦 2026-06-20
X·@emollick ⭐ · 能够分析数据并扩展论文论证 · AI能力/学术研究
It found new data, analyzed it, created reproducible files, extended the key argument...
tags: 好数字·好信源 · → daily
📷 原图
- <a id="atom-6d356b78736588ce"></a>🟦 2026-07-02
X·@EpochAIResearch · current leader on Epoch Capabilities Index · 模型发布时间
the most recent being OpenAI’s GPT-5.5 Pro
tags: 好数字 · → daily
value: date=most recent
- <a id="atom-887409801f58b01d"></a>🟦 2026-05-11
X·@hsu_steve · 在 FrontierMath Tier 4 基准测试得分 · 模型发布
OpenAI's top-tier configuration, GPT-5.5 Pro (XHigh), which scored 39.6%
tags: 好数字 · → daily
value: qty=39.6% · date=2026-05-11
📷 原图
- <a id="atom-07975633238fdc22"></a>🟦 2026-06-27
X·@yishan · radiology image interpretation benchmark score · 模型能力/基准测试
GPT-5.5 Pro did outperform the best models from before (79/100 vs 69/100)
tags: 好数字 · → daily
value: qty=79/100
- <a id="atom-2fa7cb7d68f55a56"></a>🟦 2026-06-27
X·@yishan · radiology image interpretation benchmark score vs prior best · 模型能力/基准测试
79/100 vs 69/100
tags: 好数字 · → daily
value: qty=79/100 vs 69/100
- <a id="atom-82aac528032995f8"></a>🟦 2026-06-17
X·@ArtificialAnlys · CritPt基准得分 · 技术路线
GPT-5.5 Pro tops the benchmark at 30.6% on CritPt
tags: 好数字 · → daily
value: qty=30.6% · date=2026-06-17
📷 原图
- <a id="atom-863b9b8d841e9f66"></a>🟦 2026-06-15
X·@EpochAIResearch · Epoch Capabilities Index 得分 · 模型能力/测评结果
beats out GPT-5.5 Pro by 1 point
tags: 好数字 · → daily
value: qty=160
📷 原图
- <a id="atom-76327a54dd00ed98"></a>🟦 2026-06-12
X·@quantum_aram · unit distance 求解能力 · 技术路线
GPT5.5 Pro solves unit distance already.
tags: 好数字 · → daily
- <a id="atom-e72437ee276f62cd"></a>🟦 2026-06-06
X·@chamath · 月度推理成本 · 价格动态
GPT-5.5 Pro: ~$105,000
tags: 好数字 · → daily
value: qty=$105,000 · date=每月
- <a id="atom-6f30891fca1d888d"></a>🟦 2026-05-30
X·@garybasin · 定价 · 价格动态
Because it costs a dollar a minute.
tags: 好数字 · → daily
value: qty=$1/分钟 · date=2026-05-30
- <a id="atom-3f36bd7a15d1fc23"></a>🟦 2026-05-24
X·@aramh · 推理耗时变化 · 产品发布/技术路线
a response used to take 40 minutes at minimum and now it never takes more than 4 minutes
tags: 好数字 · → daily
value: direction=decrease
- <a id="atom-74ae78bdbe8db75c"></a>🟦 2026-05-19
X·@teortaxesTex · 解决 #6 ArXivMath 问题的能力 · 模型性能
Recently got cracked with GPT 5.5 Pro
tags: 好数字 · → daily
📷 原图
- <a id="atom-d3bb6f4d1c18f181"></a>🟦 2026-05-18
X·@R2Cdev_ · API 速度提升倍数 · 技术路线/模型发布
GPT-5.5 Pro is like 1000x faster in the API.
tags: 好数字 · → daily
value: qty=1000x
- <a id="atom-a894dd3e4aed5af3"></a>🟦 2026-05-04
X·@voxelbench · VoxelBench排名 · 模型性能
GPT-5.5 Pro has ranked 1st on VoxelBench
tags: 好数字 · → daily
value: qty=第1名 · date=2026-05-04
📷 原图
- <a id="atom-90d4a14fdc9b67a7"></a>🟦 2026-05-04
X·@voxelbench · Elo分数差 · 模型性能
It scores 100+ Elo points higher than GPT-5.5!
tags: 好数字 · → daily
value: qty=100+ Elo points higher
📷 原图
- <a id="atom-0dc1fa2ea70bfb3e"></a>🟦 2026-05-03
X·@marmaduke091 · VoxelBench 排名第 1 · 模型发布/技术路线
GPT-5.5 Pro takes the #1 spot on VoxelBench!
tags: 好数字 · → daily
value: qty=#1
📷 原图
- <a id="atom-b4817724449fe702"></a>🟦 2026-04-30
X·@ArtificialAnlys · 运行CritPt使用的token数少于GPT-5.4 Pro一半 · 技术路线
uses less than half the tokens to run CritPt
tags: 好数字 · → daily
value: qty=less than half the tokens · date=2026-04-30
[图: 各大AI模型在CritPt基准测试中的Token使用量对比图 — Gemini 3 Deep Think Token使用量: 63M; GPT-5.4 Pro Token使用量: 6M; GPT-5.5 Pro Token使用量: 2.5M]
📷 原图
- <a id="atom-ec014ab4d7f6323b"></a>🟦 2026-04-30
X·@ArtificialAnlys · CritPt测试Token使用量 · 模型发布
GPT-5.5 Pro Token使用量: 2.5M
tags: 好数字 · → daily
value: qty=2.5M · date=2026-04-30
[图: 各大AI模型在CritPt基准测试中的Token使用量对比图 — Gemini 3 Deep Think Token使用量: 63M; GPT-5.4 Pro Token使用量: 6M; GPT-5.5 Pro Token使用量: 2.5M]
📷 原图
- <a id="atom-8f6f209e25fabb27"></a>🟦 2026-04-30
X·@ArtificialAnlys · CritPt得分相比GPT-5.4 Pro小幅提升 · 模型发布
GPT-5.5 Pro achieves a small bump on GPT-5.4 Pro with 60% lower cost and token use in our frontier science eval, CritPt
tags: 好数字 · → daily
value: qty=small bump · date=2026-04-30
📷 原图
- <a id="atom-3086ae61ba205f72"></a>🟦 2026-04-30
X·@ArtificialAnlys · CritPt评估成本与token使用降低60% · 技术路线
60% lower cost and token use in our frontier science eval, CritPt
tags: 好数字 · → daily
value: qty=60% · date=2026-04-30
📷 原图
- <a id="atom-f1e587b4f8c0248c"></a>🟦 2026-04-30
X·@ArtificialAnlys · xhigh配置CritPt得分比GPT-5.4 Pro高0.5个百分点 · 模型发布
GPT-5.5 Pro (xhigh) has surpassed this result by half a percentage point
tags: 好数字 · → daily
value: qty=0.5 percentage point · date=2026-04-30
📷 原图
- <a id="atom-c478720b8c59b746"></a>🟦 2026-04-28
X·@bioshok3 · ECI score · 模型发布/技术路线
GPT-5.5 Proがエポック能力指数(ECI)で159という新たな高得点を達成
tags: 好数字 · → daily
value: qty=159 · date=2026-04-29
- <a id="atom-a7d8075fbb74c3c3"></a>🟦 2026-04-28
X·@bioshok3 · FrontierMath score tier 1-3 · 模型发布/技术路线
FrontierMathでも新記録を樹立し、ティア 1-3で52% (従来 50%)
tags: 好数字 · → daily
value: qty=52% · date=2026-04-29
- <a id="atom-4f8dd9bf50d06ba0"></a>🟦 2026-04-28
X·@bioshok3 · FrontierMath score tier 4 · 模型发布/技术路线
ティア 4で40% (従来 38%) のスコアを会得
tags: 好数字 · → daily
value: qty=40% · date=2026-04-29
- <a id="atom-ce437b914d1ac5cf"></a>🟦 2026-04-25
X·@Hangsiin · KSST test score · 模型评测
This model recorded the same score as its immediate predecessor, GPT-5.4 Pro, placing it in a tie for first place.
tags: 好数字 · → daily
value: qty=same as GPT-5.4 Pro · date=2026-04-25
📷 原图
- <a id="atom-7ccc1a8494385b6d"></a>🟦 2026-04-25
X·@Hangsiin · token efficiency improvement · 推理效率
outputs that took more than 50 minutes to generate with 5.4 Pro were produced in just 10 minutes by 5.5 Pro, while maintaining a similar or even better level of quality.
tags: 好数字 · → daily
value: qty=output time reduced from >50 min to 10 min per sample, similar quality · date=2026-04-25
📷 原图
- <a id="atom-a7a074d6c641f8af"></a>🟦 2026-04-25
X·@Hangsiin · efficiency improvement per sample · 推理效率
5.4 Pro took around 20 to 50 minutes per sample, whereas 5.5 Pro completed each sample in about 5 to 10 minutes
tags: 好数字 · → daily
value: qty=5 to 10 minutes per sample vs 20-50 minutes for 5.4 Pro · date=2026-04-25
📷 原图
- <a id="atom-337cbde4e679c34c"></a>🟦 2026-07-01
X·@prz_chojecki · found counterexamples to Benes Conjecture · 技术路线/产品发布
With GPT-5.5 Pro I found counterexamples to Beneš Conjecture
tags: 好信源 · → daily
📷 原图
- <a id="atom-309959377a610119"></a>🟦 2026-07-01
X·@prz_chojecki · found counterexamples to Shuffle-Exchange Conjecture · 技术路线/产品发布
as well as related Shuffle-Exchange Conjecture in network theory
tags: 好信源 · → daily
📷 原图
- <a id="atom-1c629dfbd7351586"></a>🟦 2026-04-30
X·@ArtificialAnlys · 每token定价与GPT-5.4 Pro相同 · 价格动态
GPT-5.5 Pro has the same per-token pricing to GPT-5.4 Pro
tags: 好信源 · → daily
value: qty=same per-token pricing · date=2026-04-30
[图: 各大AI模型在CritPt基准测试中的Token使用量对比图 — Gemini 3 Deep Think Token使用量: 63M; GPT-5.4 Pro Token使用量: 6M; GPT-5.5 Pro Token使用量: 2.5M]
📷 原图
🟥 多头 takes (bullish) (7)
- <a id="atom-a6c630fa6a285786"></a>🟥 2026-05-08
X·@prz_chojecki [老化中] · open math problem solving capability · 技术路线/模型发布
From purely math perspective GPT-5.5 Pro is way better at approaching open math problems than any other model (released or not)
tags: 好观点·好信源 · → daily
- <a id="atom-1a0bf18d6834a043"></a>🟥 2026-05-28
X·@emollick · 数学证明产出 · 技术路线/模型发布
GPT-5.5 Pro is the model producing many of the novel math proofs, and also the model you should have reviewing any technical or academic paper for flaws.
tags: 好观点·好思考 · → daily
value: direction=多
- <a id="atom-c2a028d25126faf2"></a>🟥 2026-05-30
X·@MParakhin · benefit when used via Pi · 产品发布
Totally agree if you can bring Pro into the mix (through Pi, for example).
tags: 好观点 · → daily
- <a id="atom-237e9ee0dec9b423"></a>🟥 2026-04-29
X·@emollick [老化中] · task performance comparison · 产品发布/竞争格局
But here is GPT-5.5 Pro for comparison. Sadly, it took this assignment very seriously and ethically and thus was no fun at all.
tags: 好思考 · → daily
📷 原图
- <a id="atom-c4534aad3de377ca"></a>🟥 2026-05-28
X·@emollick · 作为审稿人发现错误 · 技术路线
I had to use GPT-5.5 Pro as a reviewer, it spotted one major error & some minor points
tags: 好观点 · → daily
📷 原图
- <a id="atom-c7859bf1a2dcc54e"></a>🟥 2026-05-28
X·@emollick · 发现幻觉结果并提供反馈 · 技术路线
GPT-5.5 found one issue with a hallucinated result, and had other constructive feedback.
tags: 好观点 · → daily
- <a id="atom-2aebbaa40c7a80bf"></a>🟥 2026-05-08
X·@teortaxesTex [老化中] · general capability relative to Gemini · 技术路线
5.5 pro is just generally vastly more capable than any Gemini we've seen
tags: 好观点 · → daily
🟥 空头 takes (bearish) (4)
- <a id="atom-7450003db62b4a34"></a>🟥 2026-04-28
X·@scaling01 [老化中] · ECI score comparison · 性能评估
GPT-5.5 Pro is still a bit below
tags: 好观点 · → daily
value: direction=below
- <a id="atom-6ae9a8ba75836ae9"></a>🟥 2026-06-27
X·@yishan · sufficiency for reliable medical use · 医疗可靠性
did not improve enough to be considered sufficient for reliable medical use
tags: 好思考 · → daily
value: direction=insufficient
- <a id="atom-6e973916237d7d7c"></a>🟥 2026-06-07
X·@sundeep · 能力范围评价 · 技术路线/产品发布
meeting summaries and calendar updates don’t require GPT-5.5 Pro.
tags: 好观点 · → daily
value: direction=not required for meeting summaries and calendar updates
📷 原图
- <a id="atom-1bc47e5699e5acc1"></a>🟥 2026-05-19
X·@xeophon · 写作质量 · 模型表现
gave gpt 5.5 pro ramblings for a blog … its output was so bad that i feel insulted
tags: 好观点 · → daily
⏱ 时间轴 (近 20)
- 🟦 2026-07-02 ·
fact · current leader on Epoch Capabilities Index · → daily
- 🟦 2026-07-01 ·
fact · found counterexamples to Benes Conjecture · → daily
- 🟦 2026-07-01 ·
fact · found counterexamples to Shuffle-Exchange Conjecture · → daily
- 🟦 2026-06-27 ·
fact · radiology image interpretation benchmark score · → daily
- 🟦 2026-06-27 ·
fact · radiology image interpretation benchmark score vs prior best · → daily
- 🟥 2026-06-27 ·
narrative · sufficiency for reliable medical use · → daily
- 🟦 2026-06-20 ·
fact · 能够分析数据并扩展论文论证 · → daily
- 🟦 2026-06-17 ·
fact · CritPt基准得分 · → daily
- 🟦 2026-06-15 ·
fact · Epoch Capabilities Index 得分 · → daily
- 🟦 2026-06-12 ·
fact · unit distance 求解能力 · → daily
- 🟦 2026-06-12 ·
fact · unit distance 不是一次求解 · → daily
- 🟥 2026-06-07 ·
narrative · 能力范围评价 · → daily
- 🟦 2026-06-06 ·
fact · 月度推理成本 · → daily
- 🟥 2026-05-30 ·
narrative · benefit when used via Pi · → daily
- 🟦 2026-05-30 ·
fact · 定价 · → daily
- 🟥 2026-05-28 ·
narrative · 作为审稿人发现错误 · → daily
- 🟥 2026-05-28 ·
narrative · 发现幻觉结果并提供反馈 · → daily
- 🟥 2026-05-28 ·
narrative · 数学证明产出 · → daily
- 🟦 2026-05-24 ·
fact · 推理耗时变化 · → daily
- 🟦 2026-05-24 ·
fact · 记忆泄露行为 · → daily
← 实体目录 · 系统日志