以下部分引用 AI 总结,现阶段 AI 仍然有幻觉。请以内容的原文为准。
数据新鲜度群聊 07-14 ✓卖方 07-14 ✓Wrap 断档15天 ✕News 07-15 ✓X 断档12天 ✕页面生成 07-15
Gemini 3.5 Flash
253 atoms · 跨 23 天 · 首见 2026-05-15 · 最近 2026-07-01
三色: 🟦 fact 174 · 🟥 take 79 · stance ▲81/▼49/◆45
叙事状态: 🍂 fading (退潮) · 💤 沉寂 · 策展近14d 0 atom · 🐦 X 近14d 6
状态只算低频策展源 (群/卖方); X firehose 仅作背景音量
来源: 群 1 · X 252
时态: fresh:246 · stale:7
标签: 好数字:148 · 好观点:83 · 好信源:48 · 好思考:10 · 好问题:1
🧭 拥挤度 (一人一票): ▲ 12 位作者 (KOL12) + 群 1 条 vs ▼ 14 位作者 (KOL14)
⚖️ 空头 14/14 来自KOL
🔥 核心观点 · 3 个
≥2 个独立的人各自表达 take (已排除新闻 bot) · 多空按人算 · 跨语言合并
- 性能对比结果
2 人观点 · 🟢 多 2 · X · 🥇 @scaling01 ★★★★ (2026-05-19)
<details><summary>展开 2 人 · 2026-05-19→2026-05-22</summary>
- ▲
X @scaling01 ×2 ★★★★ [2026-05-19]
- ▲ X @OfficialLoganK ★ [2026-05-22]
- price increase relative to Gemini 3 Flash
1 人观点 · 🔴 空 1 · X · +2快讯 · 🥇 @scaling01 ★★★★ (2026-05-20)
<details><summary>展开 5 人 · 2026-05-19→2026-05-20</summary>
- ▼
X @scaling01 ★★★★ [2026-05-20]
- · X @AiBattle_ ★★★★ [2026-05-19]
- · X @Presidentlin ★★ [2026-05-19]
- · X @cheatyyyy 译 [2026-05-19]
- · X @Techmeme 📢 [2026-05-19]
- AI模型性能与GPT-5.5比较
1 人观点 · ⚪ 中性 · X · +1快讯 · 🥇 @scaling01 ★★★★ (2026-05-19)
<details><summary>展开 2 人 · 2026-05-19→2026-05-19</summary>
- ◆
X @scaling01 ★★★★ [2026-05-19]
- · X @testingcatalog 📢 [2026-05-19]
📄 广泛报道的事实 · 8 个
展开 (多人转述同一事实/数字 — 确认度高, **非独立观点**, 不标多空)
- 全球上线 — 5 源报道 (含2快讯) ·
X
- pricing — 3 源报道 (含1快讯) ·
X
- output token pricing — 3 源报道 (含1快讯) ·
X
- AA-Intelligence Index 得分 — 2 源报道 ·
X
- 成本 — 2 源报道 ·
X
- 速度评价 — 2 源报道 ·
X
- 计算机使用功能发布 — 2 源报道 ·
X
- input token pricing — 2 源报道 (含1快讯) ·
X
💢 核心分歧 (4 轴)
1. 成本暴涨 3x vs 性能提升是否值回票价
*topic: 价格动态/定价/定价能力 · 9 bull vs 10 bear*
🟢 bullish 侧:
- <a id="atom-25262a80960783b2"></a>2026-05-19
X·@Presidentlin [老旧] · is cheap for a pro model · 价格动态
tags: 好观点 · → daily
📷 原图
- <a id="atom-4fe5465776a5b329"></a>2026-05-19
X·@NoamShazeer · 输出令牌速度对比 · 价格动态/产品发布 +同日1条
The metrics on 3.5 Flash are major: 4x faster than other frontier models in output tokens per second
tags: 好数字 · → daily
value: qty=4x faster than other frontier models in output tokens per second
📷 原图
- <a id="atom-6f46fc78a01028e5"></a>2026-05-19
X·@demishassabis · 速度比其他前沿模型快 4 倍 · 价格动态 +同日2条
4x faster than other frontier models
tags: 好数字 · → daily
value: qty=4x · direction=faster
📷 原图
- <a id="atom-d6ffbca0fde07f2a"></a>2026-05-20
X·@hungjng69679118 · 输入定价比较 · 定价能力 +同日1条
百万tokens输入定价...约为 GPT-5.5 ($5.00)...价格优势明显
tags: 好数字 · → daily
value: qty=$1.5 · date=2026-05-20 · direction=down
- <a id="atom-b07532f2f7ddabd4"></a>2026-05-24
X·@theinformation · 定位 · 价格动态
positioned Gemini 3.5 Flash and its coding agent Antigravity as cheaper alternatives to Anthropic’s increasingly expensive coding models.
tags: 好观点 · → daily
🔴 bearish 侧:
- <a id="atom-ef57346443de00c3"></a>2026-05-19
X·@AiBattle_ · price increase relative to Gemini 3 Flash · 价格动态
Gemini 3.5 Flash is 3 times as expensive as Gemini 3 Flash
tags: 好数字 · → daily
value: qty=3x
- <a id="atom-600089affa958fab"></a>2026-05-19
X·@Presidentlin · pricing comparison to Gemini 3 Flash · 价格动态 +同日1条
a 3x increase over Gemini 3 Flash
tags: 好数字 · → daily
value: direction=increase · qty=3x
📷 原图
- <a id="atom-403b0fce479aa678"></a>2026-05-19
X·@burkov · 定价相对于替代品 · 价格动态
priced almost like 3.1 Pro, so I'm expecting that the new 3.5 Pro will be priced even higher
tags: 好思考 · → daily
- <a id="atom-7a5a190ca3e1719a"></a>2026-05-19
X·@theo · 成本 · 价格动态 +同日1条
Gemini 3.5 Flash costs 2x more to run than Gemini 3.1 Pro on similar tasks.
tags: 好数字·好观点 · → daily
value: qty=2x
📷 原图
- <a id="atom-e1cd48e9cc06d993"></a>2026-05-20
X·@flowersslop · cost vs Gemini 3.1 Pro · 价格动态/技术路线
Gemini 3.5 Flash costing more than 3.1 Pro while performing worse
tags: 好观点 · → daily
展开 5 条中性
- <a id="atom-8f5d7cc11c3e9e00"></a>2026-05-19
X·@scaling01 ⭐ · input pricing · 价格动态
pricing is $1.5 / per mtoks
tags: 好数字·好信源 · → daily
value: qty=$1.5
📷 原图
- <a id="atom-57cd32f254a86297"></a>2026-05-19
X·@AiBattle_ · AA-Intelligence 指数运行成本 · 模型成本/价格动态
Gemini 3.5 Flash has a significantly higher cost to run the AA-Intelligence index than Gemini 3.1 Pro Cost to run: - Gemini 3.5 Flash - 1552$
tags: 好数字 · → daily
value: qty=$1,552
📷 原图
- <a id="atom-2c613a286d094b9c"></a>2026-05-20
X·@scaling01 · 价格增长原因 · 价格动态
price increase is likely because of the high interactivity (tok/s/user)
tags: 好观点 · → daily
value: direction=up
📷 原图
2. 领先速度 vs 落后推理: 快是否等于好
*topic: 技术路线/模型性能/性能指标 · 33 bull vs 19 bear*
🟢 bullish 侧:
- <a id="atom-26b3d859acae801e"></a>2026-05-15
X·@cqkten · 放置方块效果出奇好 · 产品发布/技术路线
placing blocks works surprisingly well
tags: 好观点 · → daily
- <a id="atom-0a23e3fd92b98158"></a>2026-05-19
X·@wallstengine · 性能对比Gemini 3.1 Pro · 模型发布/技术路线 +同日2条
On Google’s benchmarks, it beats Gemini 3.1 Pro across coding, real-world agentic tasks, and scaled tool use.
tags: 好数字 · → daily
value: direction=superior · comparison=beats across coding, real-world agentic tasks, and scaled tool use
📷 原图
- <a id="atom-952bc7ba04e0c7e4"></a>2026-05-19
X·@anshelsag · 性能比较 · 技术路线
Outclasses 3.1 Pro in nearly every way
tags: 好观点·好思考 · → daily
value: direction=优于 3.1 Pro
📷 原图
- <a id="atom-3591796b41051f19"></a>2026-05-19
X·@scaling01 · 在TerminalBench 2.1 上性能优于 Gemini 3.1 Pro · 模型性能 +同日3条
Gemini 3.5 Flash beats Gemini 3.1 Pro across TerminalBench 2.1
tags: 好观点 · → daily
📷 原图
- <a id="atom-77393610001dedc5"></a>2026-05-19
X·@OfficialLoganK · 技术特征描述 · 技术路线
It pushes the frontier of intelligence, speed, and cost putting 3.5 Flash in a class of its own.
tags: 好观点 · → daily
📷 原图
🔴 bearish 侧:
- <a id="atom-93b479e5c8a86f19"></a>2026-05-15
X·@pigeon_s · 标记使用可能比 pro 多 10x token · 模型发布/技术路线_
it better not have insane token usage though like gemini 3 flash that thing burns like 10x more tokens than pro just to make up for being dumber
tags: 好观点·好问题 · → daily
- <a id="atom-3e5ab7296f526309"></a>2026-05-16
X·@mohamed_yo32851 · 蒸馏和削弱前版本更优 · 模型发布/技术路线
Gemini 3.5 flash ( before distillation and nerfing )
tags: 好观点 · → daily
- <a id="atom-6d19156be07d4767"></a>2026-05-19
X·@nsdjoe · arc-agi-2 benchmark score · 模型发布/技术路线
still a bit lower on arc-agi-2 as well
tags: 好数字 · → daily
value: direction=lower
- <a id="atom-7f5558cd67c2dec9"></a>2026-05-19
X·@Angaisb_ · cost to run · 模型发布/技术路线 +同日2条
More expensive to run that Gemini 3.1 Pro (almost 2x the cost to run the Artificial Analysis Intelligence Index)
tags: 好数字 · → daily
value: qty=~2x vs Gemini 3.1 Pro
📷 原图
- <a id="atom-daef80eb56ead5b0"></a>2026-05-19
X·@scaling01 · Coding Index 得分 · 技术路线 +同日1条
Gemini 3.5 Flash scores kinda low on the Coding Index due to terrible TerminalBench-Hard scores
tags: 好数字 · → daily
value: qty=low · date=2026-05-19
📷 原图
展开 18 条中性
- <a id="atom-9eb8ecec7b5c05b7"></a>2026-05-15
X·@cqkten · 损坏和破坏方块时仍不稳定 · 产品发布/技术路线
Pretty cool! Still pretty glitchy especially damage + breaking blocks
tags: 好观点 · → daily
- <a id="atom-caeff5d31ec4f39b"></a>2026-05-19
X·@scaling01 · AI模型性能与Opus 4.7比较 · 模型性能
Gemini 3.5 Flash comparable with Opus 4.7 on the Artificial Analysis Index
tags: 好观点 · → daily
📷 原图
- <a id="atom-053e9df6c6ad5c94"></a>2026-05-19
X·@emollick · 与前沿模型比较 · 技术路线
not as powerful as a full frontier model
tags: 好观点 · → daily
📷 原图
- <a id="atom-31d9421e127b7dfb"></a>2026-05-19
X·@arcprize · ARC-AGI-2 High score · 模型性能
Gemini 3.5 Flash ARC-AGI (Verified) ARC-AGI-2: - High: 72.1%, $0.85
tags: 好数字 · → daily
value: qty=72.1% · date=ARC-AGI-2
📷 原图
- <a id="atom-fb26955874214f57"></a>2026-05-19
X·@teortaxesTex · 运行 1000 次查询花费 30 分钟 · 模型性能
Gemini 3.5 Flash is very fast though: I ran 1000 queries in 30 minutes
tags: 好数字 · → daily
value: qty=1000 queries in 30 minutes
📷 原图
3. 定位模糊: 越级打 Pro 还是降级割韭菜
*topic: 定价权/竞争格局/模型评测/定价 · 9 bull vs 11 bear*
🟢 bullish 侧:
- <a id="atom-8fd9ce89617303bf"></a>2026-05-19
X·@arena · 价格-性能前沿位置 · 产品发布/定价权
Gemini 3.5 Flash also moves the price–performance frontier as the new top Arena score in its price tier.
tags: 好数字 · → daily
📷 原图
- <a id="atom-55d5f98739182269"></a>2026-05-19
X·@NoamShazeer · 性能对比前代模型 · 竞争格局/技术路线
Outperforms our previous 3.1 Pro model on nearly all benchmarks
tags: 好数字 · → daily
value: qty=Outperforms our previous 3.1 Pro model on nearly all benchmarks
📷 原图
- <a id="atom-f58534acf6e0debe"></a>2026-05-19
X·@scaling01 · 端到端延迟对比 · 竞争格局
GPT-5.5-medium has lower end-to-end latency, uses less tokens and is overall smarter and cheaper than Gemini 3.5 Flash
tags: 好观点 · → daily
value: direction=higher
📷 原图
- <a id="atom-e88ecafc22c12b0e"></a>2026-05-19
X·@OpenRouter · 性能对比 · 竞争格局/技术路线
Beats Gemini 3.1 Pro on coding, agentic work, and tool use at Flash-tier price and speed.
tags: 好观点 · → daily
📷 原图
- <a id="atom-42d7f95528c18ac1"></a>2026-05-20
X·@wallstengine · 成本 · 模型发布/定价权
Half the cost
tags: 好数字 · → daily
value: qty=Half the cost
🔴 bearish 侧:
- <a id="atom-9d8f62def3f868d7"></a>2026-05-19
X·@theo · 成本对比 GPT-5.5 Medium · 价格动态/竞争格局 +同日1条
Gemini 3.5 Flash is more expensive than GPT-5.5 Medium.
tags: 好数字 · → daily
📷 原图
- <a id="atom-91fc8f2c0a8d46fc"></a>2026-05-20
X·@scaling01 · 输出定价对比 · 价格动态/定价权
The $9 output pricing also hurts the models position, as it didn't improve on reasoning efficiency as GPT-5.5 did over GPT-5.4
tags: 好数字·好观点 · → daily
value: qty=$9 · direction=high
📷 原图
- <a id="atom-a7990a50aecbd3de"></a>2026-05-20
X·@bridgemindai · competitive positioning · 竞争格局/模型评测
Google's newest model can't even beat budget tier competition
tags: 好观点 · → daily
📷 原图
- <a id="atom-d0523b77583b6a75"></a>2026-05-20
X·@JustinWaugh · 单次 agentic run 成本 · 成本/定价权
It's verbose, which makes it very expensive ($22.96/agentic run!!)
tags: 好数字·好思考 · → daily
value: qty=$22.96
📷 原图
- <a id="atom-9adbe8ec10d1349b"></a>2026-05-22
X·@giffmana · cost per inference compared to 3.1 Pro · 定价权
It's cheaper per token but more expensive to actually do the job
tags: 好观点 · → daily
value: direction=up
展开 4 条中性
- <a id="atom-44a0b05f09a2da0b"></a>2026-05-19
X·@arcprize · ARC-AGI performance comparison · 模型性能/竞争格局
Gemini 3.5 Flash is on par with GPT-5.5 (Medium) on ARC-AGI
tags: 好观点 · → daily
📷 原图
- <a id="atom-6d27b9a142886e08"></a>2026-05-20
X·@JustinWaugh · Pencil Puzzle Bench 得分 · 模型评测/模型发布
Final score: 43%
tags: 好数字 · → daily
value: qty=43%
📷 原图
- <a id="atom-848bac40b97e5cba"></a>2026-05-21
X·@mweinbach · 与GPT 5.5在相同技能下接近持平 · 竞争格局
With the same skills, pretty close to equal in the same harness with codex runtime
tags: 好观点 · → daily
- <a id="atom-d0c8d98be01bbf61"></a>2026-06-22
X·@ayu_walk2525 · ranking relative to competitors · 竞争格局/模型质量
Gemini 3.5 Flash sits slightly below Kimi but just above DS. When you factor in its speed, it actually occupies a pretty solid niche.
tags: 好观点 · → daily
4. 复杂基准惊艳 vs 简单任务翻车: 智能不均衡
*topic: 模型能力/基准测试/性能评测 · 6 bull vs 1 bear*
🟢 bullish 侧:
- <a id="atom-6544542977357db4"></a>2026-05-20
X·@hungjng69679118 · 输出token速率比较 · 性能评测/定价能力
输出 token速率比同档前沿模型快约 4 倍
tags: 好数字 · → daily
value: qty=约2倍 · date=2026-05-20 · direction=up
- <a id="atom-f69350bb1eafdb32"></a>2026-05-22
X·@OfficialLoganK · progress on GDPval compared to 3.1 Pro · 模型能力
Gemini 3.5 Flash has made huge progress from 3.1 Pro on GDPval
tags: 好观点 · → daily
value: direction=up
📷 原图
- <a id="atom-942273f8e2c2e692"></a>2026-05-22
X·@AxcanNathan · output speed compared to Pro · 模型能力
3.5Flash seems at least ~36% faster than Pro at giving an answer (not counting TTFT which is not computable from AA data but ofc should favor Flash, since it seems to stay smaller)
tags: 好数字 · → daily
value: qty=36% · direction=up
📷 原图
- <a id="atom-1b3413947d681af9"></a>2026-05-27
X·@mweinbach · 知识工作表现 · 模型能力
I’ve saying Gemini 3.5 Flash is good at knowledge work
tags: 好观点 · → daily
- <a id="atom-e35b52437f8cf95a"></a>2026-06-30
X·@mweinbach · value for knowledge work · 模型能力 +同日1条
Gemini 3.5 Flash is the best value model for knowledge work hands down
tags: 好观点 · → daily
🔴 bearish 侧:
- <a id="atom-78838505657bf4b1"></a>2026-05-20
X·@emollick · messed up counting letters in words (May 2026) · 模型能力
The frontier is still jagged though (here is Gemini 3.5 Flash messing up counting letters in words)
tags: 好观点 · → daily
展开 2 条中性
- <a id="atom-de81d1557e0be6cf"></a>2026-06-24
X·@trycua · Cua-Bench 平均回报 · 基准测试
Gemini 3.5 Flash's native Computer Use 在 Cua-Bench 上 posted the highest mean reward of any frontier model we tested - 0.267
tags: 好数字 · → daily
value: qty=0.267 · direction=最高
🟦 客观事实 (facts) (30)
- <a id="atom-117d40efdf473032"></a>🟦 2026-05-19
X·@ValsAI ⭐ · Finance Agent benchmark 排名 · 模型发布/基准测试
Google's Gemini 3.5 Flash is the new #1 model on our Finance Agent benchmark (v2), dethroning GPT-5.5 by six points.
tags: 好数字·好观点·好信源 · → daily
value: qty=#1 · date=2026-05-19
📷 原图
- <a id="atom-e640fce6065175ab"></a>🟦 2026-05-26
X·@scaling01 ⭐ · CAIS Text Capabilities排名 · 模型发布/技术路线
Gemini 3.5 Flash ranking 4th on CAIS Text Capabilities
tags: 好数字·好信源 · → daily
value: qty=4th
[图: AI大语言模型在Text Capabilities Index上的各项基准测试评分及排名数据面板 — GPT-5.5 平均分: 54.1; Gemini 3.1 Pro 平均分: 52.9; GPT-5.4 平均分: 49.3; Gemini 3.5 Flash 平均分: 48.8; DeepSeek 4 Pro 平均分: 32.1]
📷 原图
- <a id="atom-824dad61e9827d2f"></a>🟦 2026-05-26
X·@scaling01 ⭐ · CAIS Vision排名 · 模型发布/技术路线
Gemini 3.5 Flash ranking 1st on Vision
tags: 好数字·好信源 · → daily
value: qty=1st
[图: AI大语言模型在Text Capabilities Index上的各项基准测试评分及排名数据面板 — GPT-5.5 平均分: 54.1; Gemini 3.1 Pro 平均分: 52.9; GPT-5.4 平均分: 49.3; Gemini 3.5 Flash 平均分: 48.8; DeepSeek 4 Pro 平均分: 32.1]
📷 原图
- <a id="atom-fd41b3b9e547d4d0"></a>🟦 2026-05-20
X·@scaling01 ⭐ · 相比GPT-5.4和GPT-5.3的表现 · 模型发布/技术路线
Gemini 3.5 Flash is currently between GPT-5.4 and GPT-5.3 and between Opus 4.6 and Opus 4.7 on the Artificial Analysis Index, implying a ~2-3 month lag
tags: 好数字·好信源 · → daily
value: direction=between
📷 原图
- <a id="atom-1bf7cb48f4718a1f"></a>🟦 2026-05-20
X·@scaling01 ⭐ · 在WeirdML上的表现落后 · 模型发布/技术路线
On WeirdML it's lagging behind GPT-5.2 and Opus 4.5
tags: 好数字·好信源 · → daily
value: direction=behind
📷 原图
- <a id="atom-2b5183397d69926f"></a>🟦 2026-05-20
X·@scaling01 ⭐ · 在SWE-Bench-Pro上的表现落后 · 模型发布/技术路线
On SWE-Bench-Pro it's behind GPT-5.2
tags: 好数字·好信源 · → daily
value: direction=behind
📷 原图
- <a id="atom-f6597aee97d2949d"></a>🟦 2026-05-20
X·@scaling01 ⭐ · 在TerminalBench Hard上的表现落后 · 模型发布/技术路线
on TerminalBench Hard it's worse than Opus 4.5
tags: 好数字·好信源 · → daily
value: direction=worse
📷 原图
- <a id="atom-4ecc0d524e97ce75"></a>🟦 2026-05-20
X·@JustinWaugh ⭐ · 测试成本 · 测试成本/推理成本
Cost me ~$1k just to run 30 agentic puzzles + 300 direct ask (~$700 came from the agentic pass)
tags: 好数字·好信源 · → daily
value: qty=~$1k · context=run 30 agentic puzzles + 300 direct ask
📷 原图
- <a id="atom-8f5d7cc11c3e9e00"></a>🟦 2026-05-19
X·@scaling01 ⭐ · input pricing · 价格动态
pricing is $1.5 / per mtoks
tags: 好数字·好信源 · → daily
value: qty=$1.5
📷 原图
- <a id="atom-81bf69f3bc74fb34"></a>🟦 2026-05-19
X·@scaling01 ⭐ · output pricing · 价格动态
pricing is $9 / per mtoks
tags: 好数字·好信源 · → daily
value: qty=$9
📷 原图
- <a id="atom-a7816e8b48a8293e"></a>🟦 2026-05-19
X·@scaling01 ⭐ · inference speed · 模型发布/技术路线
Google optimized Gemini 3.5 Flash to make it run up to 12x faster (~867 tokens/s) than comparable models in AntiGravity
tags: 好数字·好信源 · → daily
value: qty=~867 tokens/s · direction=up to 12x faster than comparable models in AntiGravity
- <a id="atom-23da72b6e9371aa9"></a>🟦 2026-05-19
X·@scaling01 ⭐ · pricing confirmed · 产品发布/定价
@scaling01: Gemini 3.5 Flash Pricing confirmed at $1.5 / $9 per mtoks
tags: 好数字·好信源 · → daily
value: qty=$1.5 / $9 per mtoks
📷 原图
- <a id="atom-c60d25566088c234"></a>🟦 2026-05-19
X·@JeffDean ⭐ · 基准测试得分 · 性能对比
outscores 3.1 Pro on agentic and coding benchmarks like Terminal-Bench and MCP Atlas, while running 4x faster than other frontier models.
tags: 好数字·好信源 · → daily
value: direction=高于
📷 原图
- <a id="atom-301c8cb161384c44"></a>🟦 2026-05-19
X·@scaling01 ⭐ · vals index 排名 · 技术路线/模型发布
Gemini 3.5 Flash ranking third on vals index
tags: 好信源·好数字 · → daily
value: qty=third
📷 原图
- <a id="atom-e3c1db4bcb9cfb9e"></a>🟦 2026-05-19
X·@zephyr_z9 ⭐ · pricing vs predecessor · 价格动态/盈利能力
they wouldn't increase the cost by 3x; 3.5 Flash is based on Gemini 3 Flash
tags: 好数字·好信源 · → daily
value: direction=3x more expensive than 3.5 Flash base · date=2026-05-20
- <a id="atom-06785646dde3cc91"></a>🟦 2026-05-19
X·@Techmeme ⭐ · price comparison to Gemini 3 Flash Preview · 定价
3x the price of Gemini 3 Flash Preview
tags: 好数字·好思考 · → daily
value: direction=higher · qty=3x
- <a id="atom-2bd39a214c2f6a73"></a>🟦 2026-05-19
X·@Techmeme ⭐ · price comparison to Gemini 3.1 Flash-Lite · 定价
6x the price of Gemini 3.1 Flash-Lite
tags: 好数字·好思考 · → daily
value: direction=higher · qty=6x
- <a id="atom-963d82a6e9747d0e"></a>🟦 2026-05-28
X·@ValsAI · FinanceAgent Benchmark 排名 · 模型发布/技术路线
it hit #1 on our FinanceAgent Benchmark taking 82 steps where competitors stopped at 13
tags: 好数字·好信源 · → daily
value: qty=#1
- <a id="atom-11bebe97f07e5813"></a>🟦 2026-05-28
X·@ValsAI · 工具调用步数 · 技术路线/数据参数
82 tool calls vs competitors' 13
tags: 好数字·好信源 · → daily
value: qty=82
- <a id="atom-72f815b74863fe6f"></a>🟦 2026-05-28
X·@ValsAI · 竞品工具调用步数 · 技术路线/数据参数
competitors stopped at 13
tags: 好数字·好信源 · → daily
value: qty=13
- <a id="atom-f5e4eeb108731eec"></a>🟦 2026-05-28
X·@ValsAI · 代码性能排名变化 · 技术路线/数据参数
Coding performance: from 20th to 10th place in one generation
tags: 好数字·好信源 · → daily
value: qty=从第20名到第10名
- <a id="atom-f69a68ba022a697b"></a>🟦 2026-05-19
X·@mweinbach · 推理速度是 12x 更快 · 模型发布/技术路线
You'll be able to get the 12x faster (~800-1200 tok/s) Gemini 3.5 Flash in Antigravity today
tags: 好数字·好信源 · → daily
value: qty=800-1200 tok/s · date=2026-05-19
- <a id="atom-d36ddab87af02f73"></a>🟦 2026-05-19
X·@AiBattle_ · AA-Intelligence Index 得分 · 模型发布/技术路线
Gemini 3.5 Flash scores 55 on the AA-Intelligence Index
tags: 好数字·好信源 · → daily
value: qty=55
📷 原图
- <a id="atom-fa0380c56ab8aeaa"></a>🟦 2026-05-19
X·@LechMazur · Debate Benchmark score · 技术评估/模型性能
Gemini 3.5 Flash scores 1479 on the Debate Benchmark. Ratings are Elo-like and centered near 1500.
tags: 好数字·好信源 · → daily
value: qty=1479 · date=2026-05-20
📷 原图
- <a id="atom-b0de748e580c6142"></a>🟦 2026-05-21
X·@AdamHoltererer · is the most expensive model to run on that benchmark · 基准测试成本
Hey Logan, it might be worth noting that Gemini 3.5 Flash is also the most expensive model to run on that benchmark, despite being a Flash model.
tags: 好数字·好观点 · → daily
📷 原图
- <a id="atom-87a5e529ce0ab8e1"></a>🟦 2026-06-22
X·@ArtificialAnlys · AA-Briefcase 单任务成本 · 价格动态
Gemini 3.5 Flash while costing over 98% less
tags: 好数字 · → daily
value: qty=$0.08 · currency=USD · date=2026-06-22
📷 原图
- <a id="atom-dbcb25eac848cfed"></a>🟦 2026-05-19
X·@scaling01 · 性能对比结果 · 模型发布/技术路线
Gemini 3.5 Flash beats Gemini 3.1 Pro across TerminalBench 2.1, GDPval and MCP Atlas
tags: 好数字 · → daily
value: direction=优于 · comparison_entity=Gemini 3.1 Pro
📷 原图
- <a id="atom-c6be274a8dda461d"></a>🟦 2026-05-19
X·@scaling01 · 推理速度倍数 · 模型性能
Gemini 3.5 Flash running up to 4x faster
tags: 好数字 · → daily
value: qty=4x · direction=fast
📷 原图
- <a id="atom-1ce723c75ac132b8"></a>🟦 2026-05-19
X·@arena · Text and Code Arena Frontend 排名 · 模型发布
Gemini 3.5 Flash has landed #9 for Text and Code Arena: Frontend.
tags: 好数字 · → daily
value: qty=#9
📷 原图
- <a id="atom-a43de8802e97dd33"></a>🟦 2026-05-19
X·@arena · Code Arena Frontend 得分 · 产品发布
Scoring 1507, this is a significant +70 point improvement over Gemini-3 Flash.
tags: 好数字 · → daily
value: qty=1507
📷 原图
🟥 多头 takes (bullish) (28)
- <a id="atom-b8908a5bde697bc5"></a>🟥 2026-07-01
X·@JakeABoggs · 在 HieroglyphBench 得分是 Anthropic Fable 5 和 GPT-5.5 的两倍以上 · 技术路线/模型发布
they're both still far behind the Gemini series, where 3.5 Flash has more than double the score
tags: 好数字·好观点 · → daily
📷 原图
- <a id="atom-952bc7ba04e0c7e4"></a>🟥 2026-05-19
X·@anshelsag · 性能比较 · 技术路线
Outclasses 3.1 Pro in nearly every way
tags: 好观点·好思考 · → daily
value: direction=优于 3.1 Pro
📷 原图
- <a id="atom-942273f8e2c2e692"></a>🟥 2026-05-22
X·@AxcanNathan · output speed compared to Pro · 模型能力
3.5Flash seems at least ~36% faster than Pro at giving an answer (not counting TTFT which is not computable from AA data but ofc should favor Flash, since it seems to stay smaller)
tags: 好数字 · → daily
value: qty=36% · direction=up
📷 原图
- <a id="atom-7b6705a6feff5493"></a>🟥 2026-05-20
X·@JasonBotterill · 视频输入能力评价 · 技术路线
just going to use it for video input it’s probably the worlds best at it
tags: 好观点 · → daily
- <a id="atom-3591796b41051f19"></a>🟥 2026-05-19
X·@scaling01 · 在TerminalBench 2.1 上性能优于 Gemini 3.1 Pro · 模型性能
Gemini 3.5 Flash beats Gemini 3.1 Pro across TerminalBench 2.1
tags: 好观点 · → daily
📷 原图
- <a id="atom-d4a6426ecbdb55dc"></a>🟥 2026-05-19
X·@scaling01 · 在GDPval上性能优于 Gemini 3.1 Pro · 模型性能
Gemini 3.5 Flash beats Gemini 3.1 Pro across GDPval
tags: 好观点 · → daily
📷 原图
- <a id="atom-6d238279772f7bde"></a>🟥 2026-05-19
X·@scaling01 · 在MCP Atlas上性能优于 Gemini 3.1 Pro · 模型性能
Gemini 3.5 Flash beats Gemini 3.1 Pro across MCP Atlas
tags: 好观点 · → daily
📷 原图
- <a id="atom-19e345c7ed347e03"></a>🟥 2026-05-19
X·@scaling01 · APEX-Agents-AA score · 模型发布
the APEX-Agents-AA score is excellent
tags: 好观点 · → daily
📷 原图
- <a id="atom-f58534acf6e0debe"></a>🟥 2026-05-19
X·@scaling01 · 端到端延迟对比 · 竞争格局
GPT-5.5-medium has lower end-to-end latency, uses less tokens and is overall smarter and cheaper than Gemini 3.5 Flash
tags: 好观点 · → daily
value: direction=higher
📷 原图
- <a id="atom-da043a4437b5396c"></a>🟥 2026-05-19
群 [老旧] · 速度与成本优势 · 速度/成本
Gemini 3.5 Flash 速度与成本优势显著
tags: 好观点 · → daily
- <a id="atom-ffc52325a611f350"></a>🟥 2026-06-30
X·@mweinbach · capability assessment · 产品性能/竞争格局
Gemini 3.5 flash and it can legitimately perform full work at a fraction of the price with no tweaks to prompting, but Claude can't
tags: 好思考 · → daily
value: direction=positive
- <a id="atom-e35b52437f8cf95a"></a>🟥 2026-06-30
X·@mweinbach · value for knowledge work · 模型能力
Gemini 3.5 Flash is the best value model for knowledge work hands down
tags: 好观点 · → daily
- <a id="atom-48cabdf0b2d93a88"></a>🟥 2026-06-30
X·@mweinbach · output quality vs Opus 4.8 and GPT 5.5 · 模型能力/定价优势
It’s outputs are close to Opus 4.8 or GPT 5.5, usually slightly less pretty but numbers are the same, but at 1/10th the price
tags: 好观点 · → daily
- <a id="atom-ed506d312d259449"></a>🟥 2026-05-28
X·@mweinbach · performance · 竞争格局
flash tops everyone else and is right behind
tags: 好观点 · → daily
- <a id="atom-1b3413947d681af9"></a>🟥 2026-05-27
X·@mweinbach · 知识工作表现 · 模型能力
I’ve saying Gemini 3.5 Flash is good at knowledge work
tags: 好观点 · → daily
- <a id="atom-f902db1b5fd9c610"></a>🟥 2026-05-23
X·@OfficialLoganK · 全面性评价 · 模型发布
3.5 Flash is a very well rounded model!
tags: 好观点 · → daily
- <a id="atom-f69350bb1eafdb32"></a>🟥 2026-05-22
X·@OfficialLoganK · progress on GDPval compared to 3.1 Pro · 模型能力
Gemini 3.5 Flash has made huge progress from 3.1 Pro on GDPval
tags: 好观点 · → daily
value: direction=up
📷 原图
- <a id="atom-94cd42ee99936415"></a>🟥 2026-05-22
X·@OfficialLoganK · competing at the frontier · 竞争格局
Flash is competing at the frontier
tags: 好观点 · → daily
📷 原图
- <a id="atom-f9aa82bbb4c64e68"></a>🟥 2026-05-22
X·@e_acc24 · speed compared to 3.1 Pro · 技术路线
It's flash progress because it's faster
tags: 好观点 · → daily
value: direction=up
- <a id="atom-de66cefdd7cba407"></a>🟥 2026-05-22
X·@mweinbach [老旧] · 性能比较 · 竞争格局
I like Gemini 3.5 Flash a lot, but Kimi K2.6 is better at most things for a fraction of the price
tags: 好观点 · → daily
value: direction=inferior
- <a id="atom-504bd4cd9dc876e3"></a>🟥 2026-05-22
X·@gabor · 响应速度 · 产品发布
Gemini 3.5 Flash is very fast.
tags: 好观点 · → daily
value: direction=very fast
- <a id="atom-cab0991e26ada690"></a>🟥 2026-05-20
X·@mweinbach · Antigravity专用训练 · 技术路线/产品发布
Try it specifically in antigravity...they weren't clear enough that it was trained for that harness, and it's VERY good in it
tags: 好观点 · → daily
- <a id="atom-25262a80960783b2"></a>🟥 2026-05-19
X·@Presidentlin [老旧] · is cheap for a pro model · 价格动态
tags: 好观点 · → daily
📷 原图
- <a id="atom-701e65e1d03e2264"></a>🟥 2026-05-19
X·@DavidSZD1 · expected performance tier · 性能比较
3.5 Flash will actually be closer to the performance of 3.1 Pro
tags: 好观点 · → daily
- <a id="atom-14312c652bd4ecf9"></a>🟥 2026-05-19
X·@mweinbach · 速度评价 · 模型发布
Gemini 3.5 Flash is VERY fast and very intelligent!
tags: 好观点 · → daily
📷 原图
- <a id="atom-4b6e4fb6e33f3beb"></a>🟥 2026-05-19
X·@anshelsag · 成本优势 · 盈利能力
especially in cost
tags: 好观点 · → daily
value: direction=更优
📷 原图
- <a id="atom-0cc722f36c077828"></a>🟥 2026-05-19
X·@OriolVinyalsML · performance on agentic tasks · 性能指标/产品发布
What excites me most is 3.5 Flash's breakthrough performance on complex, multi-step agentic tasks
tags: 好观点 · → daily
📷 原图
- <a id="atom-334e14edc572b0f9"></a>🟥 2026-05-19
X·@OriolVinyalsML · capability to generate interactive simulations · 模型发布/技术路线
it generates fully tactile, interactive HTML/SVG hardware simulations, complete with bump-mapped metals, spring physics, and procedural audio, in a single shot
tags: 好观点 · → daily
🟥 空头 takes (bearish) (30)
- <a id="atom-91fc8f2c0a8d46fc"></a>🟥 2026-05-20
X·@scaling01 · 输出定价对比 · 价格动态/定价权
The $9 output pricing also hurts the models position, as it didn't improve on reasoning efficiency as GPT-5.5 did over GPT-5.4
tags: 好数字·好观点 · → daily
value: qty=$9 · direction=high
📷 原图
- <a id="atom-d0523b77583b6a75"></a>🟥 2026-05-20
X·@JustinWaugh · 单次 agentic run 成本 · 成本/定价权
It's verbose, which makes it very expensive ($22.96/agentic run!!)
tags: 好数字·好思考 · → daily
value: qty=$22.96
📷 原图
- <a id="atom-76c14056db327dab"></a>🟥 2026-05-26
X·@theo [老旧] · 比 GPT-5.5 更贵且分数仅为其一半 · 定价/模型评测
Gemini 3.5 Flash being MORE EXPENSIVE than GPT-5.5 at HALF the score is also hilarious
tags: 好数字·好观点 · → daily
📷 原图
- <a id="atom-7a5a190ca3e1719a"></a>🟥 2026-05-19
X·@theo · 成本 · 价格动态
Gemini 3.5 Flash costs 2x more to run than Gemini 3.1 Pro on similar tasks.
tags: 好数字·好观点 · → daily
value: qty=2x
📷 原图
- <a id="atom-cce19fa30bfc1f41"></a>🟥 2026-05-20
X·@scaling01 · 价格比 Gemini 3 Flash 贵 3.5 倍 · 价格动态
Gemini 3.5 Flash ... is now 3.5x more expensive on WeirdML
tags: 好数字 · → daily
value: qty=3.5x · direction=more expensive
📷 原图
- <a id="atom-41f5f4d853c35f89"></a>🟥 2026-05-20
X·@scaling01 · 性价比对比 · 价格对比/模型推理成本
Gemini 3.5 Flash is 7.46 times more EXPENSIVE than GPT-5.5-xhigh on PencilPuzzleBench
tags: 好数字 · → daily
value: qty=7.46 · direction=more expensive · comparison_entity=GPT-5.5-xhigh
📷 原图
- <a id="atom-2438d71f7a48e7f4"></a>🟥 2026-05-22
X·@giffmana · actual cost compared to PR cost claim · 价格动态
both times it was ~double price of 3.1 Pro
tags: 好数字 · → daily
- <a id="atom-9d8f62def3f868d7"></a>🟥 2026-05-19
X·@theo · 成本对比 GPT-5.5 Medium · 价格动态/竞争格局
Gemini 3.5 Flash is more expensive than GPT-5.5 Medium.
tags: 好数字 · → daily
📷 原图
- <a id="atom-c05e5ddb23db85cb"></a>🟥 2026-05-19
X·@theo · 性能 · 价格动态/定价权
gemini-3.5-flash: 73m tokens, $1,522, 55 points
tags: 好数字 · → daily
value: qty=55 points
📷 原图
- <a id="atom-acd52b87e6886638"></a>🟥 2026-06-03
X·@scaling01 · 计算机使用性能 · 产品发布
I tried Gemini 3.5 Flash for computer use today and it was worse
tags: 好观点 · → daily
value: direction=worse
- <a id="atom-c11b7debe37a36f4"></a>🟥 2026-06-03
X·@scaling01 · 指令遵循能力 · 产品发布
it's chronically overthinking and doesn't follow instructions
tags: 好观点 · → daily
value: direction=does not follow
- <a id="atom-6b96c1ae1dbf77fb"></a>🟥 2026-05-20
X·@scaling01 · 用户感知模型大小 · 模型发布/技术路线
Gemini 3.5 Flash still feels like a smaller model as it misses the intent of my questions much more often than Opus 4.7 and GPT-5.5
tags: 好观点 · → daily
value: direction=smaller
📷 原图
- <a id="atom-b293feeb2f2ed8d1"></a>🟥 2026-05-20
X·@scaling01 [老旧] · 评测表现 · 模型发布/技术路线
at best it's Minimax 2.5~2.7 level and far from the frontier for real-world coding tasks
tags: 好观点 · → daily
value: qty=Minimax 2.5~2.7 level
- <a id="atom-b9f1a21edb36dac2"></a>🟥 2026-05-19
X·@scaling01 · CritPt score expected higher · 技术路线
expected higher on CritPt
tags: 好观点 · → daily
📷 原图
- <a id="atom-93b479e5c8a86f19"></a>🟥 2026-05-15
X·@pigeon_s · 标记使用可能比 pro 多 10x token · 模型发布/技术路线_
it better not have insane token usage though like gemini 3 flash that thing burns like 10x more tokens than pro just to make up for being dumber
tags: 好观点·好问题 · → daily
- <a id="atom-403b0fce479aa678"></a>🟥 2026-05-19
X·@burkov · 定价相对于替代品 · 价格动态
priced almost like 3.1 Pro, so I'm expecting that the new 3.5 Pro will be priced even higher
tags: 好思考 · → daily
- <a id="atom-8b4e65e3d925c6dd"></a>🟥 2026-05-19
X·@theo · 成本对比 GPT-5.5 Medium · 价格动态/竞争格局
3.5 Flash is "more expensive" and "dumber" than gpt-5.5 on medium
tags: 好思考 · → daily
📷 原图
- <a id="atom-b018abfd4bf48faa"></a>🟥 2026-06-22
X·@teortaxesTex · safety risk · 安全事故/模型质量
it's outright dangerous, I would not dare using it on my live filesystem
tags: 好观点 · → daily
- <a id="atom-637d66f0f6ed8cb5"></a>🟥 2026-06-21
X·@teortaxesTex · quality垂直 · 模型质量
Gemini is horrible
tags: 好观点 · → daily
- <a id="atom-6c1e2d011759ae32"></a>🟥 2026-06-05
X·@teortaxesTex · performance in harnesses · 技术路线/模型发布
Gemini 3.5 Flash is still dogshit in most harnesses
tags: 好观点 · → daily
- <a id="atom-cb411d4bd2e6a58a"></a>🟥 2026-05-22
X·@giffmana · cost messaging compared to 3.1 Pro · 价格动态
this ain't Flash progress. Or you need a new brand to mean cheap
tags: 好观点 · → daily
value: direction=up
📷 原图
- <a id="atom-9adbe8ec10d1349b"></a>🟥 2026-05-22
X·@giffmana · cost per inference compared to 3.1 Pro · 定价权
It's cheaper per token but more expensive to actually do the job
tags: 好观点 · → daily
value: direction=up
- <a id="atom-5ef93842162e8413"></a>🟥 2026-05-21
X·@basedjensen · 系统管理任务表现 · 技术路线/模型发布
it's absolutely ass in sysadmin tasks
tags: 好观点 · → daily
value: direction=差劲
- <a id="atom-9eaab4f5231015fc"></a>🟥 2026-05-20
X·@zephyr_z9 · 在单次推理中更聪明,但工具使用能力有限 · 技术路线
In our testing, Gemini 3.5 flash is smarter in one shot reasoning, but can’t do much with tools.
tags: 好观点 · → daily
- <a id="atom-1cfd2a0c880669cf"></a>🟥 2026-05-20
X·@zephyr_z9 · 显著更差 · 技术路线
lol way worse
tags: 好观点 · → daily
- <a id="atom-104a785f5dbec3ee"></a>🟥 2026-05-20
X·@MengTo · design complexity compared to Gemini 3.1 Pro · 产品发布
Worse than Gemini 3.1 Pro in term of design complexity like layout grids, responsiveness, scroll behaviors, webgl, animations, etc.
tags: 好观点 · → daily
value: direction=worse
- <a id="atom-a7990a50aecbd3de"></a>🟥 2026-05-20
X·@bridgemindai · competitive positioning · 竞争格局/模型评测
Google's newest model can't even beat budget tier competition
tags: 好观点 · → daily
📷 原图
- <a id="atom-78838505657bf4b1"></a>🟥 2026-05-20
X·@emollick · messed up counting letters in words (May 2026) · 模型能力
The frontier is still jagged though (here is Gemini 3.5 Flash messing up counting letters in words)
tags: 好观点 · → daily
- <a id="atom-a6b2dcb90411d8d5"></a>🟥 2026-05-19
X·@Presidentlin [老旧] · pricing direction desired · 价格动态
I want them to go lower not higher
tags: 好观点 · → daily
value: direction=lower
- <a id="atom-3e5ab7296f526309"></a>🟥 2026-05-16
X·@mohamed_yo32851 · 蒸馏和削弱前版本更优 · 模型发布/技术路线
Gemini 3.5 flash ( before distillation and nerfing )
tags: 好观点 · → daily
🟥 中性 takes (neutral) (15)
- <a id="atom-d8931cea85f99d50"></a>🟥 2026-05-20
X·@scaling01 · 输出 token 使用量多于 Gemini 3 Flash · 成本结构
it's more expensive because it uses more output tokens
tags: 好思考 · → daily
value: qty=more output tokens · direction=more
📷 原图
- <a id="atom-75c0000ccc9bdfdc"></a>🟥 2026-05-19
X·@scaling01 · reasoning efficiency · 技术路线
reasoning efficiency could also be better, but it kind of depends on what setting they used. if it's the max then it's very good
tags: 好思考 · → daily
value: direction=depends on setting
📷 原图
- <a id="atom-27ee8e3e32339fb3"></a>🟥 2026-05-22
X·@scaling01 · ALE-Bench 表现随迭代提升但不如 Kimi-K2.6 · 模型表现
Gemini 3.5 Flash only gets good with multiple iterations but gets mogged by Kimi-K2.6
tags: 好观点 · → daily
📷 原图
- <a id="atom-355423acc305bbd2"></a>🟥 2026-05-21
X·@scaling01 · 模型特性描述 · 模型发布
I feel like Gemini 3.5 Flash is like a scaled up version of GPT-OSS-120B
tags: 好观点 · → daily
- <a id="atom-2c613a286d094b9c"></a>🟥 2026-05-20
X·@scaling01 · 价格增长原因 · 价格动态
price increase is likely because of the high interactivity (tok/s/user)
tags: 好观点 · → daily
value: direction=up
📷 原图
- <a id="atom-caeff5d31ec4f39b"></a>🟥 2026-05-19
X·@scaling01 · AI模型性能与Opus 4.7比较 · 模型性能
Gemini 3.5 Flash comparable with Opus 4.7 on the Artificial Analysis Index
tags: 好观点 · → daily
📷 原图
- <a id="atom-693200bf04b60df9"></a>🟥 2026-05-19
X·@scaling01 · AI模型性能与GPT-5.5比较 · 模型性能
Gemini 3.5 Flash comparable with GPT-5.5 on the Artificial Analysis Index
tags: 好观点 · → daily
📷 原图
- <a id="atom-c74a0a9e4f7b8255"></a>🟥 2026-05-19
X·@scaling01 · AI模型性能与Gemini 3.1 Pro比较 · 模型性能
Gemini 3.5 Flash comparable with Gemini 3.1 Pro on the Artificial Analysis Index
tags: 好观点 · → daily
📷 原图
- <a id="atom-d0c8d98be01bbf61"></a>🟥 2026-06-22
X·@ayu_walk2525 · ranking relative to competitors · 竞争格局/模型质量
Gemini 3.5 Flash sits slightly below Kimi but just above DS. When you factor in its speed, it actually occupies a pretty solid niche.
tags: 好观点 · → daily
- <a id="atom-2a36894be0074df5"></a>🟥 2026-05-21
X·@mweinbach · 擅长知识工作而非编码 · 技术路线
Gemini 3.5 Flash seems to excel at knowledge work, not coding
tags: 好观点 · → daily
- <a id="atom-848bac40b97e5cba"></a>🟥 2026-05-21
X·@mweinbach · 与GPT 5.5在相同技能下接近持平 · 竞争格局
With the same skills, pretty close to equal in the same harness with codex runtime
tags: 好观点 · → daily
- <a id="atom-49d13144d655ef33"></a>🟥 2026-05-20
X·@MengTo · design quality compared to GPT 5.5 · 模型发布
As good as GPT 5.5 but way cheaper.
tags: 好观点 · → daily
value: direction=equal
- <a id="atom-65477334da8af4cc"></a>🟥 2026-05-20
X·@TheAhmadOsman · 市场营销定位 · 产品发布
It’s definitely marketed as flagship
tags: 好观点 · → daily
- <a id="atom-44a0b05f09a2da0b"></a>🟥 2026-05-19
X·@arcprize · ARC-AGI performance comparison · 模型性能/竞争格局
Gemini 3.5 Flash is on par with GPT-5.5 (Medium) on ARC-AGI
tags: 好观点 · → daily
📷 原图
- <a id="atom-9eb8ecec7b5c05b7"></a>🟥 2026-05-15
X·@cqkten · 损坏和破坏方块时仍不稳定 · 产品发布/技术路线
Pretty cool! Still pretty glitchy especially damage + breaking blocks
tags: 好观点 · → daily
🟦 新闻流 (squawk · 6)
展开新闻流 (FirstSquawk / financialjuice / wallstengine / DeItaone — 快讯, 非原创 take)
- <a id="atom-923b123e3a958c98"></a>🟦 2026-05-20
X·@wallstengine · agentic benchmarks 对比 GPT-5.5 和 Claude · 模型发布/技术路线
It is beating GPT-5.5 & Claude on agentic benchmarks
tags: 好数字 · → daily
- <a id="atom-c7338db410553aec"></a>🟦 2026-05-20
X·@wallstengine · 速度 · 模型发布/技术路线
4x the speed.
tags: 好数字 · → daily
value: qty=4x
- <a id="atom-42d7f95528c18ac1"></a>🟦 2026-05-20
X·@wallstengine · 成本 · 模型发布/定价权
Half the cost
tags: 好数字 · → daily
value: qty=Half the cost
- <a id="atom-0a23e3fd92b98158"></a>🟦 2026-05-19
X·@wallstengine · 性能对比Gemini 3.1 Pro · 模型发布/技术路线
On Google’s benchmarks, it beats Gemini 3.1 Pro across coding, real-world agentic tasks, and scaled tool use.
tags: 好数字 · → daily
value: direction=superior · comparison=beats across coding, real-world agentic tasks, and scaled tool use
📷 原图
- <a id="atom-6442ab7edd8541a3"></a>🟦 2026-05-19
X·@wallstengine · 推理速度输出tokens每秒 · 模型发布/技术路线
Artificial Analysis also shows it at 289 output tokens/sec
tags: 好数字 · → daily
value: qty=289 · uom=output tokens/sec
📷 原图
- <a id="atom-1d83e026e5ffab2d"></a>🟦 2026-05-19
X·@wallstengine · 推理速度对比Claude Opus 4.7和GPT-5.5 · 模型发布/技术路线
more than 4x faster than Claude Opus 4.7 and GPT-5.5
tags: 好数字 · → daily
value: direction=faster · comparison=more than 4x faster than Claude Opus 4.7 and GPT-5.5
📷 原图
⏱ 时间轴 (近 20)
- 🟦 2026-07-01 ·
fact · took 6 minutes for a task · → daily
- 🟥 2026-07-01 ·
fact · 在 HieroglyphBench 得分是 Anthropic Fable 5 和 GPT-5.5 的两倍以上 · → daily
- 🟥 2026-06-30 ·
position · value for knowledge work · → daily
- 🟥 2026-06-30 ·
position · output quality vs Opus 4.8 and GPT 5.5 · → daily
- 🟥 2026-06-30 ·
position · capability assessment · → daily
- 🟦 2026-06-30 ·
fact · cost per run · → daily
- 🟦 2026-06-24 ·
fact · 计算机使用功能发布 · → daily
- 🟦 2026-06-24 ·
fact · Computer Use 功能发布 · → daily
- 🟦 2026-06-24 ·
fact · Cua-Bench 平均回报 · → daily
- 🟦 2026-06-24 ·
fact · KiCad 任务完成度 · → daily
- 🟦 2026-06-24 ·
fact · 速度和成本 · → daily
- 🟦 2026-06-24 ·
fact · Computer Use benchmark score · → daily
- 🟦 2026-06-24 ·
fact · pretrain revision · → daily
- 🟦 2026-06-24 ·
fact · model ID available in AI Studio · → daily
- 🟥 2026-06-22 ·
narrative · ranking relative to competitors · → daily
- 🟥 2026-06-22 ·
narrative · safety risk · → daily
- 🟦 2026-06-22 ·
fact · AA-Briefcase 单任务成本 · → daily
- 🟥 2026-06-21 ·
narrative · quality垂直 · → daily
- 🟥 2026-06-21 ·
narrative · advantages · → daily
- 🟦 2026-06-19 ·
fact · DeepSWE score · → daily
← 实体目录 · 系统日志