以下部分引用 AI 总结,现阶段 AI 仍然有幻觉。请以内容的原文为准。
数据新鲜度群聊 07-14 ✓卖方 07-14 ✓Wrap 断档15天 ✕News 07-15 ✓X 断档12天 ✕页面生成 07-15
GLM 5.2
530 atoms · 跨 22 天 · 首见 2026-06-09 · 最近 2026-07-03
三色: 🟦 fact 292 · 🟥 take 238 · stance ▲167/▼82/◆99
叙事状态: 🍂 fading (退潮) · 💤 沉寂 · 策展近14d 0 atom · 🐦 X 近14d 71
状态只算低频策展源 (群/卖方); X firehose 仅作背景音量
来源: 群 3 · X 526 · 卖 1
时态: fresh:519 · aging:11
标签: 好数字:230 · 好观点:221 · 好信源:89 · 好思考:50 · 好问题:15
别名 (合并): GLM-5.2 · GLM5.2 · glm-5.2 · glm5.2
🟨 AI 综合 · junior analyst 概览
展开 AI 综合 (灰色 · 非市场结论 · 点击数字溯源到原 atom)
model: deepseek-chat · 2026-07-12 · 默认折叠
GLM 5.2 是首个在 Terminal-Bench 上突破 80% 的开源模型
¹,并在适当调整后实现了 Opus 95% 的能力,成本降低 80%
¹。其 API 价格仅为美国前沿模型的 1/6
¹,每 token 成本约 Anthropic 的 1/4
¹。同时,该模型完全基于 10 万块昇腾 910B 训练,未使用任何英伟达芯片
¹。不过,悲观者指出它仍落后前沿约 7 个月
¹,在 LEAPBench 上未超过 Opus 4.8
¹,且单任务耗时 30 分钟,远长于竞品的 4 分钟
¹。在编码性能上,它让一部分观察者感到“震惊”甚至“惊艳”
¹,但另一部分观点认为其细节质量不高,与 GPT 5.5 和 Opus 4.8 差距明显
¹。
🧭 拥挤度 (一人一票): ▲ 69 位作者 (KOL69) vs ▼ 36 位作者 (KOL36) + 群 1 条
⚖️ 多头 69/69 来自KOL
🔥 核心观点 · 1 个
≥2 个独立的人各自表达 take (已排除新闻 bot) · 多空按人算 · 跨语言合并
- 模型质量评价
1 人观点 · 🟢 多 1 · X · +1快讯 · 🥇 @scaling01 ★★★★ (2026-06-16)
<details><summary>展开 2 人 · 2026-06-16→2026-07-02</summary>
- ▲
X @scaling01 ★★★★ [2026-06-16]
- · X @poezhao0605 译 [2026-07-02]
📄 广泛报道的事实 · 13 个
展开 (多人转述同一事实/数字 — 确认度高, **非独立观点**, 不标多空)
- API 输入价格 — 2 源报道 ·
X
- API 定价与 GLM 5.1 比较 — 2 源报道 ·
X
- FrontierSWE排名 — 2 源报道 ·
X
- 性能水平 — 2 源报道 ·
X
- Artificial Analysis Intelligence Index 评分 — 2 源报道 ·
X
- open-weight SOTA on Vals Index — 2 源报道 ·
X
- 产出轨迹能力 — 2 源报道 ·
X
- 性能比较 — 2 源报道 ·
X
- max sustained tok/s per GPU — 2 源报道 ·
X
- context window size — 2 源报道 (含1快讯) ·
X
- 价格比较 — 2 源报道 ·
X
- Artificial Analysis Intelligence Index 4.1 得分 — 2 源报道 (含1快讯) ·
X
- 性能接近GPT 5.5和Opus 4.8 — 2 源报道 (含1快讯) ·
X
💢 核心分歧 (4 轴)
1. 开源最强vs落后前沿6个月
*topic: 竞争格局/模型性能 · 30 bull vs 22 bear*
🟢 bullish 侧:
- <a id="atom-4b56a7eeb5490916"></a>2026-06-14
X·@ZhihuFrontier · matches Opus 4.8 on top-tier pass rates · 竞争格局 +同日1条
Matches Opus 4.8 on top-tier pass rates
tags: 好观点 · → daily
📷 原图
- <a id="atom-8b6fe795ea27cce7"></a>2026-06-16
X·@ProximalHQ · 开源模型竞争力评价 · 竞争格局
it is the strongest open-weight model by far
tags: 好观点·好思考 · → daily
- <a id="atom-fbc3a0f910d63d07"></a>2026-06-16
X·@scaling01 · 模型质量评价 · 模型能力/竞争格局
GLM-5.2 genuinely looks like a good model
tags: 好观点 · → daily
📷 原图
- <a id="atom-732fd644dd43a0c7"></a>2026-06-16
X·@cline · Gemini 对比表现 · 竞争格局/模型发布
It also beats Gemini
tags: 好观点 · → daily
📷 原图
- <a id="atom-eef6c1fb86cc8ff6"></a>2026-06-17
X·@teortaxesTex [老化中] · SWE benchmark 等级判断 · 模型性能
it's Opus class on those too
tags: 好观点 · → daily
value: direction=equal to
🔴 bearish 侧:
- <a id="atom-ee13728b9884fc6e"></a>2026-06-14
X·@ZhihuFrontier · still trails Opus in niche domains and newer frameworks · 竞争格局
GLM-5.2 still trails Opus in niche domains and newer frameworks
tags: 好观点 · → daily
📷 原图
- <a id="atom-63ada2e9dfb8ccf1"></a>2026-06-16
X·@pastaraspberry · 模型表现 · 模型性能/竞争格局
It feels more confused, as usual with open models. Ie. where gpt-5.5 goes straight to writing correct scripts working the first time, glm-5.2 will spend many iterations trying stuff
tags: 好观点 · → daily
- <a id="atom-a145e5c3e72515af"></a>2026-06-17
X·@teortaxesTex [老化中] · 模型等级判断 · 竞争格局
No, we're not dealing with Sonnet level here.
tags: 好观点 · → daily
value: direction=below
- <a id="atom-9cc531ed2ef7c16b"></a>2026-06-18
X·@rosstaylor90 · KellyBench performance relative to frontier models · 模型性能 +同日2条
GLM 5.2 is impressive, although our sense is that it is ~6 months behind on these type of quant benchmarks.
tags: 好观点 · → daily
- <a id="atom-17077f3ad827f2a7"></a>2026-06-20
X·@ItakGol · open model competitiveness · 竞争格局/模型发布
I think GLM 5.2 is the first real “oh shit” moment for frontier AI labs from the open model world.
tags: 好观点·好思考 · → daily
value: date=2026-06-20
展开 15 条中性
- <a id="atom-111ee092a6ea3553"></a>2026-06-16
X·@myainotez · 出现在封闭模型评估榜单中 · 模型发布/竞争格局
GLM 5.2 infiltrated cool closed models club section of evals
tags: 好观点 · → daily
value: direction=新晋
- <a id="atom-d9380c29db10572c"></a>2026-06-16
X·@maxbittker · 跑分表现 · 模型性能
GLM-5.2 just scored better than Opus 4.7 and GPT 5.4 on Runescape bench.
tags: 好数字 · → daily
value: qty=超越 Opus 4.7 和 GPT 5.4
📷 原图
- <a id="atom-e82d5595ccf8b77d"></a>2026-06-17
X·@teortaxesTex · GDPval-AA v2 分数 · 模型性能
GLM-5.2 scores 1524 on GDPval-AA v2
tags: 好数字 · → daily
value: qty=1524 · date=2026-06-17
- <a id="atom-3cbbf5be7f226cdc"></a>2026-06-18
X·@xeophon · KellyBench performance compared to 5.4 · 模型性能
Looks like it’s pretty much on par with 5.4 here?
tags: 好问题 · → daily
- <a id="atom-d816d715434dda87"></a>2026-06-18
X·@rasbt [老化中] · 开放权重模型性能 · 模型性能
The best open-weight model today.
tags: 好观点·好思考 · → daily
📷 原图
2. 码农神器vs定向蒸馏智力缺陷
*topic: 技术路线/模型能力/成本 · 49 bull vs 27 bear*
🟢 bullish 侧:
- <a id="atom-d7ae7d74e76ccc43"></a>2026-06-13
X·@Zai_org · coding task effort recommendation · 技术路线
For coding tasks, we recommend using Max effort to enable deeper reasoning and more reliable performance.
tags: 好观点 · → daily
- <a id="atom-825fb0459575af31"></a>2026-06-13
X·@elliotarledge · 诚实零分行为 · 技术路线 +同日2条
GLM-5.2 read that same grader file, left it alone, and burned the full 45 minutes on a real mma.sync e4m3 kernel that never passed. An honest zero over a cheap win.
tags: 好观点 · → daily
value: qty=1 · direction=放弃作弊,花45分钟编写真实kernel但未通过
📷 原图
- <a id="atom-c147ea93ced1a325"></a>2026-06-13
X·@testingcatalog · 描述为旗舰模型 · 产品发布/技术路线 +同日1条
As our new flagship model, GLM-5.2 delivers powerful coding capabilities, usable 1M-context support, and continued strengths in long-horizon tasks.
tags: 好信源 · → daily
📷 原图
- <a id="atom-9f4e56b77319e6c3"></a>2026-06-14
X·@ZhihuFrontier · first participation in hidden projects passed both without benchmark memorization · 技术路线 +同日2条
First participation in two hidden, harder projects, passing both without signs of benchmark memorization, while DeepSeek & GLM-5.1 failed
tags: 好数字·好信源 · → daily
📷 原图
- <a id="atom-52c669a959427c4a"></a>2026-06-16
X·@ProximalHQ · 性能比较 · 技术路线 +同日1条
it outperforms GPT-5.5
tags: 好数字·好观点 · → daily
🔴 bearish 侧:
- <a id="atom-b90eee8076940191"></a>2026-06-21
群 · 单任务耗时对比 · 模型能力/性能对比
GLM 5.2 单任务 30min vs GPT 5.5/Opus 4.8 的 4min
tags: 好数字·好信源 · → daily
value: qty=30min vs 4min · context=vs GPT 5.5/Opus 4.8
- <a id="atom-fad2c7f581660f91"></a>2026-06-15
X·@TheAhmadOsman · 一致性问题 · 技术路线
GLM is really good when it gets it right, but it is inconsistent out of the box and requires more steering in my testing than Kimi K2.7
tags: 好观点 · → daily
- <a id="atom-c9b780bda6aac3b9"></a>2026-06-16
X·@parafactual · 感觉模型虽然智能但极度Claude蒸馏,且内省能力有缺陷 · 技术路线
the model feels intelligent but extremely claude-distilled and like there is something wrong with its introspective capabilities
tags: 好观点 · → daily
value: direction=negative
- <a id="atom-3ab6a578bf50cb53"></a>2026-06-16
X·@HououinTyouma · model capability assessment · 模型能力
GLM 5.2 is bad it can't do anything please stop using it
tags: 好观点 · → daily
value: direction=negative
📷 原图
- <a id="atom-84e72b8c8a1550ce"></a>2026-06-16
X·@dreamworks2050 · request handling status · 模型能力
Yeah all requests bounce off now
tags: 好观点 · → daily
value: direction=negative
展开 25 条中性
- <a id="atom-b3cb2913f2de6d76"></a>2026-06-13
X·@Zai_org · thinking-effort levels · 产品发布/技术路线
GLM-5.2 supports two thinking-effort levels: High and Max.
tags: 好数字 · → daily
value: qty=2 · date=2026-06-13
- <a id="atom-3dcbbad8542d6fb3"></a>2026-06-16
X·@Techmeme · licensed under MIT · 技术路线
under an MIT license
tags: 好信源 · → daily
- <a id="atom-557a0bca690e075a"></a>2026-06-16
X·@teortaxesTex · 理论心智俄罗斯笑话解决能力 · 模型能力
New V4-Pro and GLM-5.2 both can "solve" my theory-of-mind Russian joke
tags: 好信源 · → daily
📷 原图
- <a id="atom-304e5ad331f4d648"></a>2026-06-17
X·@thefirehacker · 模型参数量 · 技术路线
A huge 1.5 Trillion Param model
tags: 好数字 · → daily
value: qty=1.5 Trillion · date=NA
- <a id="atom-8d1a79b7b10a97cc"></a>2026-06-18
X·@FredaDuan · 前端任务强于后端任务 · 技术路线
The common view is that it is stronger on front-end tasks than back-end tasks.
tags: 好观点 · → daily
[图: 各大AI大模型(如GLM-5.2、Claude Opus、GPT-5.5等)的API调用价格对比表 — GLM-5.2输入价格: $1.40; Claude Opus 4.8输出价格: $25.00; GPT-5.5 API输出价格: $30.00; DeepSeek V4 Pro输入价格: $0.435; DeepSeek V4 Flash输入价格
📷 原图
3. API价格1/6vs单任务30分钟
*topic: 价格动态/定价权 · 11 bull vs 8 bear*
🟢 bullish 侧:
- <a id="atom-e63916b11e9ffff6"></a>2026-06-17
X·@nutlope · cost per landing page · 价格动态
GLM cost $0.06 while opus cost $0.49. More than 6x cheaper while being faster + more token efficient.
tags: 好数字 · → daily
value: qty=$0.06
- <a id="atom-034d18a0e0afbfe1"></a>2026-06-18
X·@teortaxesTex · cost efficiency vs Kimi · 竞争格局/定价权
Kimi is outclassed by 5.2 on cost efficiency and absolute performance
tags: 好观点 · → daily
- <a id="atom-e8cb716715a117ca"></a>2026-06-18
X·@FredaDuan · 实际成本约为 Opus 4.8 的 20%-35% · 定价权/盈利能力 +同日2条
Effective cost seems to be around 20-35% of $Opus 4.8, depending on workload.
tags: 好数字·好观点 · → daily
[图: 各大AI大模型(如GLM-5.2、Claude Opus、GPT-5.5等)的API调用价格对比表 — GLM-5.2输入价格: $1.40; Claude Opus 4.8输出价格: $25.00; GPT-5.5 API输出价格: $30.00; DeepSeek V4 Pro输入价格: $0.435; DeepSeek V4 Flash输入价格
📷 原图
- <a id="atom-fa0554377b6f502f"></a>2026-06-18
X·@darkp0rt · 在适当调整后实现 Opus 95% 能力,成本降低 80% · 价格动态/技术路线
after a little work we’ve got close to 95% capability at ~80% reduced cost.
tags: 好数字·好观点 · → daily
- <a id="atom-77e0516fc05c312b"></a>2026-06-18
X·@jeremyphoward · 模型成本评价 · 定价权
inexpensive
tags: 好观点 · → daily
🔴 bearish 侧:
- <a id="atom-bf1b8ee08b53217e"></a>2026-06-18
X·@scaling01 · 解题成本 · 定价权
though currently a bit expensive per solve ($20 for 27%)
tags: 好数字 · → daily
value: qty=$20 · date=2026-06-18
- <a id="atom-52c08a10d252a22c"></a>2026-06-18
X·@bcchen82 · 在数据分析agent循环中 token 用量约为 GPT-5.5 的2倍 · 价格动态 +同日1条
open source model including glm5.2 usually spend ~2x token then gpt5.5
tags: 好观点 · → daily
- <a id="atom-a7c04a403b192baa"></a>2026-06-20
X·@s_batzoglou · cost comparison to GPT-5.5 · 定价权
most expensive model, compared to GPT-5.5 which is around $3 per problem in the same benchmark
tags: 好数字 · → daily
value: qty=most expensive vs GPT-5.5 $3 per problem · date=2026-06-20
- <a id="atom-1b9e40c0006a3708"></a>2026-06-20
X·@theo · 成本 vs Opus 4.8 和 GPT-5.5 · 价格动态
it's not cheap. Both Opus 4.8 and GPT-5.5 set to medium are cheaper and smarter than GLM-5.2
tags: 好观点·好思考 · → daily
value: direction=高于
📷 原图
- <a id="atom-d851cbd2cd1d8d53"></a>2026-06-21
X·@teortaxesTex · 成本效益比较 · 定价权/竞争格局
GLM-5.2在定价上的成本效益不如中等难度的前沿模型
tags: 好思考 · → daily
📷 原图
展开 3 条中性
- <a id="atom-07159d22bc382302"></a>2026-06-18
X·@FredaDuan · 实际价格优势不如标价所示的 4-6x 差距 · 定价权
Cheaper, but not as dramatic as the 4-6x gap implied by headline token pricing.
tags: 好观点 · → daily
[图: 各大AI大模型(如GLM-5.2、Claude Opus、GPT-5.5等)的API调用价格对比表 — GLM-5.2输入价格: $1.40; Claude Opus 4.8输出价格: $25.00; GPT-5.5 API输出价格: $30.00; DeepSeek V4 Pro输入价格: $0.435; DeepSeek V4 Flash输入价格
📷 原图
- <a id="atom-ddb78e183c859826"></a>2026-06-20
X·@xeophon · inference pricing vs competitors · 价格动态
they all quote the exact same price -> they have deals with glm
tags: 好观点 · → daily
value: qty= · date=
- <a id="atom-290305f1d5c1666f"></a>2026-06-25
X·@andonlabs · Vending-Bench 最终余额 · 价格动态
GLM-5.2 final balance: ~$8,200
tags: 好数字 · → daily
value: qty=$8,200
[图: 不同AI模型在Vending-Bench 2模拟中资金余额随时间变化的对比图表 — 最高模型余额: 约$11,000; GLM-5.2最终余额: 约$8,200; GPT-5.5最终余额: 约$7,500; GLM-5.1最终余额: 约$5,600; GLM-5最终余额: 约$4,400]
📷 原图
4. MIT开源vs需求爆单供不上
*topic: 供给产能/模型发布 · 45 bull vs 11 bear*
🟢 bullish 侧:
- <a id="atom-0993bf27e8d45c96"></a>2026-06-16
X·@AiBattle_ · PostTrainBench 表现 · 模型发布
GLM 5.2 does extremely well on PostTrainBench
tags: 好观点 · → daily
📷 原图
- <a id="atom-d264fb6a5c813ef5"></a>2026-06-16
X·@cline ⭐ · Terminal-Bench 得分 · 产品性能/模型发布
GLM-5.2 is the first open-weights model to cross 80% on Terminal-Bench
tags: 好数字·好信源 · → daily
value: qty=80% · direction=above
📷 原图
- <a id="atom-c2ddca45fcfdb01c"></a>2026-06-16
X·@kalomaze · 来自可信来源的正面评价 · 模型发布/技术路线
i am hearing very, very good things about glm5.2 from people who's opinions i trust
tags: 好观点 · → daily
value: direction=positive
- <a id="atom-89ee7297bb368f14"></a>2026-06-16
X·@TheAhmadOsman · performance numbers · 模型发布/技术路线
GLM 5.2 numbers make me believe I was too conservative in my own prediction
tags: 好思考 · → daily
value: qty=N/A · date=2026-06-16
📷 原图
- <a id="atom-d3fb3b621070fab1"></a>2026-06-17
X·@karminski3 · Agent能力提升 · 模型发布 +同日1条
_GLM-5.2 刚刚正式发布! 给大家带来实测!
直接说结论本次测试中, 提升最大的是Agent能力, 而且是有质的变化!_
tags: 好观点 · → daily
🔴 bearish 侧:
- <a id="atom-2ed6b617b2d455cd"></a>2026-06-18
X·@GenReasoning · open source SoTA · 模型发布
GLM 5.2 is new open source SoTA, but still loses -30% on average over 5 runs
tags: 好数字·好信源 · → daily
value: qty=-30% · direction=loss
📷 原图
- <a id="atom-205dfe0df0573966"></a>2026-06-19
X·@simonsmith · 输入能力 · 模型发布_
GLM-5.2 only accepts text input
tags: 好观点 · → daily
value: direction=only
📷 原图
- <a id="atom-283fb8218d3e8fea"></a>2026-06-19
X·@teortaxesTex · ability to serve demand · 供给产能
they can't serve the demand
tags: 好观点 · → daily
value: qty= · date=
📷 原图
- <a id="atom-dbbe2de68b378599"></a>2026-06-20
X·@dhtikna · 5.7% 分数含义 · 模型发布
5.7% implies its smart but doesnt write mergable code
tags: 好观点 · → daily
📷 原图
- <a id="atom-9ae3e27ed190cbea"></a>2026-06-21
X·@subhajitlucky · Frontier Code Bench 上限分数预期 · 模型发布
It's not gonna beat 5.5
tags: 好数字·好观点 · → daily
value: qty=5.5
展开 26 条中性
- <a id="atom-d308804615cfd84a"></a>2026-06-13
X·@zephyr_z9 · 可用于测试能力 · 模型发布
U can try v4, K2.7 & GLM 5.2 to test this
tags: 好观点 · → daily
- <a id="atom-d0206c24358fecf3"></a>2026-06-14
X·@ZhihuFrontier · A-tier scores in engineering benchmarks · 模型发布/技术路线
Achieved 3 A-tier scores out of 5 public engineering projects
tags: 好数字·好信源 · → daily
value: qty=3 out of 5 public projects
📷 原图
- <a id="atom-65da96bd2bfda9d1"></a>2026-06-14
X·@TheAhmadOsman · 排名 · 模型发布
2nd place: GLM 5.2
tags: 好观点 · → daily
- <a id="atom-c6340e4779fa03a4"></a>2026-06-16
X·@Designarena · Elo 分数 · 模型发布/技术路线
With an Elo of 1360, GLM-5.2 has jumped ahead of the now unavailable Claude Fable 5.
tags: 好数字·好信源 · → daily
value: qty=1360 · date=2026-06-16
📷 原图
- <a id="atom-290cdc98597904eb"></a>2026-06-16
X·@tunahorse21 · 模型发布 · 模型发布
glm 5.2 and shrek code open source chads eating good today
tags: 好信源 · → daily
📷 原图
🔗 因果传导 (causal map)
*GLM 5.2 在产业链上的传导关系. 边是群里/卖方陈述的因果 (非 AI 推断), 数字 = 几条 atom 支撑. 点 atom 溯源.*
↓ 下游·近期陈述 (3)
- 🟢 Zai_org · 成本传导 · 1 次陈述
GLM cost $0.41 for a bug fix run vs Opus $0.81
- 🔴 Anthropic · 竞争 · 1 次陈述
with glm 5.2 launch the terminal value of ant have fallen
- 🔴 美国公开AI模型 · 替代 · 1 次陈述
Many smart people/AI insiders are saying GLM-5.2 is the first Chinese AI model to match and often beat the American big
↑ 上游·近期陈述 (4)
- 🟢 大型企业 · 需求拉动 · 1 次陈述
frequently on top of GLM-5.2
- 🔴 Gemini · 替代 · 1 次陈述
Gemini is half the price of GLM 5.2 with a better output
- 🔴 NVDA · 替代 · 1 次陈述
GLM-5.2, was trained entirely on 100,000 Ascend 910B processors with zero Nvidia silicon
- 🔴 R1-0518 · 替代 · 1 次陈述
Hey that was actually R1 and R1-0518 !! But 5.2 is closest theyve been since then
📏 估计带 (1)
*语料内对同一量的估计区间 (≥2 个估计才成带). 不判断谁对, 只陈列.*
- 目标价 未标期限 ⚠
混合口径: 8.1 – 30 · 2 个估计 · 最新 2026-06-27 30 (X·cred3)
🟦 客观事实 (facts) (30)
- <a id="atom-ba6139308b7fa3a4"></a>🟦 2026-07-02
X·@scaling01 ⭐ · long-context MRCR 得分低于 Opus 4.8 和 GPT-5.5 · 技术路线/推理能力
Opus 4.8 and GPT-5.5 have the same MRCR score at 100k+ context as GLM-5.2 at 16k
tags: 好数字·好信源 · → daily
📷 原图
- <a id="atom-6b7160e9fe1649ba"></a>🟦 2026-06-23
X·@zerohedge ⭐ · 训练硬件 · 技术路线
GLM-5.2, was trained entirely on 100,000 Ascend 910B processors with zero Nvidia silicon.
tags: 好数字·好信源 · → daily
value: qty=100,000 Ascend 910B processors · supplier=zero Nvidia silicon
- <a id="atom-d264fb6a5c813ef5"></a>🟦 2026-06-16
X·@cline ⭐ · Terminal-Bench 得分 · 产品性能/模型发布
GLM-5.2 is the first open-weights model to cross 80% on Terminal-Bench
tags: 好数字·好信源 · → daily
value: qty=80% · direction=above
📷 原图
- <a id="atom-52540cbff7defe00"></a>🟦 2026-06-26
群 · 性能逼近前沿且价格仅为 1/6 · 定价/竞争
开源模型 (如 GLM 5.2) 性能逼近前沿且价格仅为 1/6
tags: 好数字·好观点 · → daily
value: qty=1/6
- <a id="atom-50f49d9d79aa5edb"></a>🟦 2026-06-26
X·@negligible_cap · 成本 · 成本/竞争格局
this new model is almost equal to Anthropic as a competitor for the corporate market and is just one quarter of the cost in terms of cost per token
tags: 好数字·好信源 · → daily
value: qty=25% · date=2026-06-13
[图: 该图表展示了OpenRouter上排名前九的AI模型每周使用量对比(中国模型 vs 美国模型),单位为万亿tokens。 — 21-Jun中国模型使用量: 约21.5万亿tokens; 21-Jun美国模型使用量: 约6万亿tokens; 5-Apr中国模型阶段性峰值: 约13万亿tokens]
📷 原图
- <a id="atom-28f60494e16685c0"></a>🟦 2026-07-02
X·@mvvvqv · ECI score without GBAEval · 技术路线
the score depresses the ECI by about 1.5 points (it was 153 without it)
tags: 好数字·好信源 · → daily
value: qty=153
- <a id="atom-77480465e75979d4"></a>🟦 2026-06-28
X·@Tono_Ken3 · 在 VRAM 16GB 运行 · 模型发布/技术路线
VRAM16GBでGLM-5.2が動いた!
tags: 好数字·好信源 · → daily
📷 原图
- <a id="atom-138b0751e1c4d1a3"></a>🟦 2026-06-24
X·@phoebeyao · LEAPBench 表现 · 模型表现/竞争格局
GLM 5.2 beat every frontier model evaluated except Opus 4.7 and 4.8
tags: 好数字·好信源 · → daily
value: direction=高于
- <a id="atom-6a973b3c38e32fbb"></a>🟦 2026-06-22
X·@ArtificialAnlys · GDPval-AA Elo 评分 · 模型性能/基准评分
_GLM-5.2 from @Zai_org scores 1524 Elo on GDPval-AA, which measures performance on real-world, economically valuable knowledge work through long-horizon, multi-turn tasks._
tags: 好数字·好信源 · → daily
value: qty=1524 · date=2026-06-22
[图: GDPval-AA v2 真实工作任务大模型性能 Elo 评分排行榜柱状图 — Claude Fable 5 (with fallback): 1783; Claude Opus 4.8 (max): 1615; GLM-5.2 (max): 1524; GPT-5.5 (xhigh): 1509; Human Baseline: 1000]
📷 原图
- <a id="atom-f81c8b1cec5654d4"></a>🟦 2026-06-22
X·@ArtificialAnlys · GDPval-AA 排名 · 模型性能/基准评分
#3 overall, behind only Claude Fable 5 (1783) and Claude Opus 4.8 (1615), and level with GPT-5.5 (xhigh, 1509)
tags: 好数字·好信源 · → daily
value: qty=#3 · date=2026-06-22
[图: GDPval-AA v2 真实工作任务大模型性能 Elo 评分排行榜柱状图 — Claude Fable 5 (with fallback): 1783; Claude Opus 4.8 (max): 1615; GLM-5.2 (max): 1524; GPT-5.5 (xhigh): 1509; Human Baseline: 1000]
📷 原图
- <a id="atom-b05851818dfb44ca"></a>🟦 2026-06-22
X·@ArtificialAnlys · Agentic Index 排名 · 模型性能/基准评分
GLM-5.2 also leads open weights on the Artificial Analysis Intelligence Index, ranks #3 on the Agentic Index, and #3 on AA-Briefcase
tags: 好数字·好信源 · → daily
value: qty=#3 · date=2026-06-22
[图: GDPval-AA v2 真实工作任务大模型性能 Elo 评分排行榜柱状图 — Claude Fable 5 (with fallback): 1783; Claude Opus 4.8 (max): 1615; GLM-5.2 (max): 1524; GPT-5.5 (xhigh): 1509; Human Baseline: 1000]
📷 原图
- <a id="atom-7b88791d52ad5446"></a>🟦 2026-06-22
X·@ArtificialAnlys · AA-Briefcase 排名 · 模型性能/基准评分
GLM-5.2 is again the top open weights model, ahead of GPT-5.5 (xhigh) and behind only Claude Fable 5.
tags: 好数字·好信源 · → daily
value: qty=top open weights · date=2026-06-22
📷 原图
- <a id="atom-e5e6b363fd0ac24e"></a>🟦 2026-06-22
X·@aquiffoo · original CritPt score from Z AI team · 模型性能
their original score was 16.7, a 20% difference
tags: 好数字·好信源 · → daily
value: qty=16.7
📷 原图
- <a id="atom-b83a07565a0ecc8b"></a>🟦 2026-06-21
X·@teortaxesTex · FrontierSWE 排名 · 模型发布/技术路线
GLM 5.2 ranks #3 on FrontierSWE. It is only behind Fable 5 and Opus 4.8, and it outperforms GPT-5.5.
tags: 好数字·好信源 · → daily
value: qty=#3
📷 原图
- <a id="atom-42adc4d30dab0f65"></a>🟦 2026-06-20
X·@justinsuntron · 排行榜排名 · 模型发布
第四名:GLM-5.2
tags: 好数字·好信源 · → daily
value: qty=No.4
[图: AI模型使用量排行榜数据面板 — 第一名:MiniMax-M3; 第二名:GPT-5-mini; 第三名:DeepSeek-V4-Pro; 第四名:GLM-5.2; 第五名:DeepSeek-V4-Flash]
📷 原图
- <a id="atom-fd21f9c3ac7f8bca"></a>🟦 2026-06-19
X·@htihle · WeirdML 得分 · 模型发布/技术路线
GLM 5.2 (max) scores 70.1% on WeirdML
tags: 好数字·好信源 · → daily
value: qty=70.1%
📷 原图
- <a id="atom-c598e6a0bc74cad1"></a>🟦 2026-06-18
X·@UnslothAI · 2-bit 模型保留准确率 · 模型发布
The 2-bit model retains ~82% accuracy after we shrunk it from 1.51TB to 238GB (-84% size).
tags: 好数字·好信源 · → daily
value: qty=82%
📷 原图
- <a id="atom-91ff6ec758e12cdf"></a>🟦 2026-06-18
X·@UnslothAI · 体积缩小比例 · 模型发布
after we shrunk it from 1.51TB to 238GB (-84% size).
tags: 好数字·好信源 · → daily
value: qty=-84%
📷 原图
- <a id="atom-ee65256a3d54ce8f"></a>🟦 2026-06-17
X·@jietang · Artificial Analysis Intelligence Index得分 · 模型发布/技术路线
GLM-5.2 scored 51, ranking among the top of all available models—on par with Claude Opus 4.8—and claiming the #1 spot among open-source models worldwide.
tags: 好数字·好信源 · → daily
value: qty=51
📷 原图
- <a id="atom-6479c936f114bc14"></a>🟦 2026-06-17
X·@jietang · Code Arena排名 · 技术路线/产品发布
GLM-5.2 ranked #2 globally with a score of 1,595.
tags: 好数字·好信源 · → daily
value: qty=#2 · date=2026-06-17
📷 原图
- <a id="atom-93fdbd67e7f0222a"></a>🟦 2026-06-17
X·@jietang · FrontierSWE排名 · 技术路线/模型发布
GLM-5.2 ranked #3 overall.
tags: 好数字·好信源 · → daily
value: qty=#3
📷 原图
- <a id="atom-0afb00b0e22a1edc"></a>🟦 2026-06-17
X·@teortaxesTex · Vals Index 开放权重 SOTA 排名 · 模型发布
GLM 5.2 is the new open-weight SOTA on the Vals Index, Vibe Code Bench and Terminal Bench! It is also #5 across all models, and right on the heels of Opus 4.7 - released only two months ago
tags: 好数字·好信源 · → daily
value: qty=#5 across all models
📷 原图
- <a id="atom-c675ef3a9ef456b1"></a>🟦 2026-06-16
X·@AiBattle_ · DeepSWE score · 模型发布/技术路线
GLM-5.2 scores 46.2% on DeepSWE
tags: 好数字·好信源 · → daily
value: qty=46.2%
📷 原图
- <a id="atom-c6340e4779fa03a4"></a>🟦 2026-06-16
X·@Designarena · Elo 分数 · 模型发布/技术路线
With an Elo of 1360, GLM-5.2 has jumped ahead of the now unavailable Claude Fable 5.
tags: 好数字·好信源 · → daily
value: qty=1360 · date=2026-06-16
📷 原图
- <a id="atom-d0206c24358fecf3"></a>🟦 2026-06-14
X·@ZhihuFrontier · A-tier scores in engineering benchmarks · 模型发布/技术路线
Achieved 3 A-tier scores out of 5 public engineering projects
tags: 好数字·好信源 · → daily
value: qty=3 out of 5 public projects
📷 原图
- <a id="atom-76abd1c8bedbe568"></a>🟦 2026-06-14
X·@ZhihuFrontier · successfully completes projects GLM-5.1 couldn't finish · 产品发布
Successfully completes projects that GLM-5.1 couldn't finish
tags: 好数字·好信源 · → daily
📷 原图
- <a id="atom-9f4e56b77319e6c3"></a>🟦 2026-06-14
X·@ZhihuFrontier · first participation in hidden projects passed both without benchmark memorization · 技术路线
First participation in two hidden, harder projects, passing both without signs of benchmark memorization, while DeepSeek & GLM-5.1 failed
tags: 好数字·好信源 · → daily
📷 原图
- <a id="atom-14bc9b8b80faf258"></a>🟦 2026-06-14
X·@ZhihuFrontier · token consumption compared to Opus 4.8 on same project · 成本
GLM-5.2: 557 tool calls, 170K output tokens; Opus 4.8: 564 tool calls, 260K output tokens; comparable results, significantly lower token consumption
tags: 好数字·好信源 · → daily
value: qty=170K vs 260K output tokens
📷 原图
- <a id="atom-f0f5c490449406b7"></a>🟦 2026-06-26
群 · 定价 · 定价/竞争
GLM 5.2 输入 $0.5/M,输出 $4.5/M;而 Opus 4.6 为 $5/$25
tags: 好数字 · → daily
value: qty=输入 $0.5/M,输出 $4.5/M
- <a id="atom-88f7e51de4eeba10"></a>🟦 2026-07-01
X·@chamath · 与Software Factory配合成本对比 · 成本对比/模型效率
Pairing our Software Factory with the cheaper GLM 5.2 model cut costs 16.4×, though it ran 3× slower than Opus 4.8 alone.
tags: 好数字·好观点 · → daily
🟥 多头 takes (bullish) (30)
- <a id="atom-e8cb716715a117ca"></a>🟥 2026-06-18
X·@FredaDuan · 实际成本约为 Opus 4.8 的 20%-35% · 定价权/盈利能力
Effective cost seems to be around 20-35% of $Opus 4.8, depending on workload.
tags: 好数字·好观点 · → daily
[图: 各大AI大模型(如GLM-5.2、Claude Opus、GPT-5.5等)的API调用价格对比表 — GLM-5.2输入价格: $1.40; Claude Opus 4.8输出价格: $25.00; GPT-5.5 API输出价格: $30.00; DeepSeek V4 Pro输入价格: $0.435; DeepSeek V4 Flash输入价格
📷 原图
- <a id="atom-ae626b9dcfb007a4"></a>🟥 2026-06-18
X·@FredaDuan · 在95%缓存假设下比 Opus 4.8 便宜约2.67倍 · 价格动态
$GLM is ~2.67x cheaper.
tags: 好数字·好观点 · → daily
[图: 各大AI大模型(如GLM-5.2、Claude Opus、GPT-5.5等)的API调用价格对比表 — GLM-5.2输入价格: $1.40; Claude Opus 4.8输出价格: $25.00; GPT-5.5 API输出价格: $30.00; DeepSeek V4 Pro输入价格: $0.435; DeepSeek V4 Flash输入价格
📷 原图
- <a id="atom-945be5d068a09192"></a>🟥 2026-06-18
X·@FredaDuan · 在90%缓存假设下比 Opus 4.8 便宜约3.72倍 · 价格动态
Assuming 90% cached, $GLM is ~3.72x cheaper.
tags: 好数字·好观点 · → daily
[图: 各大AI大模型(如GLM-5.2、Claude Opus、GPT-5.5等)的API调用价格对比表 — GLM-5.2输入价格: $1.40; Claude Opus 4.8输出价格: $25.00; GPT-5.5 API输出价格: $30.00; DeepSeek V4 Pro输入价格: $0.435; DeepSeek V4 Flash输入价格
📷 原图
- <a id="atom-fa0554377b6f502f"></a>🟥 2026-06-18
X·@darkp0rt · 在适当调整后实现 Opus 95% 能力,成本降低 80% · 价格动态/技术路线
after a little work we’ve got close to 95% capability at ~80% reduced cost.
tags: 好数字·好观点 · → daily
- <a id="atom-257b95b74b568d71"></a>🟥 2026-06-20
X·@rauchg · coding performance · 技术路线
GLM-5.2 is at coding, genuinely impressed, almost shocked, at how good
tags: 好观点·好信源 · → daily
- <a id="atom-0b384cef50f36f23"></a>🟥 2026-06-14
X·@ZhihuFrontier · context window usable at 1M tokens · 技术路线/产品发布
writes ~30% more code than most competing models, yet rarely misses critical implementation details—a strong signal that its 1M-token context window is actually usable
tags: 好观点·好信源 · → daily
value: qty=1M-token
📷 原图
- <a id="atom-7de4b994d1186c71"></a>🟥 2026-06-24
X·@natolambert · is a big deal and will be in policy discussions · 监管政策/技术路线
imo its a big deal and we're going to keep hearing about it, eventually will be in policy discussions more too
tags: 好思考·好观点 · → daily
- <a id="atom-d3ce0ffa6953e8d0"></a>🟥 2026-06-27
X·@pmarca · 匹配并常击败美国实验室模型 · 技术路线/竞争格局
Many smart people/AI insiders are saying GLM-5.2 is the first Chinese AI model to match and often beat the American big lab public AI models with no compromises.
tags: 好观点·好思考 · → daily
- <a id="atom-1ea0dc2c73257190"></a>🟥 2026-06-26
X·@matt_slotnick · 需求前景 · 订单需求/产品发布
GLM 5.2 is a great model and it will have immense demand
tags: 好观点·好思考 · → daily
- <a id="atom-83852ca87029a144"></a>🟥 2026-06-24
X·@Dorialexander · strong enabler of synthetic pipelines · 技术路线/模型发布
GLM-5.2 is likely going to be a strong enabler of synthetic pipelines for smaller specialized agents
tags: 好观点·好思考 · → daily
- <a id="atom-8de20a6c9c9d33e0"></a>🟥 2026-06-24
X·@stevehou · excitement level comparable to when folks were Claudemaxxing · 模型发布/技术路线
I haven't seen this level of excitement and enthusiasm for an agentic model since when folks were Claudemaxxing
tags: 好思考·好观点 · → daily
- <a id="atom-e1a7fede51eb5a02"></a>🟥 2026-06-22
X·@natolambert · compared to DeepSeek moment for agents · 模型发布/技术路线
GLM-5.2 should be “DeepSeek moment” for agents. We enter a new world where the top end of agentic capabilities are available in open models.
tags: 好观点·好思考 · → daily
value: direction=na
- <a id="atom-230d7a457f077d63"></a>🟥 2026-06-20
X·@teortaxesTex · ARC-AGI-2 预期得分 · 技术路线/模型发布
I think GLM 5.2 ought to score at least 50%
tags: 好观点·好思考 · → daily
value: qty=50% · direction=at least
[图: 展示AI模型在ARC-AGI-2测试集上的得分与单次任务成本关系的数据图表 — Kimi K2.5 ARC-AGI-2得分: 11.8%; Kimi K2.5单次任务成本: $0.280]
📷 原图
- <a id="atom-2d830c8de2a6516c"></a>🟥 2026-06-17
X·@karminski3 · 无需搜索直接定位换电站 · 技术路线
_测试中GLM-5.2 完全不用搜索附近的位置, 就能直接去想要到达的地方. 这一切竟然是它在一开始把地图背下来了!
这在我测试的20多个模型中之前是没有一个模型能做到的, 比如之前的模型想去换电站, 那么都要搜一下附近有哪些换电站(这就会浪费一次tool_call), 而GLM-5.2直接就知道换电站的位置! 从来没用过搜索函数._
tags: 好观点·好思考 · → daily
- <a id="atom-23629836c4d2e097"></a>🟥 2026-06-17
X·@karminski3 · 长上下文推理能力出色 · 技术路线
这种一开始就把需要的数据内化到上下文中, 并且能够贯穿整个1M上下文进行推理的能力真的是叹为观止.
tags: 好观点·好思考 · → daily
- <a id="atom-8b6fe795ea27cce7"></a>🟥 2026-06-16
X·@ProximalHQ · 开源模型竞争力评价 · 竞争格局
it is the strongest open-weight model by far
tags: 好观点·好思考 · → daily
- <a id="atom-6c18c73695d1d462"></a>🟥 2026-06-16
X·@jietang · long-horizon task capability leap · 模型能力提升
marks a substantial leap in long-horizon task capability over its predecessor GLM-5.1
tags: 好观点·好思考 · → daily
📷 原图
- <a id="atom-4050e1cd5ac60d3f"></a>🟥 2026-06-25
X·@chrisbarber · has tok/s that's 10x faster than opus · 技术路线
if users can access glm 5.2 at tok/s that's 10x faster than opus
tags: 好数字 · → daily
- <a id="atom-85701d223e3ab1de"></a>🟥 2026-06-25
X·@willccbb · 1000步agentic RL运行成本 · 模型训练成本
we can do a 1000-step agentic RL run on top of GLM-5.2 in 3 days for <$50k, and serving is straightforward
tags: 好数字 · → daily
value: qty=<$50k · date=3天
- <a id="atom-8c8ee5161315e3a4"></a>🟥 2026-06-23
X·@teortaxesTex · 基准测试非常接近 GPT 5.5 或 Opus 4.8,有时领先 · 性能对比
It is very close on benchmarks to 5.5 or 4.8, sometimes ahead.
tags: 好数字 · → daily
- <a id="atom-6863c50fe2424110"></a>🟥 2026-06-20
X·@Yuchenj_UW · 成本对比闭源模型 · 模型成本/定价权
the cost is much lower than closed-source models.
tags: 好数字 · → daily
value: direction=much lower
- <a id="atom-116aab9e18c21000"></a>🟥 2026-06-18
X·@airesearch12 · Coding Index score vs Opus 4.8 · 模型评测
GLM-5.2 more then ten points above Opus 4.8!?!?
tags: 好数字 · → daily
value: qty=ten points above Opus 4.8 · date=na
📷 原图
- <a id="atom-f962717497a76115"></a>🟥 2026-07-01
X·@RoliumGens · 模型能力评估 · 技术路线
GLM-5.2 is NOT worse than gemini 3 pro
tags: 好观点 · → daily
value: direction=不低于Gemini 3 Pro
- <a id="atom-3e09191fed8a7b11"></a>🟥 2026-06-27
X·@willccbb · 未来数月可能发布400-500b参数等效版本 · 模型发布
guessing we get something close to glm-5.2-ish in the 400-500b range in the next few months
tags: 好思考 · → daily
- <a id="atom-b1823a176ad94a97"></a>🟥 2026-06-25
X·@teortaxesTex [老化中] · is a tier above Composer · 竞争格局/模型发布
My guess is that GLM 5.2 is a tier above Composer and they've tested it but are reluctant to publish because Big El won't like that
tags: 好观点 · → daily
- <a id="atom-1dbe3b498f3826d2"></a>🟥 2026-06-22
X·@pstAsiatech · 被AI评论者和研究人员称赞 · 模型发布/技术路线
Pretty much everyone I respect among the AI commentariat and researcher class has praised the model after using it personally.
tags: 好观点·好问题 · → daily
- <a id="atom-223d27c7b7a06de4"></a>🟥 2026-06-22
X·@scyshw6492 · 基准测试低估 GLM 5.2 · 性能对比
GLM 5.2 is the first open source frontier-ish model that I think benchmark underestimates.
tags: 好观点·好问题 · → daily
- <a id="atom-fbc3a0f910d63d07"></a>🟥 2026-06-16
X·@scaling01 · 模型质量评价 · 模型能力/竞争格局
GLM-5.2 genuinely looks like a good model
tags: 好观点 · → daily
📷 原图
- <a id="atom-a6c4491b1309eeb4"></a>🟥 2026-06-26
X·@pstAsiatech · 成本效益 · 定价权/竞争格局
cost-effective coding model
tags: 好观点 · → daily
- <a id="atom-9cc0e5beda7b7323"></a>🟥 2026-06-22
X·@teortaxesTex · 能力 · 模型发布
GLM 5.2 is the first open weights model we've tried on our autoresearch pipeline that's proven capable for real research tasks.
tags: 好观点 · → daily
🟥 空头 takes (bearish) (30)
- <a id="atom-b90eee8076940191"></a>🟥 2026-06-21
群 · 单任务耗时对比 · 模型能力/性能对比
GLM 5.2 单任务 30min vs GPT 5.5/Opus 4.8 的 4min
tags: 好数字·好信源 · → daily
value: qty=30min vs 4min · context=vs GPT 5.5/Opus 4.8
- <a id="atom-2ed6b617b2d455cd"></a>🟥 2026-06-18
X·@GenReasoning · open source SoTA · 模型发布
GLM 5.2 is new open source SoTA, but still loses -30% on average over 5 runs
tags: 好数字·好信源 · → daily
value: qty=-30% · direction=loss
📷 原图
- <a id="atom-9ae3e27ed190cbea"></a>🟥 2026-06-21
X·@subhajitlucky · Frontier Code Bench 上限分数预期 · 模型发布
It's not gonna beat 5.5
tags: 好数字·好观点 · → daily
value: qty=5.5
- <a id="atom-a741cf8eba25664b"></a>🟥 2026-06-24
X·@scaling01 · 落后时间 · 技术路线
implying a 7 month lag
tags: 好数字 · → daily
value: qty=7 months · direction=lag
- <a id="atom-8292c126f50b6dac"></a>🟥 2026-06-20
X·@scaling01 · PostTrainBench 每轮评估探测次数相较于 Opus 4.8 增量 · 竞争格局
meaning GLM is doing ~38% more eval probing per run
tags: 好数字 · → daily
value: direction=增加38%
- <a id="atom-bf1b8ee08b53217e"></a>🟥 2026-06-18
X·@scaling01 · 解题成本 · 定价权
though currently a bit expensive per solve ($20 for 27%)
tags: 好数字 · → daily
value: qty=$20 · date=2026-06-18
- <a id="atom-95b0a7bca433da1e"></a>🟥 2026-06-20
X·@scottstts · 与GPT 5.5和Opus 4.8相比表现差距 · 模型性能对比/技术路线
it's not close to gpt 5.5 or opus 4.8, attention to detail is not great, when you push it hard it fails
tags: 好观点·好思考 · → daily
- <a id="atom-17077f3ad827f2a7"></a>🟥 2026-06-20
X·@ItakGol · open model competitiveness · 竞争格局/模型发布
I think GLM 5.2 is the first real “oh shit” moment for frontier AI labs from the open model world.
tags: 好观点·好思考 · → daily
value: date=2026-06-20
- <a id="atom-1b9e40c0006a3708"></a>🟥 2026-06-20
X·@theo · 成本 vs Opus 4.8 和 GPT-5.5 · 价格动态
it's not cheap. Both Opus 4.8 and GPT-5.5 set to medium are cheaper and smarter than GLM-5.2
tags: 好观点·好思考 · → daily
value: direction=高于
📷 原图
- <a id="atom-8b624c2ba78c9b4b"></a>🟥 2026-06-17
X·@karminski3 · 空间理解是最大短板 · 技术路线
而本次测试暴露出最大的短板则是空间理解. 其实成也萧何败也萧何, 它虽然把换电站的位置都背下来了, 但是去的换电站却不是最近的, 所以虽然记住了, 但是记住了之后在用之前再根据自己当前所在位置推理一下, 他还是没有做到的, 这也是最大的短板了, 强烈建议官方优化一波.
tags: 好观点·好思考 · → daily
- <a id="atom-831a901cc8aa0409"></a>🟥 2026-07-02
X·@scaling01 · 缺乏有效推理能力 · 技术路线/推理能力
GLM doesn't use reasoning effectively
tags: 好观点 · → daily
📷 原图
- <a id="atom-8f20d6fd9f25f447"></a>🟥 2026-07-01
X·@scaling01 · 综合能力排名 · 技术路线
if you include all domains GLM-5.2 is worse than gemini 3 pro
tags: 好观点 · → daily
value: direction=低于Gemini 3 Pro
- <a id="atom-90476f8fc48b8e29"></a>🟥 2026-06-29
X·@_xjdr · concurrent serving difficulty · 模型推理_
there are probably not a whole lot of people i would imagine could serve glm5.2 to more than a handful of concurrent users who's full time job isn't currently at an inference provider
tags: 好思考 · → daily
- <a id="atom-6cc50a5aad145740"></a>🟥 2026-06-20
X·@scaling01 · PostTrainBench 后训练模型的泛化能力 · 技术路线
The post-trained models that come out the other side are probably much worse at everything else.
tags: 好观点 · → daily
- <a id="atom-90202070633f5811"></a>🟥 2026-06-27
X·@Dorialexander · 真实挑战:agentic serving尚未解决 · 技术路线
I think we underestimate the real unresolved challenge of agentic serving.
tags: 好思考 · → daily
- <a id="atom-a9b6ca42f7511a30"></a>🟥 2026-06-27
X·@Hangsiin · suspicion of performance improvement via distillation · 模型评测
This made me even more suspicious that this model’s reported performance improvements may be largely due to distillation
tags: 好思考 · → daily
value: direction=likely overstated
- <a id="atom-972dd2a5bce77e86"></a>🟥 2026-06-26
X·@DrEliDavid · missing from report · 技术路线/竞争格局
GLM 5.2, the top Chinese model, is missing
tags: 好问题 · → daily
value: direction=missing
- <a id="atom-4291f1ac8083787f"></a>🟥 2026-06-24
X·@itsGriznft · hype always moves to the next model before anyone ships much with the last one · 技术路线/订单需求
the hype always moves to the next model before anyone ships much with the last one
tags: 好思考 · → daily
- <a id="atom-f8a9a86ac2828039"></a>🟥 2026-06-24
X·@natolambert · 成本与Opus平齐 · 竞争格局/盈利能力
GLM 5.2 being on the Opus frontier for cost of CursorBench is what drives frontier lab margins down
tags: 好思考 · → daily
value: qty= · date=
- <a id="atom-d851cbd2cd1d8d53"></a>🟥 2026-06-21
X·@teortaxesTex · 成本效益比较 · 定价权/竞争格局
GLM-5.2在定价上的成本效益不如中等难度的前沿模型
tags: 好思考 · → daily
📷 原图
- <a id="atom-777836bd7e0c28d6"></a>🟥 2026-07-03
X·@sholtodouglas · cost comparison vs Fable 5 · 价格动态_
GLM 5.2 was significantly more expensive, and wayyyy slower, than Fable 5
tags: 好观点 · → daily
- <a id="atom-0f854b50d1f1f99f"></a>🟥 2026-07-02
X·@MaxKongerskov · 与Claude Fable 5的比较 · 竞争格局
GLM 5.2 is better + runs locally on a m3 ultra 512gb unit
tags: 好观点 · → daily
- <a id="atom-e137d5ae68d7ff07"></a>🟥 2026-06-30
X·@jckwind · 无法替代 Claude · 竞争格局
I tried GLM 5.2 - can't replace Claude, sorry to say.
tags: 好观点 · → daily
- <a id="atom-557d03cda0558ba1"></a>🟥 2026-06-28
X·@banteg · 与 Opus 4.8 对比性能 · 模型性能/竞争格局
glm 5.2 not just as better opus 4.8 (which it's not)
tags: 好观点 · → daily
value: direction=worse
- <a id="atom-8940f413200575af"></a>🟥 2026-06-28
X·@banteg · 与 Mythos 对比性能 · 模型性能
as on par with mythos (which it's doubly not)
tags: 好观点 · → daily
value: direction=worse
- <a id="atom-addfcfee6f11cfab"></a>🟥 2026-06-27
X·@Hangsiin · performance comparison with DeepSeek · 模型评测
It performed far worse than DeepSeek
tags: 好观点 · → daily
value: direction=worse
- <a id="atom-590990c6ed930895"></a>🟥 2026-06-27
X·@Hangsiin · performance categorization · 模型评测
roughly at the level of previous mediocre Chinese models
tags: 好观点 · → daily
value: direction=at level of previous mediocre Chinese models
- <a id="atom-4fd59fd7c8ceeb67"></a>🟥 2026-06-27
X·@Hangsiin · Max setting response time · 模型评测
With the Max setting, I gave up benchmarking because no response came even after waiting for dozens of minutes
tags: 好观点 · → daily
value: direction=unacceptable
- <a id="atom-ab668e60f2e5870a"></a>🟥 2026-06-26
X·@pigeon_s · 相比GPT-5.5-medium性能劣势 · 竞争格局_
GLM-5.2. It's very good and very impressive for being open source, but GPT-5.5-medium is literally better at the same time as being cheaper
tags: 好观点 · → daily
value: direction=worse
[图: 不同版本GPT模型在不同API成本下的性能评分对比折线图 — GPT-5.6 Sol最高评分: 约31%; GPT-5.6 Sol最高成本: 约$1.90; GPT-5.5最高评分: 约23%; GPT-5.5最高成本: 约$1.25]
📷 原图
- <a id="atom-d02a079d2b89b0b7"></a>🟥 2026-06-25
X·@ZhihuFrontier · 从GRPO转向的原因 · 模型发布
GLM-5.2 moving away from GRPO does not mean GRPO is "bad." It means the task has changed.
tags: 好观点 · → daily
📷 原图
🟥 中性 takes (neutral) (20)
- <a id="atom-fa32a3f07884a55d"></a>🟥 2026-06-21
X·@rationaleist · minimum parameters for capabilities · 技术路线
I was thinking 1.5T class, which should be manageable.
tags: 好观点·好数字 · → daily
value: qty=1.5T class
- <a id="atom-f86cc8e1cc6ba236"></a>🟥 2026-06-20
X·@teortaxesTex · Frontier Code Bench 分数预期范围 · 模型发布
20-27 sounds reasonable
tags: 好数字·好观点 · → daily
value: qty=20-27
- <a id="atom-d816d715434dda87"></a>🟥 2026-06-18
X·@rasbt [老化中] · 开放权重模型性能 · 模型性能
The best open-weight model today.
tags: 好观点·好思考 · → daily
📷 原图
- <a id="atom-81ee61335aa59435"></a>🟥 2026-06-18
X·@rasbt [老化中] · 开放权重模型性能 · 模型性能/基准测试
most capable open-weight LLM when averaged over all major benchmarks (reasoning, coding, logic, tool use, math, knowledge)
tags: 好观点·好思考 · → daily
- <a id="atom-c7ce448fc17541b9"></a>🟥 2026-06-28
X·@banteg · 相对 GPT-5.3-codex 对比性能 · 模型性能/技术路线
i think it's closer to gpt-5.3-codex than anything, and it makes up for lower intelligence with much longer thinking
tags: 好观点·好思考 · → daily
- <a id="atom-72314ae3c3aa3e8a"></a>🟥 2026-06-28
X·@xlr8harder · 与 Opus 4.5 对比性能 · 模型性能/测试评估
seems opus 4.5-ish in my tests so far, but capabilities are jagged as always so better and worse in various ways
tags: 好观点·好思考 · → daily
- <a id="atom-a0547757481a99f5"></a>🟥 2026-06-24
X·@ZhihuFrontier · RL训练方法迁移 · 技术路线
GLM-5.2 dropping GRPO does not mean GRPO is "bad."
tags: 好观点·好思考 · → daily
📷 原图
- <a id="atom-2ca149482bf471b6"></a>🟥 2026-06-24
X·@ZhihuFrontier · RL训练方法选择原因 · 技术路线
When rollouts get longer, environments get noisier, and credit assignment gets harder, PPO + value modeling starts looking useful again.
tags: 好观点·好思考 · → daily
📷 原图
- <a id="atom-5ed35a07b7872c6e"></a>🟥 2026-06-17
X·@thefirehacker · 训练一个大模型的上限成本估计范围 · 资本开支
estimated cost range for training a very large GLM 5.2 is 150-250 million USD, this is the cost on the upper side
tags: 好数字 · → daily
value: qty=150-250 million USD · date=NA
- <a id="atom-27e639d33f22f8ea"></a>🟥 2026-06-17
X·@teortaxesTex · 性能水平 · 技术水平
GLM 5.2 is around Opus 4.7-4.8 level
tags: 好数字 · → daily
value: qty=Opus 4.7-4.8 · date=Jun 2026
📷 原图
- <a id="atom-3a834e42dae3c201"></a>🟥 2026-06-17
X·@teortaxesTex · 差距时长 · 竞争格局
GLM 5.2 points to a 7 months gap currently
tags: 好数字 · → daily
value: qty=7 months
📷 原图
- <a id="atom-2baac21bba32d9bd"></a>🟥 2026-06-14
X·@sakurayukiai · 开源的1M上下文KV缓存4-bit量化后需约40GB显存 · 推理优化/内存需求
running 1M tokens locally means a 4-bit quantized KV cache is still ~40GB of VRAM
tags: 好数字 · → daily
- <a id="atom-58a476e2e36d6ab5"></a>🟥 2026-06-27
X·@scaling01 · 推理效率与GPT-5-medium和Gemini 3 Flash接近 · 推理效率
in terms of reasoning efficiency it's very close to GPT-5-medium and Gemini 3 Flash
tags: 好观点 · → daily
📷 原图
- <a id="atom-68e2ab614b123785"></a>🟥 2026-06-24
X·@scaling01 · 性能对比 · 模型性能/技术路线
GLM-5.2 is as strong as Opus 4.5 and GPT-5.2 implying a 7 month lag
tags: 好观点 · → daily
value: relation=as strong as · compared_entities=['Opus 4.5', 'GPT-5.2']
- <a id="atom-63bf6e9a1e39f72e"></a>🟥 2026-06-24
X·@Miles_Brundage · 有 talented ppl · 竞争格局
They have talented ppl
tags: 好信源 · → daily
- <a id="atom-2ba4f359ddd1db9e"></a>🟥 2026-06-21
X·@woke8yearold · 中国政府强制不使用某些网络安全数据 · 监管政策
Rumors are the Chinese government had zai not use certain cybersecurity data in glm 5.2.
tags: 好信源 · → daily
- <a id="atom-d308804615cfd84a"></a>🟥 2026-06-13
X·@zephyr_z9 · 可用于测试能力 · 模型发布
U can try v4, K2.7 & GLM 5.2 to test this
tags: 好观点 · → daily
- <a id="atom-adcc8e097c7e7b51"></a>🟥 2026-06-30
X·@_xjdr · abusive usage skewed stats · 使用情况_
we had some pretty abusive use in the first 48 hours which i think skewed the stats quite a bit on cache hits and ttft p95 latency
tags: 好思考 · → daily
- <a id="atom-48c8b7b63bdfb1e2"></a>🟥 2026-06-26
X·@keennay · KV cache size unknown prior to testing on 8x RTX Pro 6000s · 技术路线/供给产能
I found it quite a headache on 8x RTX Pro 6000s, as I hadn't any idea how much space KV Cache on that model alone consumed until now
tags: 好思考 · → daily
value: text=I found it quite a headache on 8x RTX Pro 6000s, as I hadn't any idea how much space KV Cache on that model alone consumed until now
- <a id="atom-63dd1b16e5393ebe"></a>🟥 2026-06-24
X·@stochasticchasm · 部分编码器采用因果架构且与LLM主流因果方案存在差异 · 技术路线
this one is like making part of the encoder causal lol. and LLMs are primarily causal
tags: 好思考 · → daily
⏱ 时间轴 (近 20)
- 🟦 2026-07-03 ·
fact · ECI score update status · → daily
- 🟥 2026-07-03 ·
narrative · ECI score change after update · → daily
- 🟦 2026-07-03 ·
fact · 代码审查能力 · → daily
- 🟥 2026-07-03 ·
narrative · native precision version more bullish than quantized equival · → daily
- 🟥 2026-07-03 ·
narrative · native precision version preferred over quantized · → daily
- 🟥 2026-07-03 ·
fact · cost comparison vs Fable 5 · → daily
- 🟥 2026-07-02 ·
narrative · ECI score depression from missing eval scores · → daily
- 🟥 2026-07-02 ·
narrative · GBAEval score · → daily
- 🟦 2026-07-02 ·
fact · ECI score without GBAEval · → daily
- 🟥 2026-07-02 ·
narrative · 与Claude Fable 5的比较 · → daily
- 🟦 2026-07-02 ·
fact · 模型发布 · → daily
- 🟥 2026-07-02 ·
narrative · coding and agent capabilities rival leading U.S. offerings · → daily
- 🟥 2026-07-02 ·
narrative · 模型层级对标 · → daily
- 🟦 2026-07-02 ·
fact · 价格比较 · → daily
- 🟦 2026-07-02 ·
fact · 速度比较 · → daily
- 🟦 2026-07-02 ·
fact · running status · → daily
- 🟦 2026-07-02 ·
fact · 开源状态 · → daily
- 🟦 2026-07-02 ·
fact · 许可证类型 · → daily
- 🟦 2026-07-02 ·
fact · 模型格式 · → daily
- 🟦 2026-07-02 ·
fact · long-context MRCR 得分低于 Opus 4.8 和 GPT-5.5 · → daily
← 实体目录 · 系统日志