以下部分引用 AI 总结,现阶段 AI 仍然有幻觉。请以内容的原文为准。
数据新鲜度群聊 07-14 ✓卖方 07-14 ✓Wrap 断档15天 ✕News 07-15 ✓X 断档12天 ✕页面生成 07-15
GLM-5.1
35 atoms · 跨 23 天 · 首见 2026-04-17 · 最近 2026-06-22
三色: 🟦 fact 28 · 🟥 take 7 · stance ▲7/▼3/◆7
叙事状态: 🍂 fading (退潮) · 💤 沉寂 · 策展近14d 0 atom
状态只算低频策展源 (群/卖方); X firehose 仅作背景音量
来源: X 35
时态: fresh:34 · aging:1
标签: 好数字:24 · 好信源:7 · 好观点:7
🟨 AI 综合 · junior analyst 概览
展开 AI 综合 (灰色 · 非市场结论 · 点击数字溯源到原 atom)
model: deepseek-chat · 2026-07-12 · 默认折叠
GLM-5.1 在 SWE-Bench Pro 和 Coding Agent Index 上均位列开源第一 [id=d0447e8b72c0642d,cec80b1542200bd4],编码能力非常强 [id=11810fe2779eac84]。ALE-Bench 表现优于 Grok-4.3 [id=baeb3eca2b5b8c9a],整体水平已与 GPT-5.2-Codex-xhigh 相当 [id=adc3330aefaeeea3]。不过 FutureSim 显示开源与闭源模型差距巨大 [id=b73d6877f2cd9715],且其上下文窗口较短是一大短板 [id=21cd623830f686bf]。每次任务成本 $2.26 但 token 使用量高达 4.8M [id=1330148a6199d854,5559e1f1d936f8eb],CritPT 评分仅 20.9 [id=1966634d8e8bca36]。CyberGym 得分 68.7%,与 Claude Opus 4.6 相当 [id=7ad8bcaa3c51e66b],并已通过 Harvey 法律基准进行后训练 [id=57836df21cc0c49b]。
🧭 拥挤度 (一人一票): ▲ 3 位作者 (KOL3) vs ▼ 2 位作者 (KOL2)
🟦 客观事实 (facts) (28)
- <a id="atom-1966634d8e8bca36"></a>🟦 2026-06-17
X·@teortaxesTex · CritPT评分 · 模型性能
Artificial Analysis' score for GLM-5.1 on CritPT is ... 20.9
tags: 好数字·好信源 · → daily
value: qty=20.9
📷 原图
- <a id="atom-d0447e8b72c0642d"></a>🟦 2026-05-18
X·@OrcaRouter · SWE-Bench Pro 排名 · 模型发布/技术路线
#1 open-source model on SWE-Bench Pro
tags: 好数字·好信源 · → daily
value: qty=#1 · date=2026-05-18
- <a id="atom-cec80b1542200bd4"></a>🟦 2026-05-11
X·@ArtificialAnlys · 在 Coding Agent Index 得分 · 模型发布/评分对比
GLM-5.1 in Claude Code is the top open-weight result at 53
tags: 好数字 · → daily
value: qty=53
📷 原图
- <a id="atom-1330148a6199d854"></a>🟦 2026-05-11
X·@ArtificialAnlys · 每次任务成本 · 成本效率/模型发布
GLM-5.1 in Claude Code costs $2.26/task
tags: 好数字 · → daily
value: qty=$2.26/task
📷 原图
- <a id="atom-5559e1f1d936f8eb"></a>🟦 2026-05-11
X·@ArtificialAnlys · 每次任务 token 使用量 · 成本效率/模型发布
GLM-5.1 in Claude Code uses the most tokens at 4.8M/task
tags: 好数字 · → daily
value: qty=4.8M/task
📷 原图
- <a id="atom-7ad8bcaa3c51e66b"></a>🟦 2026-06-22
X·@AiBattle_ · CyberGym score · 模型发布
Among open-source models on this benchmark, GLM-5.1 is the top performer, scoring 68.7%, comparable to Claude Opus 4.6
tags: 好数字 · → daily
value: qty=68.7%
📷 原图
- <a id="atom-57836df21cc0c49b"></a>🟦 2026-06-22
X·@gabepereyra · 基于Harvey法律助手基准进行后训练 · 技术路线
post-trained the underlying GLM-5.1 model using reward signal from Harvey's Legal Agent Benchmark (LAB).
tags: 好信源 · → daily
- <a id="atom-b7450ba898d6b816"></a>🟦 2026-06-19
X·@BAI_AGI · 排名2 · 排名动态_
[图: 展示 AI 模型在平台上的流行度及排名变化的排行榜数据面板 — GLM-5.1排名: 2]
tags: 好数字 · → daily
value: qty=2 · date=2026-06-19
[图: 展示 AI 模型在平台上的流行度及排名变化的排行榜数据面板 — MiniMax-M3排名: 1; MiniMax-M3排名变化: +17; GLM-5.1排名: 2; Kimi-K2.5排名: 3; MiniMax-M2.7排名变化: +18]
📷 原图
- <a id="atom-0985902af0d24486"></a>🟦 2026-06-18
X·@teortaxesTex · score improvement between GLM-5.1 and GLM-5.2 · 模型发布/性能基准
10.2% between 5.1 and 5.2
tags: 好数字 · → daily
value: qty=10.2% · date=2026-06-18
📷 原图
- <a id="atom-711522523836b5bb"></a>🟦 2026-06-15
X·@teortaxesTex · 智能指数评分 · 模型性能/技术路线
V4-*Flash*(Max) is equal to GPT-5.4-mini (xhigh) and GLM-5.1
tags: 好数字 · → daily
value: qty=40 · date=2026-06-16
[图: Artificial Analysis 发布的各家大语言模型智能指数对比柱状图 — Claude Fable 5 (with fallback): 60; Claude Opus 4.8 (max): 56; GPT-5.5 (xhigh): 55; Gemini 3.5 Flash: 50; DeepSeek V4 Flash (Max) / GP
📷 原图
- <a id="atom-da4a43a347a0484f"></a>🟦 2026-06-12
X·@Presidentlin · 每月预估请求数 · 出货量
GLM-5.1 每月预估请求数: 4,300
tags: 好数字 · → daily
value: qty=4,300
[图: OpenCode Go 平台不同 AI 模型的服务使用额度与预估请求次数对照表 — 5小时额度限制: $12; 每周额度限制: $30; 每月额度限制: $60; DeepSeek V4 Flash 每月预估请求数: 158,150; GLM-5.1 每月预估请求数: 4,300]
📷 原图
- <a id="atom-ddd7459cc5e3d059"></a>🟦 2026-06-10
X·@mitchellh · cost per task · 价格动态
GLM cost me less than a dollar
tags: 好数字 · → daily
value: qty=less than a dollar
- <a id="atom-735491323231859f"></a>🟦 2026-06-04
X·@arena · Agent Arena 排名 · 模型发布
_#3 @Zai_org: GLM-5.1_
tags: 好数字 · → daily
value: qty=#3
📷 原图
- <a id="atom-e18dc6258d88f8e4"></a>🟦 2026-05-26
X·@vincenzoiozzo · CyberGym score · 模型对比_
GLM-5.1 to 68.7
tags: 好数字 · → daily
- <a id="atom-f329ad26c4828e17"></a>🟦 2026-05-26
X·@vincenzoiozzo · release date relative to GLM-5 · 产品发布_
only ~6 weeks between the two releases
tags: 好数字 · → daily
- <a id="atom-4ca0663b53ea4f11"></a>🟦 2026-05-21
X·@karminski3 · 输出速度 · 技术路线/模型发布
同样的脚本我测了下 glm-5.1 的接口, 输出速度只有 35 tps
tags: 好数字 · → daily
value: qty=35 tps
📷 原图
- <a id="atom-1f4ed82055137b4e"></a>🟦 2026-05-21
X·@karminski3 · 首token延迟 · 技术路线/模型发布
首 token 延迟干到了 9s
tags: 好数字 · → daily
value: qty=9s
📷 原图
- <a id="atom-0efc5f31723a9f82"></a>🟦 2026-05-21
X·@karminski3 · 单次激活参数量 · 技术路线
GLM-5.1 单次激活40B
tags: 好数字 · → daily
value: qty=40B
📷 原图
- <a id="atom-f0df63f08aad133b"></a>🟦 2026-05-21
X·@karminski3 · 显存需求 · 技术路线
按照bf16精度计算, 即使不考虑 kvcache 也要80GB的显存
tags: 好数字 · → daily
value: qty=80GB
📷 原图
- <a id="atom-042340271ac9132c"></a>🟦 2026-05-18
X·@OrcaRouter · 上下文长度 · 模型发布
200K context
tags: 好数字 · → daily
value: qty=200K
- <a id="atom-a4482a4ed4ff5491"></a>🟦 2026-05-06
X·@htihle · WeirdML 得分 · 模型评测
still well behind Kimi-k2.6 and GLM-5.1 at 56% and 57%
tags: 好数字 · → daily
value: qty=57%
📷 原图
- <a id="atom-40468c92f6cdb0ba"></a>🟦 2026-05-02
X·@scaling01 · 推理速度 · 模型性能
GLM-5.1 ... all serve at around 20-30tks/s
tags: 好数字 · → daily
value: qty=20-30 tks/s
- <a id="atom-b16bf02b33b6b869"></a>🟦 2026-05-01
X·@HeMuyu0327 · 性能相比 GLM-5 提升 · 模型发布
GLM-5.1 moving the already SOTA perf 0.25x higher than GLM-5
tags: 好数字 · → daily
value: qty=0.25x · date=2026-05-01
📷 原图
- <a id="atom-7d113c30c90e64fd"></a>🟦 2026-04-17
X·@yishan · 本地推理速度 · 技术路线/模型发布
GLM-5.1 locally is next-level. You can get 14 tok/s on 512GB ultra.
tags: 好数字 · → daily
value: qty=14 tok/s · date=local inference
- <a id="atom-85ca723dc06d9c1a"></a>🟦 2026-06-13
X·@elliotarledge · 作弊行为 · 技术路线
GLM-5.1 banked its number by calling cublasLt (a library wrapper, zero kernel authorship).
tags: 好信源 · → daily
value: qty=1 · direction=调用cublasLt替代真正kernel编写
📷 原图
- <a id="atom-c58c7806f03388bf"></a>🟦 2026-05-18
X·@OrcaRouter · 许可证类型 · 技术路线
MIT licensed
tags: 好信源 · → daily
- <a id="atom-46fc04ad601b56ed"></a>🟦 2026-05-16
X·@interconnectsai · 发布 · 模型发布
GLM-5.1
tags: 好信源 · → daily
value: date=2026-05-16
- <a id="atom-ffd4aba96b222be8"></a>🟦 2026-04-26
X·@rasbt · 发布 · 模型发布
April was a pretty strong month for LLM releases: - GLM-5.1
tags: 好信源 · → daily
📷 原图
🟥 多头 takes (bullish) (4)
- <a id="atom-adc3330aefaeeea3"></a>🟥 2026-04-27
X·@scaling01 [老化中] · 水平 · 技术路线
GLM-5.1 doing very well, now on a level with GPT-5.2-Codex-xhigh
tags: 好数字·好观点 · → daily
value: direction=on a level with GPT-5.2-Codex-xhigh
📷 原图
- <a id="atom-baeb3eca2b5b8c9a"></a>🟥 2026-05-22
X·@scaling01 · ALE-Bench 表现优于 Grok-4.3 · 模型表现
Grok-4.3 is pretty terrible, basically worse than all the frontier chinese models like ... GLM-5.1
tags: 好观点 · → daily
📷 原图
- <a id="atom-11810fe2779eac84"></a>🟥 2026-05-22
X·@CompaCompu · 编码能力 · 产品发布
glm-5.1 is very capable at coding
tags: 好观点 · → daily
📷 原图
- <a id="atom-5afc19d72c9b00a7"></a>🟥 2026-05-20
X·@dreamworks2050 · 输出质量 · 模型性能
when it works it gives good output
tags: 好观点 · → daily
value: qualifier=when it works
🟥 空头 takes (bearish) (2)
- <a id="atom-b73d6877f2cd9715"></a>🟥 2026-06-16
X·@nikhilchandak29 · FutureSim 性能 · 模型性能/竞争格局
The gap between open and closed-weights here is massive!
tags: 好观点 · → daily
value: qty=no better than GLM-5.1 on FutureSim · date=2026-06-16
📷 原图
- <a id="atom-21cd623830f686bf"></a>🟥 2026-05-22
X·@CompaCompu · 上下文窗口较短 · 技术路线
handicap of a shorter context window
tags: 好观点 · → daily
📷 原图
🟥 中性 takes (neutral) (1)
- <a id="atom-762cf1038093dc77"></a>🟥 2026-05-26
X·@vincenzoiozzo · matches Opus 4.7 on vulnerability research across all artifacts · 模型对比_
GLM-5.1, which matches Opus across the board.
tags: 好观点 · → daily
⏱ 时间轴 (近 20)
- 🟦 2026-06-22 ·
fact · CyberGym score · → daily
- 🟦 2026-06-22 ·
fact · 基于Harvey法律助手基准进行后训练 · → daily
- 🟦 2026-06-19 ·
fact · 排名_2 · → daily
- 🟦 2026-06-18 ·
fact · score improvement between GLM-5.1 and GLM-5.2 · → daily
- 🟦 2026-06-17 ·
fact · CritPT评分 · → daily
- 🟥 2026-06-16 ·
fact · FutureSim 性能 · → daily
- 🟦 2026-06-15 ·
fact · 智能指数评分 · → daily
- 🟦 2026-06-13 ·
fact · 作弊行为 · → daily
- 🟦 2026-06-12 ·
fact · 每月预估请求数 · → daily
- 🟦 2026-06-10 ·
fact · cost per task · → daily
- 🟦 2026-06-04 ·
fact · Agent Arena 排名 · → daily
- 🟥 2026-05-26 ·
fact · matches Opus 4.7 on vulnerability research across all artifa · → daily
- 🟦 2026-05-26 ·
fact · CyberGym score · → daily
- 🟦 2026-05-26 ·
fact · release date relative to GLM-5 · → daily
- 🟥 2026-05-22 ·
narrative · 编码能力 · → daily
- 🟥 2026-05-22 ·
narrative · 上下文窗口较短 · → daily
- 🟥 2026-05-22 ·
narrative · ALE-Bench 表现优于 Grok-4.3 · → daily
- 🟦 2026-05-21 ·
fact · 输出速度 · → daily
- 🟦 2026-05-21 ·
fact · 首token延迟 · → daily
- 🟦 2026-05-21 ·
fact · 单次激活参数量 · → daily
← 实体目录 · 系统日志