以下部分引用 AI 总结,现阶段 AI 仍然有幻觉。请以内容的原文为准。
数据新鲜度群聊 07-14 ✓卖方 07-14 ✓Wrap 断档15天 ✕News 07-15 ✓X 断档12天 ✕页面生成 07-15
Claude Mythos Preview
40 atoms · 跨 11 天 · 首见 2026-04-30 · 最近 2026-06-25
三色: 🟦 fact 37 · 🟥 take 3 · stance ▲12/▼1/◆17
叙事状态: 🍂 fading (退潮) · 💤 沉寂 · 策展近14d 0 atom
状态只算低频策展源 (群/卖方); X firehose 仅作背景音量
来源: X 24 · 卖 16
时态: fresh:40
标签: 好数字:33 · 好信源:10 · 好观点:2 · 好思考:1
🟨 AI 综合 · junior analyst 概览
展开 AI 综合 (灰色 · 非市场结论 · 点击数字溯源到原 atom)
model: deepseek-chat · 2026-07-12 · 默认折叠
Claude Mythos Preview 在 METR 80% 成功率基准上时间视野是次优模型的 2 倍以上
¹, 其具体时间达 3 小时 6 分钟
¹. 在 Agentic coding 基准上得 77.8%
¹, Computer use 达 85.4%
¹, 多学科推理 (无工具) 为 56.8%
¹. 开源追赶当前前沿需超 12 个月
¹, 且合作伙伴几周测试即发现数千个高危漏洞
¹, 边际成本进一步推低
¹. 较低定价的竞品价格不及该模型一半
¹, 但新检查点在两个网络靶场首次完成双杀
¹, 且是 UK AISI 低 token 限制下唯一全清任务的模型
¹.
💢 核心分歧 (1 轴)
1. 代码能力霸主 vs 作弊背叛信任
*topic: 技术路线/安全事故/模型能力/产品能力/产品发布 · 9 bull vs 1 bear*
🟢 bullish 侧:
- <a id="atom-c50eca1759a0683c"></a>2026-05-07
X·@AnthropicAI · 模型能力描述 · 模型发布/技术路线 +同日1条
Claude Mythos Preview is our most powerful coding model
tags: 好信源 · → daily
- <a id="atom-589d199417313ceb"></a>2026-05-08
X·@alexalbert_ · METR 80% success rate benchmark time horizon · 模型性能/技术路线_
An early Claude Mythos Preview snapshot we provided METR has a time horizon of more than 2x the next best model on their 80% success rate benchmark
tags: 好数字·好信源 · → daily
value: qty=~3 hours · date=2026-05-08
[图: 展示不同大语言模型(LLM)在80%成功率下可完成软件任务的人类工作时间跨度随发布时间变化的趋势图表 — Claude Mythos Preview (early): 约3小时; Gemini 3.1 Pro: 约1.5小时; GPT-5.2 (high): 约1.1小时; o3: 约30分钟]
📷 原图
- <a id="atom-5ba65096de1615cb"></a>2026-05-13
X·@bioshok3 · 首次完成两个网络靶场中的第二个 · 模型能力
モデルが2つのサイバーレンジのうち2つ目を完了したのは今回が初めて。
tags: 好数字 · → daily
📷 原图
- <a id="atom-aac13bc5ace1a831"></a>2026-05-13
X·@logangraham · XBOW 评价 · 技术路线/产品发布 +同日3条
XBOW tested it on their offensive security benchmarks, finding "token-for-token, unprecedented precision." It's the only model to succeed at subtle V8 sandbox work.
tags: 好信源 · → daily
- <a id="atom-b9d9dc802fd15ddc"></a>2026-06-04
X·@FundaAI · 对漏洞发现边际成本影响 · 技术路线
Claude Mythos Preview 可能将漏洞发现的边际成本进一步推低
tags: 好观点 · → daily
value: direction=下降
📷 原图
🔴 bearish 侧:
- <a id="atom-8172f25d2e9b4037"></a>2026-05-07
X·@AnthropicAI · cheated on coding task · 安全事故/技术路线
Claude Mythos Preview cheated on a coding task by breaking rules, then added misleading code as a coverup
tags: 好信源 · → daily
📷 原图
展开 2 条中性
- <a id="atom-68b81110e8175d79"></a>2026-04-30
X·@teortaxesTex · benchmark pass rate average · 模型发布/技术路线
Mythos Preview 68.6% (±8.7%)
tags: 好数字 · → daily
value: qty=68.6% · date=2026-04-30
📷 原图
- <a id="atom-96415fc68d16ab28"></a>2026-05-08
X·@scaling01 · METR time horizons · 模型发布/技术路线
Claude Mythos Preview's METR time horizons AT LEAST 16 hours
tags: 好数字 · → daily
value: qty=16 hours · direction=at least
📷 原图
🟦 客观事实 (facts) (30)
- <a id="atom-399228da1b97198e"></a>🟦 2026-06-09
X·@btibor91 ⭐ · 定价对比 · 定价
less than half the price of Claude Mythos Preview
tags: 好数字·好信源 · → daily
value: qty=less than half the price of Claude Mythos Preview
📷 原图
- <a id="atom-3916abd010ebca0b"></a>🟦 2026-05-13
X·@bioshok3 ⭐ · 新检查点完成两个网络靶场 · 模型能力
新しいMythos Previewチェックポイントが、サイバーレンジの両方を完了。「The Last Ones」レンジは10回の試行のうち6回(前のチェックポイントでは3回)で、これまで未解決だった「Cooling Tower」レンジは10回の試行のうち3回で解決。
tags: 好数字·好信源 · → daily
value: qty=10回试行的6回和3回 · date=2026-05-13
📷 原图
- <a id="atom-049932ffe9933f4f"></a>🟦 2026-05-13
X·@logangraham ⭐ · UK AISI 低 token 上限下表现 · 技术路线/产品发布
Mythos is also the only model that clears every one of their tasks estimated over 8 hours under their deliberately low 2.5M-token cap.
tags: 好数字·好信源 · → daily
value: qty=2.5M
- <a id="atom-acb9ea96157737c1"></a>🟦 2026-05-08
X·@scaling01 ⭐ · 80% METR time horizons · 模型发布
80% METR time horizons for Claude Mythos Preview: 3 hours 6 minutes
tags: 好数字·好信源 · → daily
value: qty=3 hours 6 minutes · date=2026-05-08
[图: 大语言模型在METR软件任务上的时间跨度(Time Horizon)发展趋势图表 — Claude Mythos Preview (early) 时间跨度: 3小时6分钟; Gemini 3.1 Pro 时间跨度: 约1.8小时; GPT-5.2 (high) 时间跨度: 约1.1小时]
📷 原图
- <a id="atom-589d199417313ceb"></a>🟦 2026-05-08
X·@alexalbert_ · METR 80% success rate benchmark time horizon · 模型性能/技术路线_
An early Claude Mythos Preview snapshot we provided METR has a time horizon of more than 2x the next best model on their 80% success rate benchmark
tags: 好数字·好信源 · → daily
value: qty=~3 hours · date=2026-05-08
[图: 展示不同大语言模型(LLM)在80%成功率下可完成软件任务的人类工作时间跨度随发布时间变化的趋势图表 — Claude Mythos Preview (early): 约3小时; Gemini 3.1 Pro: 约1.5小时; GPT-5.2 (high): 约1.1小时; o3: 约30分钟]
📷 原图
- <a id="atom-37258bb51abe8f8a"></a>🟦 2026-06-25
卖·MS Tom Wigg · Agentic coding (SWE-Bench Pro) score
Claude Mythos Preview achieved a score of 77.8% on the Agentic coding (SWE-Bench Pro) benchmark.
tags: 好数字 · → daily
value: qty=77.8%
- <a id="atom-d962779146bd8a82"></a>🟦 2026-06-25
卖·MS Tom Wigg · Computer use (OSWorld-Verified) score
Claude Mythos Preview achieved a score of 85.4% on the Computer use (OSWorld-Verified) benchmark.
tags: 好数字 · → daily
value: qty=85.4%
- <a id="atom-205e88a53d4f9dae"></a>🟦 2026-06-25
卖·MS Tom Wigg · Multidisciplinary reasoning (Humanity's Last Exam, no tools) score
Claude Mythos Preview achieved a score of 56.8% on the Multidisciplinary reasoning (Humanity's Last Exam, no tools) benchmark.
tags: 好数字 · → daily
value: qty=56.8%
- <a id="atom-97cf87ceb87a727c"></a>🟦 2026-06-25
卖·MS Tom Wigg · Multidisciplinary reasoning (Humanity's Last Exam, with tools) score
Claude Mythos Preview achieved a score of 64.7% on the Multidisciplinary reasoning (Humanity's Last Exam, with tools) benchmark.
tags: 好数字 · → daily
value: qty=64.7%
- <a id="atom-d110d8a48114a854"></a>🟦 2026-06-25
卖·MS Tom Wigg · Biology (BioMysteryBench, hard) score
Claude Mythos Preview achieved a score of 29.6% on the Biology (BioMysteryBench, hard) benchmark.
tags: 好数字 · → daily
value: qty=29.6%
- <a id="atom-5f47bb4c8a1f7a9c"></a>🟦 2026-06-25
卖·MS Tom Wigg · Biology (BioMysteryBench, human solved) score
Claude Mythos Preview achieved a score of 82.6% on the Biology (BioMysteryBench, human solved) benchmark.
tags: 好数字 · → daily
value: qty=82.6%
- <a id="atom-ab70f0f52c0c0567"></a>🟦 2026-06-25
卖·MS Tom Wigg · Cybersecurity (ExploitBench (Cap%)) score
Claude Mythos Preview achieved a score of 69.0% on the Cybersecurity (ExploitBench (Cap%)) benchmark.
tags: 好数字 · → daily
value: qty=69.0%
- <a id="atom-b3b7a53f70ba6594"></a>🟦 2026-06-25
卖·MS Tom Wigg · Health (HealthBench Professional) score
Claude Mythos Preview achieved a score of 64.7% on the Health (HealthBench Professional) benchmark.
tags: 好数字 · → daily
value: qty=64.7%
- <a id="atom-e88a572b3c3a08f2"></a>🟦 2026-06-10
卖·MS Tom Wigg · Agentic coding SWE-Bench Pro score
Claude Mythos Preview scored 77.8% on Agentic coding SWE-Bench Pro.
tags: 好数字 · → daily
value: qty=77.8%
- <a id="atom-b2363738b4dbadb2"></a>🟦 2026-06-10
卖·MS Tom Wigg · Computer use OSWorld-Verified score
Claude Mythos Preview scored 85.4% on Computer use OSWorld-Verified.
tags: 好数字 · → daily
value: qty=85.4%
- <a id="atom-8f1f9808e3b540d7"></a>🟦 2026-06-10
卖·MS Tom Wigg · Multidisciplinary reasoning Humanity's Last Exam (no tools) score
Claude Mythos Preview scored 56.8% on Multidisciplinary reasoning Humanity's Last Exam (no tools).
tags: 好数字 · → daily
value: qty=56.8%
- <a id="atom-86a0ac8c8492a0f1"></a>🟦 2026-06-10
卖·MS Tom Wigg · Multidisciplinary reasoning Humanity's Last Exam (with tools) score
Claude Mythos Preview scored 64.7% on Multidisciplinary reasoning Humanity's Last Exam (with tools).
tags: 好数字 · → daily
value: qty=64.7%
- <a id="atom-6b78bd77fd0aaa6f"></a>🟦 2026-06-10
卖·MS Tom Wigg · Biology BioMysteryBench (hard) score
Claude Mythos Preview scored 29.6% on Biology BioMysteryBench (hard).
tags: 好数字 · → daily
value: qty=29.6%
- <a id="atom-a956886281ed459c"></a>🟦 2026-06-10
卖·MS Tom Wigg · Biology BioMysteryBench (human solved) score
Claude Mythos Preview scored 82.6% on Biology BioMysteryBench (human solved).
tags: 好数字 · → daily
value: qty=82.6%
- <a id="atom-fea7880ad6b55f20"></a>🟦 2026-06-10
卖·MS Tom Wigg · Cybersecurity ExploitBench (Cap%) score
Claude Mythos Preview scored 69.0% on Cybersecurity ExploitBench (Cap%).
tags: 好数字 · → daily
value: qty=69.0%
- <a id="atom-b511a9a1eb300284"></a>🟦 2026-06-10
卖·MS Tom Wigg · Health HealthBench Professional score
Claude Mythos Preview scored 64.7% on Health HealthBench Professional.
tags: 好数字 · → daily
value: qty=64.7%
- <a id="atom-5ba65096de1615cb"></a>🟦 2026-05-13
X·@bioshok3 · 首次完成两个网络靶场中的第二个 · 模型能力
モデルが2つのサイバーレンジのうち2つ目を完了したのは今回が初めて。
tags: 好数字 · → daily
📷 原图
- <a id="atom-96415fc68d16ab28"></a>🟦 2026-05-08
X·@scaling01 · METR time horizons · 模型发布/技术路线
Claude Mythos Preview's METR time horizons AT LEAST 16 hours
tags: 好数字 · → daily
value: qty=16 hours · direction=at least
📷 原图
- <a id="atom-a22f66ac52ef9e07"></a>🟦 2026-05-08
X·@METR_Evals · 50%-time-horizon 估计 · 技术评估
We estimated a 50%-time-horizon of at least 16hrs (95% CI 8.5hrs to 55hrs) on our task suite
tags: 好数字 · → daily
value: qty=至少 16 小时 · date=March 2026
📷 原图
- <a id="atom-126d33e3ab6e752f"></a>🟦 2026-06-09
X·@eliebakouch · 定价 · 价格动态
mythos preview: $25 / input MTok $125 / output MTok
tags: 好数字 · → daily
value: qty=$25/input MTok, $125/output MTok
[图: OpenAI GPT-5.5及5.4系列模型API代币调用价格表 — gpt-5.5-pro短上下文输入价格: $30.00/1M tokens; gpt-5.5-pro短上下文输出价格: $180.00/1M tokens; gpt-5.5-pro长上下文输入价格: $60.00/1M tokens; gpt-5.5-pro长上下文输出价格: $2
📷 原图
- <a id="atom-b33616cc4c9542d3"></a>🟦 2026-06-08
X·@eliebakouch · RSI eval Kernel task multiple · 模型发布
Claude Mythos Preview Kernel task: 399.42×
tags: 好数字 · → daily
value: qty=399.42×
[图: AI研发自动评估规则排除摘要表,对比了Claude Opus 4.5、4.6和Claude Mythos Preview等模型在不同任务上的表现及阈值 — Claude Mythos Preview Kernel task: 399.42×; Claude Mythos Preview LLM training: 51.91×; Claude Myt
📷 原图
- <a id="atom-1733da6d015252a4"></a>🟦 2026-06-08
X·@eliebakouch · RSI eval LLM training multiple · 模型发布
Claude Mythos Preview LLM training: 51.91×
tags: 好数字 · → daily
value: qty=51.91×
[图: AI研发自动评估规则排除摘要表,对比了Claude Opus 4.5、4.6和Claude Mythos Preview等模型在不同任务上的表现及阈值 — Claude Mythos Preview Kernel task: 399.42×; Claude Mythos Preview LLM training: 51.91×; Claude Myt
📷 原图
- <a id="atom-048d12e61a3ef966"></a>🟦 2026-06-08
X·@eliebakouch · RSI eval Quadruped RL score · 模型发布
Claude Mythos Preview Quadruped RL: 30.87
tags: 好数字 · → daily
value: qty=30.87
[图: AI研发自动评估规则排除摘要表,对比了Claude Opus 4.5、4.6和Claude Mythos Preview等模型在不同任务上的表现及阈值 — Claude Mythos Preview Kernel task: 399.42×; Claude Mythos Preview LLM training: 51.91×; Claude Myt
📷 原图
- <a id="atom-357a16173a7b5d87"></a>🟦 2026-06-08
X·@eliebakouch · RSI eval Novel Compiler percentage · 模型发布
Claude Mythos Preview Novel Compiler: 77.2%
tags: 好数字 · → daily
value: qty=77.2%
[图: AI研发自动评估规则排除摘要表,对比了Claude Opus 4.5、4.6和Claude Mythos Preview等模型在不同任务上的表现及阈值 — Claude Mythos Preview Kernel task: 399.42×; Claude Mythos Preview LLM training: 51.91×; Claude Myt
📷 原图
- <a id="atom-11203f061dc86042"></a>🟦 2026-06-08
X·@eliebakouch · RSI eval Internal suite 2 score · 模型发布
Claude Mythos Preview Internal suite 2: 0.65
tags: 好数字 · → daily
value: qty=0.65
[图: AI研发自动评估规则排除摘要表,对比了Claude Opus 4.5、4.6和Claude Mythos Preview等模型在不同任务上的表现及阈值 — Claude Mythos Preview Kernel task: 399.42×; Claude Mythos Preview LLM training: 51.91×; Claude Myt
📷 原图
🟥 多头 takes (bullish) (3)
- <a id="atom-e5a21e81c5e13e28"></a>🟥 2026-05-15
X·@scaling01 ⭐ · 当前性能被开源模型追赶时间 · 模型性能差距/开源追赶速度
it will take >12 months (>April 7th 2027) to catch the current frontier that is Claude Mythos Preview
tags: 好观点·好思考 · → daily
value: date=>April 7th 2027
- <a id="atom-20f7c0244d7cc4e3"></a>🟥 2026-05-13
X·@logangraham · 合作伙伴发现的漏洞数量 · 产品发布/技术路线
In a few weeks of testing, Mythos Preview has helped them find many thousands of (estimated) high + critical severity vulnerabilities, sometimes double what they'd normally find in a year.
tags: 好数字 · → daily
value: qty=many thousands
- <a id="atom-b9d9dc802fd15ddc"></a>🟥 2026-06-04
X·@FundaAI · 对漏洞发现边际成本影响 · 技术路线
Claude Mythos Preview 可能将漏洞发现的边际成本进一步推低
tags: 好观点 · → daily
value: direction=下降
📷 原图
⏱ 时间轴 (近 20)
- 🟦 2026-06-25 ·
fact · Agentic coding (SWE-Bench Pro) score · → daily
- 🟦 2026-06-25 ·
fact · Computer use (OSWorld-Verified) score · → daily
- 🟦 2026-06-25 ·
fact · Multidisciplinary reasoning (Humanity's Last Exam, no tools) · → daily
- 🟦 2026-06-25 ·
fact · Multidisciplinary reasoning (Humanity's Last Exam, with tool · → daily
- 🟦 2026-06-25 ·
fact · Biology (BioMysteryBench, hard) score · → daily
- 🟦 2026-06-25 ·
fact · Biology (BioMysteryBench, human solved) score · → daily
- 🟦 2026-06-25 ·
fact · Cybersecurity (ExploitBench (Cap%)) score · → daily
- 🟦 2026-06-25 ·
fact · Health (HealthBench Professional) score · → daily
- 🟦 2026-06-10 ·
fact · Agentic coding SWE-Bench Pro score · → daily
- 🟦 2026-06-10 ·
fact · Computer use OSWorld-Verified score · → daily
- 🟦 2026-06-10 ·
fact · Multidisciplinary reasoning Humanity's Last Exam (no tools) · → daily
- 🟦 2026-06-10 ·
fact · Multidisciplinary reasoning Humanity's Last Exam (with tools · → daily
- 🟦 2026-06-10 ·
fact · Biology BioMysteryBench (hard) score · → daily
- 🟦 2026-06-10 ·
fact · Biology BioMysteryBench (human solved) score · → daily
- 🟦 2026-06-10 ·
fact · Cybersecurity ExploitBench (Cap%) score · → daily
- 🟦 2026-06-10 ·
fact · Health HealthBench Professional score · → daily
- 🟦 2026-06-09 ·
fact · 定价 · → daily
- 🟦 2026-06-09 ·
fact · 定价对比 · → daily
- 🟦 2026-06-08 ·
fact · RSI eval Kernel task multiple · → daily
- 🟦 2026-06-08 ·
fact · RSI eval LLM training multiple · → daily
← 实体目录 · 系统日志