55 atoms · 跨 24 天 · 首见 2026-04-15 · 最近 2026-07-03
三色: 🟦 fact 36 · 🟥 take 19 · stance ▲12/▼5/◆25
叙事状态: 🍂 fading (退潮) · 💤 沉寂 · 策展近14d 0 atom · 🐦 X 近14d 3
状态只算低频策展源 (群/卖方); X firehose 仅作背景音量
来源: X 35 · 卖 20
时态: fresh:50 · aging:4 · stale:1
标签: 好数字:39 · 好观点:13 · 好信源:3 · 好问题:2 · 好思考:1
🧭 拥挤度 (一人一票): ▲ 9 位作者 (KOL9) vs ▼ 5 位作者 (KOL5)
⚖️ 多头 9/9 来自KOL
XX·@scaling01 ⭐ · 80% METR time horizons · 模型发布qty=about 1.8 hours · date=2026-05-08X·@HuggingPapers · Claw-Eval-Live benchmark pass rate · 基准测试/模型性能qty=53.3% · date=2026-05-03卖·MS Tom Wigg · Agentic coding (SWE-Bench Pro) scoreqty=54.2%卖·MS Tom Wigg · Knowledge work (GDPval-AA) scoreqty=1314卖·MS Tom Wigg · Knowledge work vision (GDP.pdf) scoreqty=16.7%卖·MS Tom Wigg · Spatial reasoning (Blueprint-Bench 2) scoreqty=26.5%卖·MS Tom Wigg · Tool use (AutomationBench) scoreqty=9.6%卖·MS Tom Wigg · Computer use (OSWorld-Verified) scoreqty=76.2%卖·MS Tom Wigg · Legal (Legal Agent Benchmark) scoreqty=0.0%卖·MS Tom Wigg · Multidisciplinary reasoning (Humanity's Last Exam, no tools) scoreqty=44.4%卖·MS Tom Wigg · Multidisciplinary reasoning (Humanity's Last Exam, with tools) scoreqty=51.4%卖·MS Tom Wigg · Agentic coding (Terminal-Bench 2.1) scoreqty=70.7%卖·MS Tom Wigg · Agentic coding SWE-Bench Pro scoreqty=54.2%卖·MS Tom Wigg · Knowledge work GDPval-AA scoreqty=1314卖·MS Tom Wigg · Knowledge work vision GDP.pdf scoreqty=16.7%卖·MS Tom Wigg · Spatial reasoning Blueprint-Bench 2 scoreqty=26.5%卖·MS Tom Wigg · Tool use AutomationBench scoreqty=9.6%卖·MS Tom Wigg · Computer use OSWorld-Verified scoreqty=76.2%卖·MS Tom Wigg · Legal Agent Benchmark scoreqty=0.0%卖·MS Tom Wigg · Multidisciplinary reasoning Humanity's Last Exam (no tools) scoreqty=44.4%卖·MS Tom Wigg · Multidisciplinary reasoning Humanity's Last Exam (with tools) scoreqty=51.4%卖·MS Tom Wigg · Agentic coding Terminal-Bench 2.1 scoreqty=70.7%X·@Techmeme · 政治提示左倾回应频率 · AI模型direction=左倾X·@ArtificialAnlys · 在 Coding Agent Index 得分 · 模型发布/评分对比qty=43X·@hsu_steve · 在 FrontierMath Tier 4 基准测试自主得分 · 模型发布qty=19% · date=2026-05-11X·@ArtificialAnlys · AA-Omniscience 指数评分 · 模型能力qty=33 · date=2026-04-18X·@ArtificialAnlys · 输出 token 消耗(Intelligence Index 测试) · 模型能力qty=57M · date=2026-04-18X·@andonlabs · 亏损金额 · 盈利能力qty=$6k · currency=USD · date=2026-07-01T16:23:23X·@jumperz · DeepSWE测试通过率 · 模型发布/技术路线qty=10%X·@theo · 在 Skatebench 2 的视觉得分 · 测评基准/模型表现qty=98%X·@dejavucoder · 平均基准提升排名 · 基准测试X·@scaling01 · ALE-Bench 表现优于 Gemini 3.5 Flash · 模型表现X·@wenhaocha1 · OpenDeepThink 框架下的表现 · 技术路线/模型发布X·@KuittinenPetri · niche knowledge · 竞争格局X·@themmyleke · 作为实现工具足够好 · 模型能力/工具适用X·@mweinbach · benchmark category performance · 竞争格局X·@ZackKorman · consistently says 'yea that's good' on Rothko reproduction critique · 模型评估X·@teortaxesTex [老旧] · 视觉能力评价 · 模型性能X·@emollick [老化中] · model quality assessment · 产品发布/技术路线X·@MechanizeWork · Game Boy Advance emulator 工作结果 · 模型评测X·@KuittinenPetri · agentic coding suitability · 技术路线X·@anshelsag · benchmark comparison · 竞争格局X·@4m473r45u · post-training quality · 技术路线X·@morqon · thinks everything user does is fantastic, quite unusable for critique · 模型评估X·@teortaxesTex · 曾是强大模型 · 技术路线/产品发布date=2023-2024narrative · agentic coding suitability · → dailynarrative · niche knowledge · → dailyfact · 亏损金额 · → dailyfact · Agentic coding (SWE-Bench Pro) score · → dailyfact · Knowledge work (GDPval-AA) score · → dailyfact · Knowledge work vision (GDP.pdf) score · → dailyfact · Spatial reasoning (Blueprint-Bench 2) score · → dailyfact · Tool use (AutomationBench) score · → dailyfact · Computer use (OSWorld-Verified) score · → dailyfact · Legal (Legal Agent Benchmark) score · → dailyfact · Multidisciplinary reasoning (Humanity's Last Exam, no tools) · → dailyfact · Multidisciplinary reasoning (Humanity's Last Exam, with tool · → dailyfact · Agentic coding (Terminal-Bench 2.1) score · → dailyfact · 政治提示左倾回应频率 · → dailynarrative · testing status · → dailyfact · Agentic coding SWE-Bench Pro score · → dailyfact · Knowledge work GDPval-AA score · → dailyfact · Knowledge work vision GDP.pdf score · → dailyfact · Spatial reasoning Blueprint-Bench 2 score · → dailyfact · Tool use AutomationBench score · → daily