以下部分引用 AI 总结,现阶段 AI 仍然有幻觉。请以内容的原文为准。
数据新鲜度群聊 07-14 ✓卖方 07-14 ✓Wrap 断档15天 ✕News 07-15 ✓X 断档12天 ✕页面生成 07-15
ArtificialAnlys
X · 45 atoms · 2026-06-23 → 2026-06-25 · 信用档先验 cred3
🟦 fact 44 · 🟥 take 1 · stance ▲0/▼0/◆1
覆盖实体 (top 10)
话题分布 (top 8)
任务耗时 (6) · 能力领跑 (6) · 模型评估 (4) · 基准测试 (4) · 模型表现 (4) · 模型排名 (4) · 模型评测排名 (3) · 智能体表现 (2)
样例 atoms
- 2026-06-24 · AA-Briefcase · Agentic knowledge work can take frontier models over 20 minutes per task, as measured in AA-Briefcase, our new benchmark
- 2026-06-24 · Claude Opus 4.8 · Claude Opus 4.8 is the highest-scoring available model, but it is also one of the slowest, taking ~23 minutes per task on average
- 2026-06-24 · Claude Opus 4.8 · Claude Opus 4.8 is the highest-scoring available model, but it is also one of the slowest, taking ~23 minutes per task on average
- 2026-06-24 · GPT-5.5 (xhigh) · GPT-5.5 (xhigh) in particular stands out as one of the most efficient top-performing models, using around half the time per task of Opus 4.8 (11 minutes) while ranking top 5 on the
- 2026-06-24 · GPT-5.5 (xhigh) · GPT-5.5 (xhigh) in particular stands out as one of the most efficient top-performing models, using around half the time per task of Opus 4.8 (11 minutes) while ranking top 5 on the
← Author 索引 · 实体目录