以下部分引用 AI 总结,现阶段 AI 仍然有幻觉。请以内容的原文为准。
数据新鲜度群聊 07-14 ✓卖方 07-14 ✓Wrap 断档15天 ✕News 07-15 ✓X 断档12天 ✕页面生成 07-15
EpochAIResearch
22 atoms · 跨 6 天 · 首见 2026-04-23 · 最近 2026-07-02
三色: 🟦 fact 16 · 🟥 take 6 · stance ▲1/▼4/◆12
来源: X 22
时态: fresh:21 · aging:1
标签: 好信源:11 · 好数字:9 · 好观点:4 · 好思考:3 · 好问题:1
叙事 (narrative) (6)
- 2026-07-02 · claims AI on-the-fly learning has economic and safety implications
If AI can learn on the fly, it becomes much more general-purpose. This has economic implications (learning on the job) as well as safety consequences (developing dangerous capabilities post-release)
tags: 好观点·好思考 · → daily
- 2026-06-12 · future of math benchmarking lies in open problems from real research
We believe the future of math benchmarking lies in open problems drawn from real research, like those we’ve collected in FrontierMath: Open Problems.
tags: 好观点·好思考 · → daily
- 2026-06-12 · FrontierMath Tier 1-4 approaching saturation
FrontierMath: Tiers 1–4 is now approaching saturation.
tags: 好思考 · → daily
- 2026-06-12 · Anthropic 数学能力快速提升的趋势
This continues a streak of Anthropic models improving rapidly at math.
tags: 好观点 · → daily
- 2026-06-19 · pretraining data scrutiny question
They say “Look at your data!" But when is the last time you looked at the pretraining data?
tags: 好问题 · → daily
- 2026-05-05
[老化中] · 发布观点:经典推理基准测试需要演化
Reasoning benchmarks aren’t dead, but they do need to evolve.
tags: 好观点 · → daily
事实 (fact) (16)
- 2026-06-12 ⭐ · Claude Fable 5 FrontierMath Tiers 1–3 得分
Claude Fable 5 scores very well on FrontierMath: Tiers 1–3 (v2), reaching 87%
tags: 好数字·好信源 · → daily
value: qty=87%
- 2026-06-12 ⭐ · Claude Fable 5 FrontierMath Tier 4 得分
88% on Tier 4
tags: 好数字·好信源 · → daily
value: qty=88%
- 2026-07-02 · 披露2026年6月发现约1500个高等严重CVE
In June 2026, 21 notable organizations disclosed ~1,500 high- and critical-severity CVEs, over 3.5× the previous monthly record set before
tags: 好数字·好信源 · → daily
value: qty=~1500 · date=2026-06
- 2026-04-23 · 数据点置信度声明
we are only 80% confident that any given data point is within 6 months of the true value
tags: 好信源·好数字 · → daily
value: qty=80%
- 2026-07-02 · introduces EBR-bench
Introducing EBR-bench, our new benchmark to measure on-the-fly learning
tags: 好信源 · → daily
- 2026-07-02 · EBR-bench shows no on-the-fly learning across playthroughs
We see no on-the-fly learning
tags: 好信源 · → daily
- 2026-07-02 · models fall short of expert human performance on fatigue mechanic
Models do better than random, but fall short of expert human performance
tags: 好信源 · → daily
- 2026-07-02 · models explore only fraction of deck archetypes
Models explore only a fraction of them
tags: 好信源 · → daily
- 2026-07-02 · models show no ability to get better with practice even with strategy guide
Even if we give them a full strategy guide ... models improve only modestly and still show no ability to get better with practice
tags: 好信源 · → daily
- 2026-06-29 · 提供免费AI加速器市场份额数据
@EpochAIResearch has free data on AI accelerator market share, compute ownership, DC builds and costs, lab P&Ls, etc.
tags: 好信源 · → daily
- 2026-06-12 · FrontierMath Tier 1-4 v2 audit resolved errors in 42% of problems
We concluded an audit that addressed errors in 42% of problems.
tags: 好数字 · → daily
value: qty=42% · date=2026-06-12
- 2026-06-12 · FrontierMath Tier 1-3 leader GPT-5.5 xhigh score 85%
The current leaders are GPT-5.5 (xhigh) with 85% on Tiers 1–3
tags: 好数字 · → daily
value: qty=85% · date=2026-06-12
- 2026-06-12 · FrontierMath Tier 4 leader Google AI co-mathematician score 76%
Google’s AI co-mathematician with 76% on Tier 4
tags: 好数字 · → daily
value: qty=76% · date=2026-06-12
- 2026-06-12 · OpenAI funded FrontierMath Tiers 1-4 development and has exclusive access to 80%
Note that OpenAI funded the development of Tiers 1–4 and has exclusive access to about 80% of it
tags: 好信源 · → daily
value: qty=80% · date=2026-06-12
- 2026-06-12 · FrontierMath Tier 1-3 removed 5 problems (2%)
We also removed 5 problems (2%) from Tiers 1–3
tags: 好数字 · → daily
value: qty=2% · date=2026-06-12
- 2026-06-12 · FrontierMath Tier 4 removed 7 problems (15%)
removed 7 (15%) from Tier 4
tags: 好数字 · → daily
value: qty=15% · date=2026-06-12
时间轴 (近 20)
- 2026-07-02 ·
fact · introduces EBR-bench · → daily
- 2026-07-02 ·
narrative · claims AI on-the-fly learning has economic and safety implic · → daily
- 2026-07-02 ·
fact · EBR-bench shows no on-the-fly learning across playthroughs · → daily
- 2026-07-02 ·
fact · models fall short of expert human performance on fatigue mec · → daily
- 2026-07-02 ·
fact · models explore only fraction of deck archetypes · → daily
- 2026-07-02 ·
fact · models show no ability to get better with practice even with · → daily
- 2026-07-02 ·
fact · 披露2026年6月发现约1500个高等严重CVE · → daily
- 2026-06-29 ·
fact · 提供免费AI加速器市场份额数据 · → daily
- 2026-06-19 ·
narrative · pretraining data scrutiny question · → daily
- 2026-06-12 ·
fact · FrontierMath Tier 1-4 v2 audit resolved errors in 42% of pro · → daily
- 2026-06-12 ·
fact · FrontierMath Tier 1-3 leader GPT-5.5 xhigh score 85% · → daily
- 2026-06-12 ·
fact · FrontierMath Tier 4 leader Google AI co-mathematician score · → daily
- 2026-06-12 ·
fact · OpenAI funded FrontierMath Tiers 1-4 development and has exc · → daily
- 2026-06-12 ·
fact · FrontierMath Tier 1-3 removed 5 problems (2%) · → daily
- 2026-06-12 ·
fact · FrontierMath Tier 4 removed 7 problems (15%) · → daily
- 2026-06-12 ·
narrative · FrontierMath Tier 1-4 approaching saturation · → daily
- 2026-06-12 ·
narrative · future of math benchmarking lies in open problems from real · → daily
- 2026-06-12 ·
fact · Claude Fable 5 FrontierMath Tiers 1–3 得分 · → daily
- 2026-06-12 ·
fact · Claude Fable 5 FrontierMath Tier 4 得分 · → daily
- 2026-06-12 ·
narrative · Anthropic 数学能力快速提升的趋势 · → daily
← 实体目录 · 系统日志