以下部分引用 AI 总结,现阶段 AI 仍然有幻觉。请以内容的原文为准。
数据新鲜度群聊 07-14 ✓卖方 07-14 ✓Wrap 断档15天 ✕News 07-15 ✓X 断档12天 ✕页面生成 07-15
GPT-5.2
9 atoms · 跨 7 天 · 首见 2026-04-16 · 最近 2026-06-25
三色: 🟦 fact 5 · 🟥 take 4 · stance ▲4/▼2/◆2
来源: X 9
时态: fresh:8 · aging:1
标签: 好数字:5 · 好思考:3 · 好观点:2
叙事 (narrative) (2)
- 2026-05-04
[老化中] · 接近真正拐点
Close to the real inflection point
tags: 好思考 · → daily
- 2026-05-28 · 早期访问
Had early access to GPT-5.2.
tags: 好观点 · → daily
value: qty= · date=
预测 (forecast) (2)
- 2026-06-20 · ARC-AGI-2 得分比较
it should be well above GPT-5.2
tags: 好观点·好思考 · → daily
value: direction=well above
- 2026-05-14 · ARC-AGI curve shape compared to Ring 2.5
I think it should look like the curve of GPT-5.2 but with worse scaling with higher tokens
tags: 好思考 · → daily
事实 (fact) (5)
- 2026-05-21 · matches or outperforms human experts in 70.9% of tasks across 44 occupations
GPT-5.2 now matches or outperforms human experts in 70.9% of tasks across 44 real occupations, law, finance, engineering, and medicine up from 38% with GPT-5.
tags: 好数字 · → daily
value: qty=70.9% · date=2026
- 2026-05-21 · performance improvement over GPT-5 of 70.9% vs 38% across occupations
GPT-5.2 now matches or outperforms human experts in 70.9% of tasks across 44 real occupations...up from 38% with GPT-5.
tags: 好数字 · → daily
value: qty=70.9% vs 38% · date=2026
- 2026-06-25 · safety failure rate under TVD framework
Result: 95.3% average safety failure rate across GPT-5.2 & Claude Sonnet 4.5—far exceeding standard jailbreak attacks.
tags: 好数字 · → daily
value: qty=95.3%
- 2026-06-20 · launch date
GPT 5.2 was launched on Dec 11, 2025
tags: 好数字 · → daily
value: date=Dec 11, 2025
- 2026-04-16 · benchmark score on OccuBench
15 frontier models tested: GPT-5.2 leads at 79.6% but no model dominates all industries. Implicit faults (missing data) prove harder than explicit errors like timeouts.
tags: 好数字 · → daily
value: qty=79.6%
时间轴 (近 20)
- 2026-06-25 ·
fact · safety failure rate under TVD framework · → daily
- 2026-06-20 ·
fact · launch date · → daily
- 2026-06-20 ·
forecast · ARC-AGI-2 得分比较 · → daily
- 2026-05-28 ·
narrative · 早期访问 · → daily
- 2026-05-21 ·
fact · matches or outperforms human experts in 70.9% of tasks acros · → daily
- 2026-05-21 ·
fact · performance improvement over GPT-5 of 70.9% vs 38% across oc · → daily
- 2026-05-14 ·
forecast · ARC-AGI curve shape compared to Ring 2.5 · → daily
- 2026-05-04 ·
narrative · 接近真正拐点 · → daily
- 2026-04-16 ·
fact · benchmark score on OccuBench · → daily
← 实体目录 · 系统日志