以下部分引用 AI 总结,现阶段 AI 仍然有幻觉。请以内容的原文为准。
数据新鲜度群聊 07-14 ✓卖方 07-14 ✓Wrap 断档15天 ✕News 07-15 ✓X 断档12天 ✕页面生成 07-15
PostTrainBench
6 atoms · 跨 4 天 · 首见 2026-06-11 · 最近 2026-07-02
三色: 🟦 fact 4 · 🟥 take 2 · stance ▲0/▼1/◆0
来源: X 6
时态: fresh:6
标签: 好观点:3 · 好数字:2 · 好信源:1
叙事 (narrative) (2)
- 2026-06-20 · 模型触发的常见作弊方式
_there are many other benchmark hacks: - repeated official eval probing + checkpoint/hyperparameter selection - exploiting stochastic or underspecified eval settings - editing model-side generation_config.json / tokenizer / EOS / stop-token behavior - training to exact parser/scorer quirks - synthetic data that mirrors benchmark schemas, styles, or rubrics - judge/rubric hacking for Arena and HealthBench_
tags: 好观点 · → daily
- 2026-06-20 · 系统提示词倾向
I think the biggest issue is that models are encouraged to do cheat: 'We want to train the small LLM {model} to excel at {benchmark}.' 'You should perform automated research and development to post-train {model} to achieve maximum performance on {benchmark}.'
tags: 好观点 · → daily
事实 (fact) (4)
- 2026-06-21 · 新裁判机制
we are finalizing a new judge, btw, that will penalize more behaviors!
tags: 好信源 · → daily
- 2026-07-02 · benchmark ranking
GLM 5.2 tops PostTrainBench
tags: 好数字 · → daily
value: direction=top
- 2026-06-20 · 作弊检测范围
The judge that is supposed to stop cheating on PostTrainbench mostly checks for direct contamination/model substitution
tags: 好观点 · → daily
- 2026-06-11 · Opus 4.8 max reasoning score
Opus 4.8 with max reasoning is a new best model on PostTrainBench by a very large margin: 37.2% vs. 28.6% of Opus 4.7.
tags: 好数字 · → daily
value: qty=37.2% · date=2026-06-11
时间轴 (近 20)
- 2026-07-02 ·
fact · benchmark ranking · → daily
- 2026-06-21 ·
fact · 新裁判机制 · → daily
- 2026-06-20 ·
fact · 作弊检测范围 · → daily
- 2026-06-20 ·
narrative · 模型触发的常见作弊方式 · → daily
- 2026-06-20 ·
narrative · 系统提示词倾向 · → daily
- 2026-06-11 ·
fact · Opus 4.8 max reasoning score · → daily
← 实体目录 · 系统日志