以下部分引用 AI 总结,现阶段 AI 仍然有幻觉。请以内容的原文为准。
数据新鲜度群聊 07-14 ✓卖方 07-14 ✓Wrap 断档15天 ✕News 07-15 ✓X 断档12天 ✕页面生成 07-15
Large-scale RL runs
5 atoms · 跨 2 天 · 首见 2026-05-30 · 最近 2026-06-04
三色: 🟦 fact 0 · 🟥 take 5 · stance ▲0/▼1/◆4
来源: X 5
时态: fresh:5
标签: 好思考:3 · 好数字:2
叙事 (narrative) (3)
- 2026-06-04 · inference side: max concurrency approximation formula
Inference side, I find using “(aggregate memory capacity - model weights) / average sequence length” a good approx of max concurrency, which affects how fast you can feed the trainer
tags: 好思考 · → daily
- 2026-05-30 · production run: trainer GPU allocation vs wall-clock time practical sweet spot ratio
then for prod runs, you have one clean knob: trainer GPU allocation vs wall-clock time. theoretical limit is 3:1 if inference is FLOP-bound, but more realistically it's 1:3 or 1:4 as a practical sweet spot.
tags: 好思考 · → daily
value: qty=1:3 or 1:4
- 2026-05-30 · actual bottleneck: flops for trainer, membw for inference
flops for trainer, membw for inference
tags: 好思考 · → daily
预测 (forecast) (2)
- 2026-05-30 · production run: 4-week run estimated steps and rollouts
what this very roughly gets you is that a 4-week prod run might look something like 100K steps, each with 100K rollouts, optimistically.
tags: 好数字 · → daily
value: qty=100K steps, each with 100K rollouts · date=4-week prod run
- 2026-05-30 · production run: overhead estimate
in practice there's prob 0.5-1 OOMs of overhead which creep in somewhere, pulling down both your batch size + step count a bit.
tags: 好数字 · → daily
value: qty=0.5-1 OOMs
时间轴 (近 20)
- 2026-06-04 ·
narrative · inference side: max concurrency approximation formula · → daily
- 2026-05-30 ·
narrative · production run: trainer GPU allocation vs wall-clock time pr · → daily
- 2026-05-30 ·
forecast · production run: 4-week run estimated steps and rollouts · → daily
- 2026-05-30 ·
forecast · production run: overhead estimate · → daily
- 2026-05-30 ·
narrative · actual bottleneck: flops for trainer, membw for inference · → daily
← 实体目录 · 系统日志