以下部分引用 AI 总结,现阶段 AI 仍然有幻觉。请以内容的原文为准。
数据新鲜度群聊 07-14 ✓卖方 07-14 ✓Wrap 断档15天 ✕News 07-15 ✓X 断档12天 ✕页面生成 07-15
Microsoft MAI
26 atoms · 跨 2 天 · 首见 2026-06-02 · 最近 2026-06-03
三色: 🟦 fact 24 · 🟥 take 2 · stance ▲2/▼0/◆24
来源: X 26
时态: fresh:26
标签: 好数字:25 · 好观点:1 · 好信源:1
叙事 (narrative) (1)
- 2026-06-02 · tech report transparency
microsoft MAI tech report is a gold mine, one of the most transparent for a model at this scale.
tags: 好观点·好信源 · → daily
事实 (fact) (25)
- 2026-06-03 · training hardware
they are using GB200 for training
tags: 好数字 · → daily
- 2026-06-03 · serving chip
serving the model on their chip lead to 40% throughput increase for the same 'rack power budget'
tags: 好数字 · → daily
value: qty=40% throughput increase
- 2026-06-02 · no synthetic data or distillation
this model uses zero synthetic data or distillation from previous models.
tags: 好数字 · → daily
- 2026-06-02 · model size
it's a 1T model with 35B active, trained on 33.5T tokens (30T pre-training, 3.55T mid-training).
tags: 好数字 · → daily
value: qty=1T total, 35B active, 33.5T tokens (30T pre-training 3.55T mid-training)
- 2026-06-02 · architecture choices
interleaved dense/MoE layers, local/global sliding window attention, LatentMoE
tags: 好数字 · → daily
- 2026-06-02 · layer architecture
they alternate dense and MoE layers.
tags: 好数字 · → daily
- 2026-06-02 · scaling ladder variable
the only knob they change is model depth (number of layers), everything else is derived from it with heuristics
tags: 好数字 · → daily
- 2026-06-02 · hidden size heuristic
hidden size = L * 256/3
tags: 好数字 · → daily
- 2026-06-02 · FFN hyperparameters
FFN expansion is 2x, latentMoE hyperparameters are 2x compression -> 3x expansion
tags: 好数字 · → daily
- 2026-06-02 · ablation tokens per parameter
ablation run at 100/200 TPP which is around 'chinchilla optimal'
tags: 好数字 · → daily
value: qty=100/200 TPP
- 2026-06-02 · Efficiency Gain metric
they have this Efficiency Gain (EG) metric which basically quantifies 'to reach the loss our candidate got, how much more compute would the baseline have needed?'
tags: 好数字 · → daily
- 2026-06-02 · loss definition components
it's a NLL private set with: 50% code, 17.5% STEM, 17.5% Math, 10% General knowledge, 5% Multilingual
tags: 好数字 · → daily
- 2026-06-02 · pre-training data sources
the data comes from both common crawl and private sources, no synthetic data
tags: 好数字 · → daily
- 2026-06-02 · pre-training optimizer
AdamW with slightly different betas (usually we see 0.9, 0.95), weight decay 0.1 (0.01 attention 0.005 embedding), LR decay to 10% with cosine schedule, dropout 0.15 added before residual
tags: 好数字 · → daily
- 2026-06-02 · pre-training init
N(0,0.02) init, on pre residual proj scaled down by number of residual connections
tags: 好数字 · → daily
- 2026-06-02 · batch size
global batch of 134M
tags: 好数字 · → daily
value: qty=134M
- 2026-06-02 · training sequence length
training sequence length is 16k
tags: 好数字 · → daily
value: qty=16k
- 2026-06-02 · no cold start phase
no synthetic data so no cold start phase
tags: 好数字 · → daily
- 2026-06-02 · GRPO differences
length penalty, entropy-based outer clip, no KL term, normalization is global instead of per response
tags: 好数字 · → daily
- 2026-06-02 · final data mixture
50%+ code, 15% stem, math not specified
tags: 好数字 · → daily
- 2026-06-02 · mid-training mixture
mid-training mixture increases stem and math to 35%
tags: 好数字 · → daily
- 2026-06-02 · RL training strategy
they do some rounds of RL -> self distillation SFT to recover -> RL, they claim this might be because of train/inference mismatch
tags: 好数字 · → daily
- 2026-06-02 · final training infra numbers
good infra numbers about the final training run
tags: 好数字 · → daily
- 2026-06-02 · inference throughput per watt
40% higher throughput per Watt
tags: 好数字 · → daily
value: qty=40% higher
- 2026-06-02 · shared expert use
they rely less on shared expert (and hence remove it) and this interleaving scheme leads to better efficiency.
tags: 好数字 · → daily
时间轴 (近 20)
- 2026-06-03 ·
fact · training hardware · → daily
- 2026-06-03 ·
fact · serving chip · → daily
- 2026-06-02 ·
narrative · tech report transparency · → daily
- 2026-06-02 ·
fact · no synthetic data or distillation · → daily
- 2026-06-02 ·
fact · model size · → daily
- 2026-06-02 ·
fact · architecture choices · → daily
- 2026-06-02 ·
fact · layer architecture · → daily
- 2026-06-02 ·
fact · shared expert use · → daily
- 2026-06-02 ·
fact · scaling ladder variable · → daily
- 2026-06-02 ·
fact · hidden size heuristic · → daily
- 2026-06-02 ·
fact · FFN hyperparameters · → daily
- 2026-06-02 ·
fact · ablation tokens per parameter · → daily
- 2026-06-02 ·
fact · Efficiency Gain metric · → daily
- 2026-06-02 ·
fact · loss definition components · → daily
- 2026-06-02 ·
fact · pre-training data sources · → daily
- 2026-06-02 ·
fact · pre-training optimizer · → daily
- 2026-06-02 ·
fact · pre-training init · → daily
- 2026-06-02 ·
fact · batch size · → daily
- 2026-06-02 ·
fact · training sequence length · → daily
- 2026-06-02 ·
fact · no cold start phase · → daily
← 实体目录 · 系统日志