以下部分引用 AI 总结,现阶段 AI 仍然有幻觉。请以内容的原文为准。
数据新鲜度群聊 07-14 ✓卖方 07-14 ✓Wrap 断档15天 ✕News 07-15 ✓X 断档12天 ✕页面生成 07-15
MiMo
23 atoms · 跨 7 天 · 首见 2026-04-27 · 最近 2026-06-21
三色: 🟦 fact 15 · 🟥 take 8 · stance ▲2/▼2/◆5
来源: X 23
时态: fresh:17 · aging:6
标签: 好数字:14 · 好观点:7 · 好思考:1 · 好信源:1
叙事 (narrative) (7)
- 2026-04-28
[老化中] · 设计目标
this was the explicit goal with MiMo
tags: 好思考 · → daily
value: qty=
- 2026-04-27
[老化中] · V2.5-Pro 思考 12 分钟
Thought for 12 minutes, long enough to run out of output tokens lmao
tags: 好数字 · → daily
- 2026-05-27 · 对公司盲目降价的看法
We previously advised LLM companies not to "blindly cut prices" precisely because very few model architectures and inference optimizations can keep API costs from running at a loss.
tags: 好观点 · → daily
- 2026-04-27
[老化中] · pretraining data pipeline quality
I actually like the mimo pretraining data pipeline most out of others who have mentioned it. Super clean and scalable
tags: 好观点 · → daily
- 2026-04-27
[老化中] · V2.5 有 CoT 模式
MiMo V2.5 sure has some… entertaining CoT motifs
tags: 好观点 · → daily
- 2026-04-27
[老化中] · V2.5 的 RL 来自自家
but I can tell it's their own RL
tags: 好观点 · → daily
- 2026-04-27
[老化中] · Pro 有 Gemini 风格的循环
Pro has gemini-style "I must stop procrastinating" loops but gets out of them
tags: 好观点 · → daily
事实 (fact) (16)
- 2026-05-27 · API Input (Cache Hit) 最大价格降幅
The deepest price cut, up to 99%, is for Input (Cache Hit).
tags: 好数字 · → daily
value: qty=99% · direction=down
- 2026-05-27 · 推理框架支持层级 KV cache 优化
our inference framework now supports hierarchical KV cache optimization for SWA
tags: 好信源 · → daily
- 2026-05-27 · 生产推理引擎测试缓存 token 容量提升
Production inference engine tests show this optimization increases cached token capacity by 5x, equivalent to an 80% reduction in caching costs.
tags: 好数字 · → daily
value: qty=5x · direction=up
- 2026-05-27 · API Input (Cache Miss) 和 Output 价格降低幅度
Prices for Input (Cache Miss) and Output are also reduced by 60%-80%.
tags: 好数字 · → daily
value: qty=60%-80% · direction=down
- 2026-05-27 · 模型架构 Full:SWA 稀疏比
this mainly benefits from the extreme 1:7 Full:SWA sparsity ratio brought by the model architecture (the prefill compute of the 70-layer MiMo-V2.5-Pro roughly equals a 10-layer GQA model).
tags: 好数字 · → daily
value: qty=1:7
- 2026-05-27 · 此前 API 定价利润率
This kept our original inference costs well below the industry average, naturally leaving a 2x-3x profit margin in pricing.
tags: 好数字 · → daily
value: qty=2x-3x
- 2026-05-27 · 降价后推理引擎产能利用率
Operating at these newly reduced API prices, our production inference engine is running at near full capacity
tags: 好数字 · → daily
value: qty=near full capacity · direction=up
- 2026-05-27 · 降价后运营盈亏判断
and we can still essentially break even.
tags: 好数字 · → daily
value: qty=break even
- 2026-04-29 · V2.5-Pro 输入缓存未命中价格
MiMo-V2.5-Pro (256K-1M)输入缓存未命中价格: ¥14/1M tokens
tags: 好数字 · → daily
value: qty=¥14/1M tokens · date=2026-04-29
- 2026-04-29 · V2.5-Pro 输出价格
MiMo-V2.5-Pro (256K-1M)输出价格: ¥42/1M tokens
tags: 好数字 · → daily
value: qty=¥42/1M tokens · date=2026-04-29
- 2026-06-21 · pretraining tokens
MiMo did 47T
tags: 好数字 · → daily
value: qty=47T
- 2026-05-29 · token price cut to match DeepSeek
permanently cut API prices to match DeepSeek
tags: 好数字 · → daily
value: qty=permanently · date=May
- 2026-05-27 · monthly token usage with $16 sub
Got the $16 sub and its on track for around 0.4B-0.5B tokens
tags: 好数字 · → daily
value: qty=0.4B-0.5B tokens
- 2026-05-14 · first LLM release date
their first (7B dense) llm was released exactly a year ago
tags: 好数字 · → daily
value: date=1 year ago · direction=before 2026-05-14
- 2026-05-27 · caching performance vs DeepSeek
MiMo caching is no where near as good DS so its much more expensive even with the sub
tags: 好观点 · → daily
- 2026-04-27 · pretraining data pipeline missing augmented synthetic data
Although it didn’t have too many details and was lacking in augmented synthetic data like Kimi does very well (defo added it in the newer models)
tags: 好观点 · → daily
时间轴 (近 20)
- 2026-06-21 ·
fact · pretraining tokens · → daily
- 2026-05-29 ·
fact · token price cut to match DeepSeek · → daily
- 2026-05-27 ·
fact · caching performance vs DeepSeek · → daily
- 2026-05-27 ·
fact · monthly token usage with $16 sub · → daily
- 2026-05-27 ·
fact · API Input (Cache Hit) 最大价格降幅 · → daily
- 2026-05-27 ·
fact · 推理框架支持层级 KV cache 优化 · → daily
- 2026-05-27 ·
fact · 生产推理引擎测试缓存 token 容量提升 · → daily
- 2026-05-27 ·
fact · API Input (Cache Miss) 和 Output 价格降低幅度 · → daily
- 2026-05-27 ·
fact · 模型架构 Full:SWA 稀疏比 · → daily
- 2026-05-27 ·
fact · 此前 API 定价利润率 · → daily
- 2026-05-27 ·
fact · 降价后推理引擎产能利用率 · → daily
- 2026-05-27 ·
fact · 降价后运营盈亏判断 · → daily
- 2026-05-27 ·
narrative · 对公司盲目降价的看法 · → daily
- 2026-05-14 ·
fact · first LLM release date · → daily
- 2026-04-29 ·
fact · V2.5-Pro 输入缓存未命中价格 · → daily
- 2026-04-29 ·
fact · V2.5-Pro 输出价格 · → daily
- 2026-04-28 ·
narrative · 设计目标 · → daily
- 2026-04-27 ·
narrative · pretraining data pipeline quality · → daily
- 2026-04-27 ·
fact · pretraining data pipeline missing augmented synthetic data · → daily
- 2026-04-27 ·
narrative · V2.5 有 CoT 模式 · → daily
← 实体目录 · 系统日志