以下部分引用 AI 总结,现阶段 AI 仍然有幻觉。请以内容的原文为准。
数据新鲜度群聊 07-14 ✓卖方 07-14 ✓Wrap 断档15天 ✕News 07-15 ✓X 断档12天 ✕页面生成 07-15
dzhulgakov
X · 11 atoms · 2026-06-24 → 2026-06-27 · 信用档先验 cred4
🟦 fact 3 · 🟥 take 8 · stance ▲2/▼0/◆6
覆盖实体 (top 10)
话题分布 (top 8)
动态推测 (3) · 强化学习进展 (1) · 模型训练 (1) · 推理速度 (1) · 吞吐量提升 (1) · 时延公式 (1) · 推测平衡 (1) · GPU流水线 (1)
样例 atoms
- 2026-06-27 · DeepSeek DSpark · Cheap, but not free, if running MTP for 5 tokens it’s like running 5 extra later plus sampling 5 times. If model is 60 layers, it maybe becomes 10%. But I was making the point for
- 2026-06-27 · DeepSeek DSpark · DSpark from @deepseek_ai ingeniously integrates many speculative decoding ideas to achieve 1.5x to 5x higher throughput in a real production system
- 2026-06-27 · DeepSeek DSpark · time_per_token = (num_tokens_drafted * drafter_time + verify_time(num_tokens_drafted)) / num_tokens_accepted
- 2026-06-27 · DeepSeek DSpark · Slow to run speculator or drafting too many tokens with a low guess rate can be hurtful. The right balance is needed
- 2026-06-27 · DeepSeek DSpark · What num_draft_tokens should be? It varies: * some requests (e.g. coding) are easier to predict than others * optimal length depends on server load (batch size). Speculate more wit
← Author 索引 · 实体目录