以下部分引用 AI 总结,现阶段 AI 仍然有幻觉。请以内容的原文为准。
数据新鲜度群聊 07-14 ✓卖方 07-14 ✓Wrap 断档15天 ✕News 07-15 ✓X 断档12天 ✕页面生成 07-15
OpenRouter Fusion
15 atoms · 跨 2 天 · 首见 2026-06-13 · 最近 2026-06-22
三色: 🟦 fact 6 · 🟥 take 9 · stance ▲6/▼0/◆0
来源: X 15
时态: fresh:15
标签: 好数字:7 · 好观点:5 · 好信源:5 · 好思考:1
叙事 (narrative) (1)
- 2026-06-22 · 第二个自适应编排API实例
This is interesting, because it's the second instance, after OpenRouter Fusion, of an adaptive orchestration API. And I believe not the last, as this category will only grow from here.
tags: 好观点 · → daily
预测 (forecast) (1)
- 2026-06-13 · 未来工作计划
it calls for future work to benchmark both on long-horizon tasks
tags: 好思考 · → daily
事实 (fact) (13)
- 2026-06-13 · 面板模型 vs 单模型表现
Panels of models consistently outperform individual models
tags: 好数字·好观点 · → daily
- 2026-06-13 · 预算模型组合能力
Panels of budget models can surpass frontier models at a much lower cost
tags: 好数字·好观点 · → daily
- 2026-06-13 · 预算面板性能对比
the budget panel was comparable with Claude Fable 5 in performance
tags: 好数字·好观点 · → daily
- 2026-06-13 · 融合性能提升来源比例
roughly three quarters of the lift that Fusion provides comes from synthesis, and one quarter from diversity
tags: 好数字 · → daily
value: qty=~3/4 from synthesis, 1/4 from diversity
- 2026-06-13 · 合成性能提升幅度
Synthesis lift: +6.7%
tags: 好数字 · → daily
value: qty=+6.7%
- 2026-06-13 · 多样性性能提升幅度
Diversity lift: +2.1%
tags: 好数字 · → daily
value: qty=+2.1%
- 2026-06-13 · 预算面板与Claude Fable 5性能差距
It landed within 1% of Fable 5 while costing roughly half the price
tags: 好数字 · → daily
value: qty=1% · direction=slightly below
- 2026-06-13 · DRACO基准测试范围
We ran it on the DRACO deep research benchmark by Perplexity: 100 deep research tasks across 10 domains, from law and medicine to finance and product comparison
tags: 好信源 · → daily
- 2026-06-13 · 基准测试数据处理
We excluded those domains across every model with a one-line config change to the OpenRouter web search tool config, then re-ran everything. All published numbers come from the clean setup.
tags: 好信源 · → daily
- 2026-06-13 · 测试时间
We ran these benchmarks earlier in the week, before Fable was taken down
tags: 好信源 · → daily
- 2026-06-13 · 成本对比方式
Cost comparisons were performed including cache hits
tags: 好信源 · → daily
- 2026-06-13 · 基准测试局限
we have only evaluated one deep research benchmark so far, which did not include long-horizon tasks
tags: 好信源 · → daily
- 2026-06-13 · 性能可达前沿水平
Beyond-frontier performance can be achieved with frontier panels
tags: 好观点 · → daily
时间轴 (近 20)
- 2026-06-22 ·
narrative · 第二个自适应编排API实例 · → daily
- 2026-06-13 ·
fact · 面板模型 vs 单模型表现 · → daily
- 2026-06-13 ·
fact · 性能可达前沿水平 · → daily
- 2026-06-13 ·
fact · 预算模型组合能力 · → daily
- 2026-06-13 ·
fact · 融合性能提升来源比例 · → daily
- 2026-06-13 ·
fact · 合成性能提升幅度 · → daily
- 2026-06-13 ·
fact · 多样性性能提升幅度 · → daily
- 2026-06-13 ·
fact · 预算面板性能对比 · → daily
- 2026-06-13 ·
fact · 预算面板与Claude Fable 5性能差距 · → daily
- 2026-06-13 ·
fact · DRACO基准测试范围 · → daily
- 2026-06-13 ·
fact · 基准测试数据处理 · → daily
- 2026-06-13 ·
fact · 测试时间 · → daily
- 2026-06-13 ·
fact · 成本对比方式 · → daily
- 2026-06-13 ·
fact · 基准测试局限 · → daily
- 2026-06-13 ·
forecast · 未来工作计划 · → daily
← 实体目录 · 系统日志