以下部分引用 AI 总结,现阶段 AI 仍然有幻觉。请以内容的原文为准。
数据新鲜度群聊 07-14 ✓卖方 07-14 ✓Wrap 断档15天 ✕News 07-15 ✓X 断档12天 ✕页面生成 07-15
DeepSWE
13 atoms · 跨 7 天 · 首见 2026-05-26 · 最近 2026-06-23
三色: 🟦 fact 9 · 🟥 take 4 · stance ▲1/▼3/◆3
来源: X 13
时态: fresh:12 · stale:1
标签: 好数字:5 · 好观点:4 · 好信源:3 · 好问题:1 · 好思考:1
叙事 (narrative) (4)
- 2026-06-12 · 使用了mini-swe-agent作为引擎
DeepSWE also used mini-swe-agent as the harness
tags: 好信源 · → daily
- 2026-06-03 · evaluation leaves much to be desired
the actual evaluation leaves much to be desired
tags: 好观点 · → daily
- 2026-06-01 · 是第一个合理的agentic代码基准
DeepSWE is the first agentic code bench that makes sense
tags: 好观点 · → daily
- 2026-05-27 · task horizon assessment
Wouldn’t say that it’s long horizon, tasks are relatively small/short
tags: 好观点 · → daily
立场 (position) (1)
展开 1 条老旧 / 已过期
- 2026-05-26
[老旧] · 基准列表未包含Grok
wtf is this list where is grok
tags: 好问题 · → daily
事实 (fact) (8)
- 2026-06-01 · scaling from xhigh to max 效果发现
DeepSWE found very little scaling from xhigh to max
tags: 好数字·好思考 · → daily
value: direction=minimal
- 2026-06-23 · benchmark performance vs open source frontier
on Toolathlon, DeepSWE etc it's below the open source frontier
tags: 好数字 · → daily
value: direction=below
- 2026-06-01 · 完整报告已发布
Full report: https://t.co/RglaGGablq
tags: 好信源 · → daily
- 2026-05-30 · benchmark task count
DeepSWE has 113 tasks.
tags: 好数字 · → daily
value: qty=113
- 2026-05-26 · 发布新基准
Today we’re releasing DeepSWE, a new standard for agentic coding benchmarks.
tags: 好信源 · → daily
value: date=2026-05-26
- 2026-05-26 · 排名表现
DeepSWE blows up the AI coding leaderboard
tags: 好观点 · → daily
value: direction=超越
- 2026-05-26 · false-positive rate on SWE-Bench Pro
0.3% false-positive vs SWE-Bench Pro's 8.5%
tags: 好数字 · → daily
value: qty=0.3%
- 2026-05-26 · active repos count
91 active repos vs SWE-Bench Pro Public's 11 and SWE-Bench Verified's 12
tags: 好数字 · → daily
value: qty=91
时间轴 (近 20)
- 2026-06-23 ·
fact · benchmark performance vs open source frontier · → daily
- 2026-06-12 ·
narrative · 使用了mini-swe-agent作为引擎 · → daily
- 2026-06-03 ·
narrative · evaluation leaves much to be desired · → daily
- 2026-06-01 ·
narrative · 是第一个合理的agentic代码基准 · → daily
- 2026-06-01 ·
fact · scaling from xhigh to max 效果发现 · → daily
- 2026-06-01 ·
fact · 完整报告已发布 · → daily
- 2026-05-30 ·
fact · benchmark task count · → daily
- 2026-05-27 ·
narrative · task horizon assessment · → daily
- 2026-05-26 ·
fact · 发布新基准 · → daily
- 2026-05-26 ·
position · 基准列表未包含Grok · → daily
- 2026-05-26 ·
fact · 排名表现 · → daily
- 2026-05-26 ·
fact · false-positive rate on SWE-Bench Pro · → daily
- 2026-05-26 ·
fact · active repos count · → daily
← 实体目录 · 系统日志