以下部分引用 AI 总结,现阶段 AI 仍然有幻觉。请以内容的原文为准。
数据新鲜度群聊 07-14 ✓卖方 07-14 ✓Wrap 断档15天 ✕News 07-15 ✓X 断档12天 ✕页面生成 07-15
Gemini models
4 atoms · 跨 4 天 · 首见 2026-05-05 · 最近 2026-05-29
三色: 🟦 fact 0 · 🟥 take 4 · stance ▲0/▼2/◆2
来源: X 4
时态: fresh:4
标签: 好观点:3 · 好思考:1
叙事 (narrative) (3)
- 2026-05-29 · LisanBench 表现问题
Gemini models perform a bit worse than their Anthropic and OpenAI counterparts, but they have by far the longest outputs - they are a bit delusional and keep yapping; they don't realize and stop when they made a mistake
tags: 好观点 · → daily
- 2026-05-21 · benchmark 表现问题陈述
Benchmarks were never the issue for Gemini models.
tags: 好观点 · → daily
- 2026-05-20 · Previous models were bad for coding
the gemini models were so bad for coding.
tags: 好观点 · → daily
事实 (fact) (1)
- 2026-05-05 · knowledge evaluation methodology
Geminis were used for calibrating this test
tags: 好思考 · → daily
value: qualifier=used for calibrating this test
时间轴 (近 20)
- 2026-05-29 ·
narrative · LisanBench 表现问题 · → daily
- 2026-05-21 ·
narrative · benchmark 表现问题陈述 · → daily
- 2026-05-20 ·
narrative · Previous models were bad for coding · → daily
- 2026-05-05 ·
fact · knowledge evaluation methodology · → daily
← 实体目录 · 系统日志