以下部分引用 AI 总结,现阶段 AI 仍然有幻觉。请以内容的原文为准。
数据新鲜度群聊 07-14 ✓卖方 07-14 ✓Wrap 断档15天 ✕News 07-15 ✓X 断档12天 ✕页面生成 07-15
Gemini Flash 3.5
21 atoms · 跨 3 天 · 首见 2026-05-19 · 最近 2026-05-26
三色: 🟦 fact 9 · 🟥 take 12 · stance ▲4/▼12/◆2
来源: X 21
时态: fresh:20 · stale:1
标签: 好观点:12 · 好数字:5 · 好信源:4
叙事 (narrative) (12)
- 2026-05-19 · 纳入 CursorBench 评测
Gemini Flash 3.5 is now on CursorBench, our main coding agent eval.
tags: 好信源 · → daily
- 2026-05-20 · 是历史上最差的主要实验室模型更新
This might be the worst major lab model drop of all time. Llama 4 tier. Insane.
tags: 好观点 · → daily
- 2026-05-20 · is a useless model
3.5 is a useless model that should not be used for, well, anything as far as I can tell
tags: 好观点 · → daily
- 2026-05-20 · In a CursorBench evaluation against Composer and other models
I think this mostly just shows how good the Composor models are in Cursor's own benchmarks. It's still beating GPT 5.5 low and Kimi K2.6 which lots of people love as their favorite models..
tags: 好观点 · → daily
- 2026-05-20 · 编码表现不如 GPT 5.5 medium,价格相当
5.5 medium is way better and roughly same price
tags: 好观点 · → daily
- 2026-05-20 · 编码表现不如 GPT 5.5 low,但价格更贵
5.5 low is 'roughly as good', follows instructions WAY better, writes broken code less often, and is roughly half as expensive.
tags: 好观点 · → daily
- 2026-05-20 · 编码时经常写出有问题的代码
I have tried it in many tools and it aggressively writes broken code. It performed terribly in every bench I ran it on and every task I threw at it. It is not a good model.
tags: 好观点 · → daily
- 2026-05-20 · 是一个优秀的模型
It's a fantastic model. No idea where this is coming from
tags: 好观点 · → daily
- 2026-05-20 · 定价现状
Not sure why is Gemini flash 3.5 even considered a flash model based on its pricing
tags: 好观点 · → daily
- 2026-05-20 · 与用户付费意愿比较
Seems pretty far away from what users would pay for production workflow use cases
tags: 好观点 · → daily
- 2026-05-20 · 代码能力比较
it seems far behind compared to gpt 5.5/opus 4.7 even with all that benchmaxing
tags: 好观点 · → daily
- 2026-05-20 · 智能水平评价
it’s not a pro model (that one is still cooking), but it is a very smart flash model
tags: 好观点 · → daily
立场 (position) (1)
展开 1 条老旧 / 已过期
- 2026-05-20
[老旧] · 成本不应超过 frontier 模型
I expect it to not cost more than frontier models 🙃🙃
tags: 好观点 · → daily
事实 (fact) (8)
- 2026-05-20 · cursorBench 评测得分
Gemini Flash 3.5 is now on CursorBench, our main coding agent eval.
tags: 好数字 · → daily
- 2026-05-20 · Jeff Dean 称其是最强编码模型
Jeff Dean himself says this is their “strongest model for coding”
tags: 好信源 · → daily
- 2026-05-26 · 产品发布
Gemini Flash 3.5
tags: 好信源 · → daily
value: date=2026-05-26
- 2026-05-20 · Scored worse than Composer 2
Oh my god it scored worse than Composer 2! Not even 2.5! And it cost 4x to run!!!
tags: 好数字 · → daily
- 2026-05-20 · 推理成本是 Composer 2 的 4 倍
And it cost 4x more to run!!!
tags: 好数字 · → daily
- 2026-05-20 · is beating GPT 5.5 low and Kimi K2.6
It's still beating GPT 5.5 low and Kimi K2.6 which lots of people love as their favorite models
tags: 好数字 · → daily
- 2026-05-20 · 早期测试用户被告知该模型用于编码和代理工作
I was an early access tester and they explicitly said it was for coding and to use it for agentic work.
tags: 好信源 · → daily
- 2026-05-20 · 在 Skatebench 2 的视觉得分
Flash 3.5 gets a 78%
tags: 好数字 · → daily
value: qty=78%
时间轴 (近 20)
- 2026-05-26 ·
fact · 产品发布 · → daily
- 2026-05-20 ·
fact · cursorBench 评测得分 · → daily
- 2026-05-20 ·
fact · Scored worse than Composer 2 · → daily
- 2026-05-20 ·
fact · 推理成本是 Composer 2 的 4 倍 · → daily
- 2026-05-20 ·
narrative · 是历史上最差的主要实验室模型更新 · → daily
- 2026-05-20 ·
narrative · is a useless model · → daily
- 2026-05-20 ·
position · 成本不应超过 frontier 模型 · → daily
- 2026-05-20 ·
narrative · In a CursorBench evaluation against Composer and other model · → daily
- 2026-05-20 ·
fact · is beating GPT 5.5 low and Kimi K2.6 · → daily
- 2026-05-20 ·
fact · Jeff Dean 称其是最强编码模型 · → daily
- 2026-05-20 ·
fact · 早期测试用户被告知该模型用于编码和代理工作 · → daily
- 2026-05-20 ·
narrative · 编码表现不如 GPT 5.5 medium,价格相当 · → daily
- 2026-05-20 ·
narrative · 编码表现不如 GPT 5.5 low,但价格更贵 · → daily
- 2026-05-20 ·
narrative · 编码时经常写出有问题的代码 · → daily
- 2026-05-20 ·
narrative · 是一个优秀的模型 · → daily
- 2026-05-20 ·
fact · 在 Skatebench 2 的视觉得分 · → daily
- 2026-05-20 ·
narrative · 定价现状 · → daily
- 2026-05-20 ·
narrative · 与用户付费意愿比较 · → daily
- 2026-05-20 ·
narrative · 代码能力比较 · → daily
- 2026-05-20 ·
narrative · 智能水平评价 · → daily
← 实体目录 · 系统日志