one year on
Google launches Gemini 2.5 Flash with controllable thinking budgets in preview
The new model lets developers toggle thinking on or off and set a thinking budget, promising a cost-efficient alternative to its bigger sibling and other leading models.
Google today released an early preview of Gemini 2.5 Flash, the company’s first model with fully controllable reasoning. Announced via the Google Developers Blog by Director of Product Management Tulsee Doshi, the model is available in the Gemini API through Google AI Studio and Vertex AI.
Doshi describes 2.5 Flash as a ‘hybrid reasoning model’ that lets developers toggle thinking on or off, and set a thinking budget of up to 24,576 tokens for the reasoning phase. The model is trained to decide how long to think per prompt, using the full budget only when needed. With thinking set to zero, it maintains speeds comparable to 2.0 Flash while still offering improved performance. Google claims 2.5 Flash sits second only to 2.5 Pro on LMArena’s Hard Prompts benchmark, at a fraction of the cost.
The Hacker News thread that accompanied the launch — 1076 points and 560 comments — captured an industry in flux. Multiple users report cancelling Anthropic subscriptions in favor of Gemini, citing better reasoning and less sycophantic behavior. But the sentiment is far from unanimous: while developers praise Gemini’s architectural insight, many still prefer Claude for hands-on coding and agentic tasks, and some complain about the Gemini web app’s reliability. The thread captures a moment where the cost-performance frontier is shifting, and no single provider has locked down developer loyalty.
Calls Gemini 2.5 Pro a step up from the free models they've used before, praising its willingness to push back rather than obsequiously agree.
Cancelled Claude subscription after side-by-side coding comparisons, says Gemini now his go-to.
Acknowledges Anthropic's Claude Code tooling is still superior for file manipulation and syntax fixing.
Switched from Claude to Gemini, saying Claude 3.7 overengineers far beyond what's asked.
Finds Gemini better for architectural discussions in Rust, but Claude stronger for hands-on coding.
Thinks Google has pulled ahead of Anthropic with 2.5 Pro, though notes it is slower than Claude 3.7.
Criticizes the Gemini web app as slow and buggy, while AI Studio fares better.
One year later — open only if you can handle spoilers
Gemini 2.5 Flash went on to become one of the most widely adopted cost-efficient reasoning models of 2025, with its thinking budget approach later emulated by OpenAI and Anthropic. Google's bet on controllable reasoning helped it capture significant share of the developer API market by year-end, though Claude Code remained the gold standard for agentic workflows well into 2026.
The Weekly Replay · free by email
This week, one year ago — every Sunday.
One email each Sunday: the week's replayed AI news, with the one-year-later annotations included. Written like it's breaking — dated like it isn't.
Free · double opt-in · unsubscribe anytime · privacy