hotAI

2 min read

Google bets on cheap Gemini as flagship still slips

Google launched three lower-cost Gemini models, including a security-focused Cyber variant, while its delayed 3.5 Pro remains missing.

Image: TNW

Google’s response to a summer of being outpaced is not a larger flagship model, but cheaper Gemini releases. On Tuesday, a day before Alphabet reports earnings, the company introduced three new Gemini models at its fast, low-cost Flash tier, signaling a push for efficiency over raw model power.

The main release is Gemini 3.6 Flash. Google says it improves on its predecessor for coding and knowledge work while using about 17% fewer output tokens. It also cuts pricing to $7.50 per million output tokens, down from $9, and extends its knowledge cutoff to March 2026.

Google also launched Gemini 3.5 Flash-Lite, which it says is the fastest in the family at 350 tokens per second and is even cheaper. The strategy is straightforward: most enterprise AI workloads, Google is betting, do not need a frontier-class model.

“Companies are already blowing through their annual token budgets, and it’s only May.”

Sundar Pichai, chief executive

Pichai told Business Insider earlier this year that a mix of Flash models could save companies more than $1 billion a year.

Recommended reading

Why AI chatbots feel like fictional companions

The most targeted release is Gemini 3.5 Flash Cyber, tuned to find and patch software vulnerabilities inside Google’s CodeMender agent. Google describes it as a “cost-efficient” alternative to large security models. The implied rival is Anthropic’s Mythos, priced at $10 per million input tokens and $50 per million output tokens, according to The Verge. Google claims Flash Cyber matches frontier-level performance on a key benchmark at a fraction of that cost.

Because a vulnerability-finding model could also help attackers, Google is limiting access. Flash Cyber is available only to governments and trusted partners in a limited pilot.

What stands out just as much is what Google did not ship. Gemini 3.5 Pro, the flagship model the company had promised for June, is still in testing and was reportedly delayed after underperforming on coding. Google currently has no model in the public top ten.

That gap has become harder to ignore. In roughly a week, xAI’s Grok 4.5, three versions of OpenAI’s GPT-5.6, and Moonshot’s Kimi K3 all launched. Anthropic’s Fable 5 now leads the leaderboards, as Reuters reported.

Google’s longer-term answer is Gemini 4. The company says it has begun its “most ambitious pre-training run yet” for the model, but that is still a statement of intent rather than a shipping product. Google is also working on a custom chip to run Gemini more cheaply.

For now, the company’s case rests on cost. One caution remains: the benchmark claims in this release are Google’s own, and no independent party has verified them yet.

Ava Chen

AI Editor

Ava covers the rapidly evolving world of artificial intelligence, from foundational models and research labs to the real-world economics of intelligence. With a background in computational linguistics, she cuts through the hype to find out what actually works. She firmly believes that benchmarks are just marketing until reproduced in the wild.

via TNW

// Keep reading