hotAI

2 min read

Google ships Gemini 3.6 Flash, delays 3.5 Pro again

Google unveiled Gemini 3.6 Flash, 3.5 Flash-Lite, and a cyber model on July 21, while saying Gemini 3.5 Pro is still only in partner testing.

Image: Mashable

Google has launched two new Gemini Flash models and a specialized cybersecurity model, but Gemini 3.5 Pro still is not here. In a July 21 blog post, the company introduced Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber, while saying the delayed Pro model is “currently testing with partners” and will reach general availability “as soon as it’s ready.” Google also said it has already started training runs for Gemini 4.

That update stands out because at Google I/O in May, Alphabet CEO Sundar Pichai said Gemini 3.5 Pro would launch in June. Since then, rivals have moved ahead: Anthropic released Fable 5, OpenAI launched GPT-5.6 Sol, and Chinese lab Moonshot introduced Kimi K3, which Mashable described as a cheaper but competitive open-source model.

Gemini 3.6 Flash pricing and performance

Google is positioning Gemini 3.6 Flash as its general-purpose “workhorse” model. According to the company, it improves on 3.5 Flash in coding, knowledge work, and multimodal tasks.

Google also cut pricing versus 3.5 Flash:

  • $1.50 per million input tokens
  • $7.50 per million output tokens
Chart showing benchmark performance of the Google AI model Gemini 3.6 Flash
Chart showing benchmark performance of the Google AI model Gemini 3.6 Flash

The company published benchmark results showing gains across multiple domains. Google also said 3.6 Flash adds expanded protections against jailbreak attempts tied to chemical, biological, radiological, and nuclear misuse, as well as cyber offenses, while trying to avoid blocking legitimate requests.

Recommended reading

Why AI chatbots feel like fictional companions

Chart showing benchmark performance and token cost of the Google AI model Gemini 3.6 Flash
Chart showing benchmark performance and token cost of the Google AI model Gemini 3.6 Flash

Gemini 3.5 Flash-Lite is the cheaper, faster option. Google said it is the fastest model in the 3.5 series, capable of 350 output tokens per second, citing Artificial Analysis figures. Pricing is set at $0.3 per million input tokens and $2.5 per million output tokens. Google says it is aimed at high-throughput, latency-sensitive tasks such as agentic search and document processing, and that it beats the previous Flash-Lite generation on agentic benchmarks.

The most tightly controlled release is Gemini 3.5 Flash Cyber, built to find and patch software vulnerabilities inside CodeMender, Google’s security agent. Google said it will not be broadly available, and instead will be offered only to governments and select partners in a limited-access pilot because of the dual-use risks of a model that can identify security flaws.

Ava Chen

AI Editor

Ava covers the rapidly evolving world of artificial intelligence, from foundational models and research labs to the real-world economics of intelligence. With a background in computational linguistics, she cuts through the hype to find out what actually works. She firmly believes that benchmarks are just marketing until reproduced in the wild.

via Mashable

// Keep reading