hotAI

2 min read

Google readies Frozen v2 chip for Gemini efficiency leap

Google is reportedly building a new server chip for Gemini that could make inference 6–10 times more efficient by 2028.

Image: ITzine

Alphabet is reportedly developing a new server chip, Frozen v2, aimed at sharply reducing the cost of running Gemini and other company models. According to The Information, the accelerator could make inference 6–10 times more efficient than Google’s current systems when measured by tokens per unit of energy. The chip is expected to arrive by 2028.

This is about what happens after training, when users are actively sending prompts. That stage now consumes a huge share of resources as services scale, driving up electricity bills, cooling demands, and data center requirements. For Google, that pressure is already tangible: the company has faced capacity shortages and at one point paused onboarding for some enterprise customers on Google Cloud.

Frozen v2 is said to rely on unusually tight hardware-software integration. Parts of the Gemini architecture are reportedly being built directly into the chip to cut unnecessary computation and reduce data movement between servers. In generative AI systems, that is one of the most expensive parts of the stack. Less data shuffling means lower power use and less heat, making the project as much about economics as engineering.

Google has been pursuing custom silicon for years through its TPUs, both internally and in the cloud, and Frozen v2 appears to be the next step: a chip tuned more aggressively for a specific model. That broader shift is already visible across the market, as cloud and AI companies move away from general-purpose accelerators toward specialized designs that squeeze more work out of every watt.

There is also a financial driver. Alphabet has already told investors to expect capital expenditures of $180–190 billion in 2026. Against that backdrop, even modest gains in energy efficiency can matter not just operationally, but in the company’s financial reporting. The cheaper each Gemini request becomes, the easier it is to justify massive AI infrastructure spending.

Recommended reading

AI explains discoveries but fails to predict them

Google is also playing catch-up in a market where in-house chips are becoming standard. Amazon has Trainium and Inferentia, Microsoft is backing Maia, and Meta is investing in MTIA. Nvidia still sells the most flexible and sought-after accelerators, but that is exactly why major customers are trying to shift at least part of the load away from an expensive, supply-constrained external supplier.

“Co-designing hardware and software from the ground up is key to maximizing system performance.”

Google spokesperson, comment to TechCrunch

Investors noticed the report too: Alphabet shares rose 3% after the information surfaced. If Frozen v2 gets close to its promised efficiency gains, Google will have a stronger case that it can rein in the cost of generative AI by the time the project reaches actual hardware in 2028.

Ava Chen

AI Editor

Ava covers the rapidly evolving world of artificial intelligence, from foundational models and research labs to the real-world economics of intelligence. With a background in computational linguistics, she cuts through the hype to find out what actually works. She firmly believes that benchmarks are just marketing until reproduced in the wild.

via ITzine

// Keep reading