3 min read

Google may hardwire Gemini into a new AI chip

Google is reportedly developing a Gemini-specific chip that could be 6 to 10 times more efficient than its current AI silicon.

Image: TNW

Google is reportedly working on a new kind of AI chip that would hardwire Gemini’s neural-network architecture into the silicon itself, rather than loading the model onto a general-purpose accelerator. The project, informally known as “Frozen v2,” was first reported by The Information and later picked up by Reuters and Bloomberg Law. Alphabet shares rose as much as 3.7% after the reports.

Google has not confirmed the project, and the chip is reportedly still years away, with deployment targeted as early as 2028. Still, the idea points to a more aggressive approach to AI infrastructure.

Today’s AI chips typically keep a model in memory and move data back and forth as it runs. Frozen v2 would instead bake Gemini’s architecture directly into the circuitry. Engineers could still update the model’s weights, but the core structure would remain fixed — effectively “frozen.” According to the report, Google is still deciding how much of the model to hardwire.

Recommended reading

AI’s next bottleneck may be materials, not chips

The goal is efficiency. The Information says the chip could be 6 to 10 times more efficient than Google’s latest custom AI chips when measured by tokens served per unit of power. It would reportedly be a separate silicon line rather than a replacement for TPUs.

Why Google is pursuing a fixed-model chip

The timing appears tied to an internal AI capacity crunch. The Information reports that the pressure has been strong enough for Google Cloud to turn away some outside customers, while also creating internal tensions. At data-center scale, lower power use directly affects cost, and a chip tuned for one model can strip out much of the overhead that comes with a more flexible design.

A fixed architecture could also reduce latency, making it better suited to real-time services such as voice assistants. Strategically, it would push Google further away from dependence on Nvidia, extending a long-running effort that already includes Google’s in-house TPUs and a broader supplier base.

Google is not alone in exploring this direction. Startup Taalas is already pitching a similar concept with a chip called Hardcore, which it says prints a model’s weights and architecture directly onto the hardware. Taalas claims its chip can serve up to 17,000 tokens a second, compared with roughly 150 per user on a top Nvidia GPU, and says it does not require costly high-bandwidth memory.

The tradeoff is obvious: less flexibility. AI models evolve quickly, and a chip built around today’s Gemini architecture could look outdated by 2028. Google’s reported design would soften that by allowing weight updates, but the architecture itself would still be locked in.

A Google spokesperson did not confirm the project, saying only that the company’s teams experiment with high-efficiency ideas and that not every lab effort reaches production. For now, Frozen v2 looks less like a product announcement than a sign of where custom AI silicon may be heading next.

Marcus Vance

Enterprise Editor

Marcus follows the money. He covers enterprise software, cloud architecture, and the tectonic shifts in Big Tech strategy. He translates dense earnings calls and complex M&A activity into actionable insights about where the industry is actually heading. If a tech giant makes a silent pivot, Marcus is usually the first to notice.

via TNW

// Keep reading