Google is building a Gemini-only chip that could cut AI serving costs by an order of magnitude

Alphabet is developing a server chip that does something no Google processor has done before: it embeds the architecture of Gemini directly into silicon. The chip, referred to internally as Frozen v2, is designed exclusively for Google’s flagship AI model family, trading away the general-purpose flexibility of a TPU for a steep efficiency gain.

The Information reported on July 20 that the project targets a 2028 deployment and could deliver six to ten times more tokens per unit of power than Google’s latest TPU generation, the TPU 8t for training and TPU 8i for inference launched earlier this year. Alphabet stock rose roughly 3% on the news.

The name borrows a concept from training terminology: “freezing” a model’s parameters so they stop changing. Applied to chip design, it means hardwiring parts of Gemini’s decision-making into the physical silicon rather than computing them each time. The result is less electrical work per response, which lowers the cost of serving AI at scale, one of the central tensions in an industry where every new generation of model becomes more expensive to run.

“The cost of serving AI at scale has become one of the central tensions in the industry,” CNBC noted in its own reporting on the story, citing analyst commentary that the efficiency race may matter more than the capability race over the next two years.

Google declined to confirm the project’s specifics but offered a statement: “Our teams are constantly researching and experimenting with new innovations to deliver maximum performance and efficiency for our users and customers. While not every project moves into production, this rigorous exploration is central to our full stack approach.”

Frozen v2 is part of a broader industry shift. OpenAI announced its first custom inference chip, codenamed Jalapeño, in June. Anthropic is reportedly in early discussions with Samsung on chip manufacturing. The common thread: as AI models grow more capable, dependence on off-the-shelf silicon from Nvidia and others becomes both a cost burden and a supply-chain risk. Google already builds its own TPUs, but Frozen v2 goes a step further by marrying chip design to a single model family, a bet that Gemini’s architecture is stable enough to justify the investment.

A 6–10x improvement on tokens per watt would, if realized, make Gemini substantially cheaper to operate at scale. In a world where inference costs determine which features get shipped and which remain experimental, that kind of efficiency gap can become a strategic advantage, or a vulnerability if the architecture assumptions prove wrong.

Sources: Google is working on a new AI chip designed to make Gemini more efficient (TechCrunch, July 2026); Google’s Frozen v2 chip targets Gemini efficiency (NoWrap.ai, July 2026)

Scroll to Top