Google’s Frozen v2 chip could make Gemini dramatically cheaper to run

Home Semiconductor News Google’s Frozen v2 chip could make Gemini dramatically cheaper to run
Google Gemini Frozen v2

Frozen v2 bakes Gemini’s architecture into silicon for up to 10x efficiency

In sum – what we know:

  • Architecture in silicon – Frozen v2 bakes parts of Gemini’s architecture directly into an inference-only chip, cutting the data movement that dominates large-model serving costs.
  • A big efficiency claim – Google’s internal estimates project a 6–10x gain in tokens served per unit of power over its latest TPUs, though the chip is unfinished.
  • A cautious 2028 target – Google plans a limited internal rollout in 2028, treating the specialized chip as a test bed rather than a TPU-scale flagship.

According to a new report from The Information, Google is working on a new AI chip, internally codenamed “Frozen v2,” and if reports are accurate, it could have a dramatic impact on the cost of AI inference. The idea is to “freeze” parts of the Gemini model permanently into the hardware itself — baking chunks of the model’s architecture directly into silicon rather than loading everything flexibly at runtime. It’s an inference chip, not a training accelerator, meaning its entire job would be serving Gemini-generated responses to users and enterprise customers as cheaply and quickly as possible.

The project reportedly originated with Jeff Dean, Google’s chief scientist, who has long pushed for deep co-design between the company’s AI research and hardware teams. Frozen v2 is a more extreme expression of that philosophy — an experiment in just how specialized AI hardware can get, framed internally as a trial run rather than a committed platform.

The new chip would represent an entirely new processor line, distinct from Google’s TPU series. TPUs are general-purpose AI accelerators designed to run many different models. Frozen v2 is built to run one, but do so extremely efficiently.

A unique approach

The conventional approach is to build a flexible accelerator, then load model weights onto it and execute them through software frameworks. That flexibility comes with a cost. Generic accelerators carry execution overhead precisely because they need to work with whatever model you throw at them. Frozen v2 sidesteps that by pre-embedding key aspects of Gemini’s architecture, its data-processing patterns, and the blueprint of its layers directly into the chip.

Frozen v2, as the name suggests, isn’t the first crack at this. The original Frozen — also a Jeff Dean project — was more radical still. Rather than just hardwiring Gemini’s architecture, it would have burned the model’s actual weights directly into the chip. That would have squeezed out even more efficiency, but it carried a fatal flaw: a chip fused to one specific set of weights can only ever run one specific version of Gemini. The moment Google shipped a model update, that expensive silicon would be dead weight. The lifecycle was simply too short to justify, and Google shelved the approach. Frozen v2 is the pragmatic retreat — freeze the architecture, which evolves slowly, while keeping the weights loadable so the chip can survive across multiple Gemini generations.

How far Google takes this is still an open question, and the reporting suggests engineers haven’t settled it either. There’s a spectrum here. Hardwiring the architecture while keeping weights loadable preserves some room to update the model; baking in weights or data-processing steps buys more efficiency but locks things down further. Google is actively navigating that trade-off, and the answer will largely determine how useful the chip remains as Gemini evolves.

The team is also reportedly exploring hardware-level optimizations like operator fusion, where multiple steps that would run separately in software get merged into single hardware operations. The larger goal is to keep big portions of a Gemini model — perhaps all of it — resident on-chip, cutting down data movement between memory and compute units. That’s worth dwelling on, because data movement, not arithmetic, is the dominant bottleneck and energy cost in large-model inference. A chip that barely has to move data has a structural advantage no amount of raw compute can match.

Big on efficiency

The headline number is striking. Internal estimates reportedly forecast a 6–10x efficiency gain over Google’s latest TPUs, measured in tokens served per unit of power. That’s Google’s own projection for an unfinished chip, so treat it accordingly — but even if the real-world figure lands at the bottom of that range, or below it, a multi-fold improvement would be significant at Google’s scale.

The practical upside is twofold. Lower latency means faster Gemini responses across Search, Workspace, and the rest of Google’s product stack. And a lower marginal cost per query makes it cheaper to embed Gemini features across more products and user segments — including the free tiers where serving costs quietly eat margins.

There’s an environmental angle too. AI inference is a growing share of data center energy consumption, and a chip that serves the same queries on a fraction of the power would materially help on that front. Google doesn’t need to frame Frozen v2 as a sustainability play for it to function as one.

A few years away

None of this is imminent. Google is reportedly targeting 2028 for initial deployment in its own data centers, and even then the plan is a limited rollout for select internal workloads — not mass production. Expect Frozen v2 to ship in much smaller quantities than TPUs, at least initially, handling a subset of Gemini inference where the specialization pays off most.

There are plenty of risks too. Hardwiring a model’s architecture into silicon is a bet that the architecture stays relatively stable. If future Gemini versions shift their structure meaningfully, the chip could become suboptimal or outright obsolete — and unlike a software stack running on TPUs or GPUs, you can’t patch silicon. Debugging and upgrading a chip this specialized is also considerably harder than pushing a software update to a general accelerator. That’s the double-edged sword of specialization, and it’s presumably why Google is treating this as a test bed rather than a flagship.

The commercial positioning is unclear as well. Google hasn’t decided whether Frozen v2 will eventually be branded as a Google Cloud offering or kept strictly as internal infrastructure sitting behind Gemini APIs. Given the limited production plans, the internal route seems the likelier starting point — with any cloud marketing following only if the technology proves itself.

The compute crunch

The context behind all of this is a genuine capacity problem. Demand for Gemini-powered products has surged to the point that Google, like others, is facing an internal AI compute crunch, and it’s reportedly caused friction. Google Cloud has had to decline deals with some external customers due to limited AI computing capacity, which has fueled internal tensions over how resources get allocated. Frozen v2 is, in part, an answer to that — a way to serve far more Gemini queries from the same power and infrastructure footprint, and to support larger, more capable Gemini variants at similar or lower operating cost.

It also chips away at Google’s dependence on Nvidia. GPUs remain expensive and supply-constrained, and every workload Google can shift onto its own silicon is leverage it doesn’t have to buy. Against OpenAI and Microsoft, the efficiency gains would sharpen Google’s competitive position on both performance and price for generative AI services — assuming, of course, the projections hold.

Zoom out and Frozen v2 fits a pattern. Amazon has Inferentia and Trainium, Microsoft is developing custom accelerators alongside its AMD partnership, and Meta has its MTIA chips. Every hyperscaler is pursuing vertical integration to control cost and scale. Google is simply taking the idea further than anyone else.

What you need to know in 5 minutes

Join 37,000+ professionals receiving the AI Infrastructure Daily Newsletter

This field is for validation purposes and should be left unchanged.

This website uses cookies to improve your experience. We'll assume you're ok with this, but you can opt-out if you wish. Accept Read More