# Google Is Building a Chip That Bakes Gemini's Architecture Into Silicon

Engineers project the "Frozen v2" server part could serve six to ten times more tokens per watt than Google's newest tensor processing units, trading flexibility for efficiency.

- Published: 2026-07-21T05:32:28.125Z
- Canonical: https://polylog.news/ai/2026-07-21/google-is-building-a-chip-that-bakes-gemini-s-architecture-i
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [CNBC](https://www.cnbc.com/2026/07/20/alphabet-googl-stock-ai-chip-report.html), [TechCrunch](https://techcrunch.com/2026/07/20/google-is-working-on-a-new-ai-chip-designed-to-make-gemini-more-efficient/), [The Decoder](https://the-decoder.com/googles-frozen-v2-chip-reportedly-bakes-geminis-architecture-directly-into-silicon-for-efficiency-gains/), [Polylog editors](https://polylog.news)

Google is developing a server accelerator, internally called Frozen v2, that would permanently implement parts of the Gemini model architecture in hardware rather than running it as software on a general-purpose chip, according to [reporting from The Information relayed by CNBC](https://www.cnbc.com/2026/07/20/alphabet-googl-stock-ai-chip-report.html) and [TechCrunch](https://techcrunch.com/2026/07/20/google-is-working-on-a-new-ai-chip-designed-to-make-gemini-more-efficient/). The name describes the design, because the model's computing logic is fixed permanently into the silicon. Alphabet shares rose after the report.

The key technical distinction is that Frozen v2 hardwires the architecture, not the weights. New Gemini weights can still be loaded, as long as future models keep the same underlying structure. Google engineers estimate the part could serve between six and ten times more tokens per unit of power than the company's latest tensor processing units (TPUs), [The Decoder reported](https://the-decoder.com/googles-frozen-v2-chip-reportedly-bakes-geminis-architecture-directly-into-silicon-for-efficiency-gains/). That efficiency comes at the cost of generality, because an application-specific integrated circuit (ASIC) tuned to one architecture becomes useless the moment the architecture changes.

Several caveats limit the claim. Google reportedly views Frozen v2 as a trial run rather than a TPU-scale program, with deployment expected around 2028. The projections are internal and unverified by independent benchmarks, and no silicon has been shown. What matters is the direction. A large cloud company is judging that model architectures are now stable enough to justify fixing them into hardware for the inference workloads that dominate its serving costs.

## What this means

The exposed party is the general-purpose accelerator model that underpins Nvidia's margins. If serving a fixed architecture on a bespoke ASIC really delivers six to ten times the tokens per watt, the marginal cost of running frontier inference falls sharply for whoever controls both the model and the chip, and Google controls both. The gain depends on Gemini's architecture staying stable through 2028, so Google is assuming that model design has become stable enough to treat as a fixed manufacturing specification rather than an active research question.

## What to watch

- Whether Google publishes or leaks independent tokens-per-watt figures for Frozen v2 against a named TPU baseline, which would separate a real efficiency step from an internal projection.
- Whether other labs and clouds move toward model-specific silicon, signaling that architectures are converging enough to hardwire rather than still changing every cycle.
