Monday, August 10, 2026
DarkSubscribe
AI Infrastructure · News & Analysis
HomeChips & HardwareReport
Chips & Hardware · Report

Google develops Frozen v2 custom chip that embeds Gemini architecture directly into silicon, targeting 6-10x efficiency gains over standard GPUs.

First-party silicon optimization could materially reduce Google's per-inference costs and pressure Nvidia's role in scaling LLM workloads.
Trade pressSlicast · July 21, 2026 · US · Source: Google News
importance 80

Google is building a new server chip internally called "Frozen v2" that embeds the Gemini AI model's architecture directly into silicon. According to sources cited by The Information, the chip could be 6 to 10 times more efficient at serving AI responses than Google's current TPU chips. Google plans to deploy it starting in 2028 and views Frozen v2 as a test run for specialized chips, with a smaller production volume than its TPU line.

Unlike Google's TPUs, which work with many models, Frozen v2 has parts of Gemini's model structure built directly into the hardware. The name follows the same logic as "freezing" parameters in AI models, where you lock values so they stop changing. With Frozen v2, a portion of the model gets permanently frozen into the chip itself, which cuts down on compute steps and speeds up responses.

The original idea reportedly came from Jeff Dean, Google DeepMind's chief scientist. His first Frozen design called for embedding the model weights directly into the chip—the specific settings that determine how an AI model responds to queries. Google abandoned that approach because the chip would have only worked with a single Gemini version and would have become outdated too quickly.

Frozen v2 takes a more flexible path by embedding the model architecture instead of weights, meaning the underlying blueprint rather than the tuned parameters. New weights can still be loaded onto the chip, though how much of the architecture will actually be hardcoded remains undecided, according to The Information.

Because the chip only works as long as Google maintains the same model architecture, it will likely never become a product for outside customers. Google already leases its TPUs to Meta, offers them to external cloud customers, and positions them through its "TPU@Premises" program as an alternative to Nvidia with an internal goal of capturing ten percent of Nvidia's annual revenue. Frozen v2, by contrast, is designed to ease Google's internal crunch on AI compute capacity.

If the chip delivers on its promise, it could still become a major competitive edge. In the AI business, how well companies optimize inference costs increasingly determines their margins. Google could use Frozen v2 to run powerful models at lower prices and capture market share from OpenAI and Anthropic.

Read the original
Google develops Frozen v2 custom chip that… · Slicast