Google develops Frozen v2 AI inference chip that runs Gemini models up to 10 times more efficiently than prior generation.
Google is reportedly developing a custom AI chip internally codenamed "Frozen v2" designed to run Gemini models six to ten times more power-efficiently than its latest Tensor Processing Units (TPUs). Rather than replacing the existing TPU lineup, Frozen v2 is expected to form a new class of highly specialized AI accelerators by embedding parts of Gemini's architecture directly into the hardware.
The project reflects Google's growing focus on optimizing AI infrastructure as demand for Gemini-powered services surges. Reports suggest the company faces internal compute constraints that have, at times, forced Google Cloud to decline external AI infrastructure deals. By creating hardware tailored to Gemini, Google aims to dramatically improve inference efficiency while lowering energy consumption and operating costs.
Unlike general-purpose AI accelerators designed to run a wide variety of models, Frozen v2 is engineered specifically around the Gemini family. This approach—where the processor is optimized for a particular neural network architecture rather than supporting a broad range of models—resembles model-specific AI hardware concepts. Since AI inference has become one of the largest operating expenses for companies deploying large language models at scale, embedding Gemini directly into the hardware could eliminate redundant processing steps, allowing the chip to generate more AI output while consuming significantly less power. This would improve the economics of serving billions of AI queries across products including Search, Workspace, Android, YouTube, and Google Cloud.
Google is expected to maintain a dual-chip strategy: programmable TPUs for training future Gemini models, and specialized Frozen chips for deploying mature production models at much lower cost. The reported development comes as Google faces rapidly growing demand for AI infrastructure. Shortages in computing capacity have created internal tensions and limited Google Cloud's ability to accept certain external customer workloads. Frozen v2 is intended to alleviate these bottlenecks by dramatically increasing inference efficiency without requiring proportional growth in data center infrastructure.
Improving inference efficiency could reduce Google's dependence on third-party AI hardware while maximizing its in-house silicon strategy—a different approach from Microsoft, which is testing cheaper open-weight models for Copilot instead of building model-specific silicon.
Specialized chips carry trade-offs: because Frozen v2 targets deployment around 2028, engineers must ensure the hardware remains compatible with future Gemini model architectures. India, among the largest user bases for Google Search, Android, YouTube, and Workspace, represents a significant share of Gemini inference. Cheaper tokens per watt translate directly into more affordable AI features for Indian consumers and lower cloud bills for Indian startups and IT services firms using Google Cloud. Power efficiency also matters locally, as data center capacity in India is constrained by electricity availability and cooling costs.
Custom AI silicon depends on advanced foundry capacity and export rules that are increasingly politicized, as seen in Beijing's proposed export controls on AI models and chips. Google has not announced any India-specific deployment plans for Frozen v2, and the chip remains an unconfirmed internal project.
Google's reported Frozen v2 project highlights the next phase of the AI infrastructure race, where companies move beyond general-purpose AI accelerators toward chips optimized for specific foundation models. If Frozen v2 delivers its reported six-to-tenfold improvement in tokens per watt, it could significantly reduce the cost of running Gemini while expanding Google's AI capacity. Although the project remains under development and targeted for deployment as early as 2028, it underscores Google's long-term commitment to vertically integrating AI software and hardware. As competition intensifies among Google, OpenAI, Microsoft, Meta, and Anthropic, specialized AI silicon may become an increasingly important differentiator in delivering faster, more efficient, and more cost-effective generative AI services.