Google is developing a new server chip that directly integrates the design of its Gemini AI model, The Information reported on July 21, citing two people familiar with internal matters.
Inside Google, the chip is called "Frozen v2" and is focused on addressing the company's AI computing power shortage.
The Information reported that the shortage of AI compute has fueled internal conflict and even led Google Cloud to turn down external customer contracts.
Google employees involved in the chip's development expect Frozen v2 to be 6 to 10 times more efficient than the latest TPU, based on the number of tokens processed per unit of power.
The name Frozen comes from the concept of permanently etching parts of a model into silicon. Google plans to deploy the Frozen chip as early as 2028.
The Frozen project is not meant to replace TPUs. Like Nvidia GPUs, TPUs are designed to work across multiple AI models, and they take time because they must determine a new processing method each time the model changes.
Frozen v2, by contrast, embeds many of the decisions needed for the Gemini model into the chip itself, reducing processing steps and the amount of data movement.
The Information reported that the Frozen chip can be built to respond to queries faster and could become a catalyst for Google to roll out new AI application services.
But the Frozen chip is designed on the assumption that Google will continue to maintain the current Gemini model architecture. The chip can continue to be used only when future Gemini versions are built on the same architecture.
Google has not yet decided how much of the model weights to reflect in the chip. Earlier, Google reviewed an initial Frozen design that would etch the model weights themselves into the chip, but it put it on hold because it would shorten the chip's lifespan. This initial design was led by Google DeepMind chief scientist Jeff Dean (제프 딘).