News

Google secretly builds a "Frozen v2" dedicated chip: Gemini is burned into the silicon chip to increase the token power per unit power consumption by 6 to 10 times. The model is directly welded into the chip, allowing the silicon chip to save computing power and data transmission for the algorithm. Google is following this path and quietly polishing a new server chip codenamed Frozen v2.

3 min read
Google secretly builds a "Frozen v2" dedicated chip: Gemini is burned into the silicon chip, which increases the token capacity per unit power consumption by 6 to 10 times. The model is directly welded into the chip, allowing the silicon chip to save computing power and data transmission for the algorithm. Google is following this path and quietly polishing a new server chip codenamed Frozen v2. According to multiple media reports, the goal of this chip is to run the Gemini model more efficiently. Reports say that it will permanently embed part of the Gemini architecture into the chip, thereby reducing the amount of calculation and data handling required to answer user questions. Google engineers estimate that compared to its latest generation general-purpose AI chip TPU (tensor processor), the Frozen chip can provide 6 to 10 times the token processing power per unit power consumption. The debut of this chip is scheduled for 2028, but Google has made it clear that it will not replace the general-purpose TPU, but will become a more specialized and targeted branch of the custom chip product line. After the news was announced, the capital market immediately reacted: Google Class A shares (GOOGL) U.S. stocks rose nearly 3.7% in early trading, and Class C shares (GOOG) rose as high as about 3.9% during the session. In response to media inquiries, Google responded in a statement that the company's team continues to research and experiment with various innovative technologies, aiming to provide users and customers with the best performance and highest efficiency, and emphasized that although not all projects will enter mass production, this rigorous exploration is a core component of its full-stack strategy. Google also said that by co-designing hardware and software from the bottom up, the entire system is deeply integrated and highly optimized for real workloads. However, Google currently positions Frozen v2 as an experimental project and has no plans for mass production like TPU. Analysts believe that behind this project is Google’s increasingly tight computing power dilemma. The previous lack of computing power not only intensified the competition for internal resources, but also allegedly forced Google Cloud to reject some external customer business. Just last month, Google agreed to pay SpaceX nearly $1 billion a month to fill the computing power gap and fulfill its computing power commitments to enterprise customers. Frozen v2 is clearly a step towards mitigating this gap. The price of solidifying the architecture into the chip is the shrinkage of flexibility. The report pointed out that this chip will be compatible with subsequent versions only if Google's future Gemini model continues to use the same underlying architecture - once the underlying model is significantly changed, the part of the design baked into it may become invalid. What is more pressing than the chip is the series of challenges that Google's AI business is currently facing: There is news that the release of the next generation Gemini Pro has been postponed, and many senior AI researchers have flowed to competitors. when