Cobo Agentic Wallet

Google Develops Gemini-Specific AI Chip as Custom Silicon Race Intensifies

Alphabet is reportedly developing a new server chip codenamed Frozen v2 that embeds Gemini architecture directly into silicon, targeting 6-10x efficiency gains over current TPUs with deployment expected in 2028.

Cobo Newsroom
Cobo NewsroomJul 21, 2026
Key takeaways
  • Alphabet is developing a custom AI chip codenamed Frozen v2 that permanently embeds Gemini neural network architecture into silicon circuitry
  • The chip is projected to deliver 6-10 times more tokens per unit of power compared to Google current TPU chips, with deployment targeted for 2028
  • The frozen design hardwires the model structure into hardware to reduce computation and data movement, though weights can still be updated
  • The project responds to severe AI capacity constraints at Google, which has reportedly turned away some Google Cloud customers due to shortages
  • Major AI companies including OpenAI and Anthropic are pursuing similar custom chip initiatives to reduce dependence on Nvidia
  • The development signals a broader industry shift from general-purpose AI hardware toward model-specific optimized silicon

Summary

Alphabet is reportedly developing a new server chip codenamed Frozen v2 that embeds Gemini architecture directly into silicon, targeting 6-10x efficiency gains over current TPUs with deployment expected in 2028.

From General-Purpose Chips to Model-Specific Hardware

Alphabet is developing a new type of server chip designed specifically for its Gemini large language models, according to a report by The Information. Internally codenamed Frozen v2, the chip represents a radical departure from conventional AI hardware: rather than loading models into memory at runtime, it permanently embeds Gemini neural network architecture directly into the silicon circuitry.

The news drove Alphabet shares up 1.51% on Monday, with intraday gains reaching as high as 3.7%. While Google did not directly confirm the project in statements to CNBC and TechCrunch, the company said its teams are constantly researching and experimenting with new innovations to deliver maximum performance and efficiency and emphasized that co-designing hardware and software from the ground up is central to its full-stack approach.

Current mainstream AI chips, including Nvidia GPUs and Google own Tensor Processing Units, follow a general-purpose design philosophy. They store models in memory and constantly shuttle data between processors and memory during operation. This flexibility allows the same chip to run various different models, but comes at a cost in power consumption and processing time.

Frozen v2 takes a fundamentally different approach. By freezing Gemini architecture into the hardware itself, the chip can dramatically reduce data movement and redundant calculations. Engineers can still update the model by loading new weight parameters, but the underlying network structure remains fixed in silicon. This design trades flexibility for substantial efficiency gains.

The concept bears some resemblance to application-specific integrated circuits used in cryptocurrency mining, where hardwiring specific algorithms into silicon delivers performance levels unattainable with general-purpose processors. However, Frozen v2 represents a more sophisticated implementation, maintaining some degree of updateability while capturing the efficiency benefits of specialization.

Six to Ten Times Efficiency Improvement

According to internal projections cited by The Information, Google engineers estimate that Frozen v2 could serve between six and ten times more tokens per unit of power compared to the company current TPU chips. Tokens are the fundamental units of data processed by AI models, typically corresponding to words or word fragments in large language models. Higher token processing efficiency means serving more user requests with the same energy consumption, or dramatically reducing energy costs for the same service volume.

This efficiency leap carries significant implications in today AI competitive landscape. As large language models continue to scale, the energy consumption of training and inference has become an industry focal point. Investors increasingly scrutinize AI companies capital expenditures and operating costs.

The report indicates that Frozen v2 will not replace Google existing TPU product line, which is co-designed with Broadcom. Instead, it will operate as a separate chip line dedicated to running Gemini models. This dual-track strategy maintains the flexibility of general-purpose AI chips while pursuing maximum efficiency for core business operations.

The chip is currently targeted for deployment as early as 2028, still several years away. Google is still determining which portions of the model architecture to hardwire into silicon and how to balance efficiency against flexibility. The long development timeline reflects both the complexity of chip design and the need to ensure the frozen architecture remains relevant as AI technology evolves.

Industry observers note that this efficiency improvement, if realized, could provide Google with a substantial competitive advantage in serving AI workloads at scale. The ability to deliver the same performance with a fraction of the power consumption translates directly to lower operating costs and improved profit margins on AI services.

Capacity Constraints Drive Innovation Pressure

The Information report also reveals a critical driver behind the Frozen v2 project: Google is facing severe AI capacity constraints internally. The shortage has become acute enough that Google Cloud has turned away some external customers, and has created internal tensions within the organization.

This situation highlights the infrastructure challenges confronting the AI industry. Despite Google operating one of the world largest data center networks and investing in custom AI chips for years, the inference demands of models like Gemini still exceed available capacity. This capacity bottleneck not only affects product experience but also limits the company ability to scale in the AI race.

Custom chips represent an important tool for alleviating this pressure. By dramatically improving efficiency per unit of compute, Frozen v2 could significantly expand service capacity without corresponding increases in hardware investment. This is particularly critical for consumer-facing AI products that must serve hundreds of millions of users simultaneously.

The capacity crunch at Google is not unique. OpenAI has reportedly faced similar constraints, leading to periodic service slowdowns and the need to limit access to certain features. Anthropic, despite raising billions in funding, has also grappled with scaling challenges. The common thread is that AI model capabilities have outpaced the growth of available compute infrastructure, creating a fundamental supply-demand imbalance.

Efficiency improvements also serve as an important selling point in AI product competition. As market concerns about AI spending intensify, companies that can deliver services at lower cost gain competitive advantages. This explains why the Frozen v2 news drove Alphabet stock price higher, as investors see potential for long-term cost control and improved profitability.

Beyond immediate capacity relief, the project reflects a broader strategic imperative. As AI becomes central to Google product portfolio across Search, Cloud, and consumer applications, controlling the full technology stack from silicon to software becomes a source of sustainable competitive advantage. Companies that depend entirely on third-party chip suppliers face both cost and supply chain vulnerabilities.

Tech Giants Accelerate Chip Independence

Alphabet Frozen v2 project is part of a broader industry trend: major AI companies are accelerating custom chip development to reduce dependence on third-party suppliers like Nvidia.

In June, OpenAI announced its first custom chip, an inference processor codenamed Jalapeño. Earlier this month, reports emerged that Anthropic is in discussions with Samsung about a new chipmaking partnership. Amazon and Microsoft, both major cloud service providers, are also developing their own AI chip product lines. Amazon Trainium and Inferentia chips target training and inference workloads respectively, while Microsoft has announced custom AI accelerators for its Azure platform.

This independence drive stems from multiple factors. First is supply chain security: over-reliance on a single supplier creates risk, especially amid ongoing global chip supply constraints. Second is cost control: custom chips can avoid paying third-party vendor premiums, potentially reducing operating costs significantly over time. Third is performance optimization: chips tailored to specific models and workloads can achieve efficiency levels unattainable with general-purpose alternatives.

Nvidia, while still dominant in the AI chip market, faces growing challenges to its position. The custom chip projects from major tech companies represent not just technical innovation but assertions of strategic independence. For companies like Google that operate both cloud services and consumer AI products, chip autonomy is particularly important.

The competitive dynamics are complex. Nvidia maintains advantages in software ecosystems, developer tools, and manufacturing partnerships that will be difficult to replicate. However, for workloads that companies run at massive scale internally, the economics of custom silicon become increasingly compelling. The question is not whether custom chips will replace Nvidia entirely, but rather how the market will segment between general-purpose and specialized solutions.

Analysts note that this trend could reshape the AI hardware landscape over the next five years. Companies with the resources to develop custom chips, primarily large tech firms with massive AI workloads, may pull away from those dependent on commercial offerings. This could create a two-tier system where the largest players enjoy cost and performance advantages that smaller competitors cannot match.

A Paradigm Shift in AI Infrastructure

The Frozen v2 project represents an important direction in AI infrastructure development: the evolution from general-purpose hardware toward model-specific silicon. This evolution parallels earlier computing history, such as the progression from general-purpose processors to specialized accelerators like graphics processing units.

In AI early development phase, generality was the key requirement. Researchers needed to rapidly iterate across various model architectures, making hardware flexibility more important than maximum efficiency. But as certain model architectures like Transformers became de facto standards, and model scales reached hundreds of billions of parameters, efficiency began to supersede flexibility as the primary consideration.

Baking model architectures into silicon represents a logical extension of this evolution. When a model architecture is sufficiently mature and stable, and deployment scale is sufficiently large, custom hardware becomes economically viable. This specialization can achieve efficiency levels unattainable with general-purpose chips, much as ASICs in Bitcoin mining far exceed general GPU performance for specific applications.

However, this approach carries risks. If AI model architectures undergo major transformations, specialized chips could rapidly become obsolete. This is why Google is maintaining its TPU and other general-purpose AI chip product lines. The dual-track strategy bets on current architecture continuity while preserving flexibility to respond to future changes.

The technical challenges of frozen architecture chips are substantial. Designers must predict which aspects of model architecture will remain stable and which might evolve. They must balance the degree of hardwiring against the need for some adaptability. Manufacturing complexity increases as more logic is embedded directly in silicon rather than implemented in software.

From a broader perspective, the Frozen v2 project reflects the AI industry transition from an exploration phase to an optimization phase. As technical paths become clearer, competitive focus shifts from can it be done to how can it be done more efficiently. Engineering capabilities like custom hardware, algorithm optimization, and system-level co-design become key differentiators among AI companies.

For observers tracking AI infrastructure development, this trend merits close attention. Chip design cycles typically require several years, meaning today technical choices will shape the competitive landscape for years to come. When Google deploys Frozen v2 in 2028, the AI industry may look significantly different, but the race for efficiency and specialization has already begun. The companies that successfully navigate this transition, balancing specialization with adaptability and efficiency with flexibility, will likely emerge as the infrastructure leaders of the next AI era.

Source: link

AI

About Cobo

Cobo is an institutional digital asset infrastructure provider founded in 2017. The Cobo Agentic Wallet extends Cobo's MPC custody platform to autonomous onchain agents.

Press inquiries: [email protected] · Media kit, executive bios, and additional materials available on request.
Agentic Economy by Cobo

Get this in your inbox every Friday.

The weekly newsletter from the Cobo team — unpacking the most consequential stories in crypto, AI & payments through the lens of institutional custody.