Tech

Why is Google building its own AI chip, and what would it mean for Gemini

TechCrunch9 h ago
A close-up of a computer processor chip on a circuit board
A close-up of a computer processor chip on a circuit boardPhoto: Jakub Pabis / Pexels

Google's parent company Alphabet is reportedly working on a new chip specifically designed to make its Gemini AI models run more efficiently, according to people familiar with the effort. While the company has not formally detailed the project, the report fits a pattern that has become increasingly central to the competitive strategy of every major AI lab: control over the underlying hardware that runs your models.

To understand why this matters, it helps to separate two different jobs a chip can do for an AI system. Training a large language model, the process of feeding it enormous amounts of data so it learns patterns, is extremely computationally intensive and has traditionally been the domain of specialized graphics processing units, or GPUs, the vast majority of which are produced by Nvidia. Running a trained model to answer a user's question, known as inference, is a separate and increasingly dominant cost as AI products scale to hundreds of millions of users.

Inference is where efficiency chips like the one reportedly in development matter most. Every time someone asks Gemini a question, the company running the model pays for the computing power, electricity and cooling required to generate that answer. At Google's scale, even small efficiency gains per query translate into enormous savings when multiplied across billions of daily interactions, which is why hardware specifically optimized for inference has become a priority across the industry.

Google is not new to this space. The company has developed its own AI chips, called Tensor Processing Units or TPUs, for close to a decade, initially for internal use and later offered to external customers through its cloud computing division. TPUs have given Google a degree of hardware independence that few of its AI competitors can match, since most other labs rely heavily on purchasing Nvidia GPUs, a market where demand has far outstripped supply and prices have risen accordingly.

A new chip reportedly optimized specifically for Gemini would extend that strategy further, tailoring hardware not just to AI workloads in general but to the specific architecture and computational patterns of Google's own models. This kind of tight integration between model design and chip design, sometimes called co-design, can unlock efficiency gains that a general-purpose chip cannot match, since the hardware is built around the exact operations the model actually performs most often.

The strategic logic extends beyond cost savings. Reducing reliance on Nvidia GPUs, which remain in tight supply amid surging global AI demand, gives Google more control over how quickly it can scale Gemini's capacity without being constrained by another company's production schedule or pricing decisions. It also strengthens Google's negotiating position in an industry where GPU access has become a genuine bottleneck for smaller AI labs and cloud customers alike.

Other major AI players have pursued similar strategies. Amazon has developed its own Trainium and Inferentia chips for its cloud division, Microsoft has its Maia chip effort, and even OpenAI has reportedly explored custom silicon partnerships, underscoring how central chip independence has become to the largest players in the industry, even as Nvidia continues to dominate the broader GPU market by a wide margin.

For everyday users, the direct effect of a more efficient chip is unlikely to be visible in any dramatic way, gains typically show up as faster response times, lower operating costs that companies may or may not pass on to customers, and the ability to run more capable models without proportionally higher energy costs. The efficiency gains also matter for Google's broader push to embed Gemini across its product lineup, from search to workplace tools to Android, where running AI features cheaply at massive scale is essential to making them viable.

Energy consumption is another dimension where efficiency gains carry weight beyond the balance sheet. Data centers running AI workloads have drawn growing scrutiny over their electricity and water use, and chips that deliver more computation per watt directly reduce the environmental footprint of running models like Gemini at the scale Google operates.

Google has not confirmed a timeline for the reported chip or detailed how it differs technically from existing TPU generations, and reports describing early-stage hardware projects at major tech companies do not always translate into shipped products on the timeline initially suggested. Still, the effort reflects a broader industry reality: as AI shifts from a race over which company has the smartest model to a race over which company can run that model most cheaply and at the largest scale, the chip underneath has become just as strategically important as the software on top of it.

This article is an AI-curated summary based on TechCrunch. The illustration is a stock photo by Jakub Pabis from Pexels.

Read next