A funny thing has been happening in AI where the companies that built their names making hardware are writing software and the companies that built software are busy designing hardware.
NVIDIA is one of the more recent and well-known examples of a chip company that now has AI models and software platforms. Google, which most people know for search, Gemini and cloud software, is working on new AI chips. Meta has built its own AI models but is also designing its own AI accelerators – and Anthropic is doing the same.
The old line between hardware and software is becoming less obvious. AI has turned them into two sides of the same product. Building the model is only one element. Running that model quickly and cheaply has become just as important.
Why Are Software Companies Suddenly Designing Chips?
Google has spent years building its own Tensor Processing Units, known as TPUs. Now, according to The Information, it is also working on a new server chip called Frozen v2 that is designed specifically for Gemini. The report says the chip could launch in 2028.
The hardware is separate from Google’s existing TPUs. According to the report, Frozen v2 could run six to 10 times more efficiently than Google’s current custom chips and answer queries more quickly. Alphabet has been developing both semiconductor hardware and AI models as it builds a fully integrated AI system.
Meta is also designing its own route into AI chips and reports say Samsung Foundry is in talks to manufacture Meta’s third generation Meta Training and Inference Accelerator chips using its 2 nanometre process. Meta wants AI data centres with a combined capacity of 5GW before 2030. The company also wants to rely less on AI accelerators from AMD and Nvidia and intends to release a new generation of AI chips every six months.
Anthropic has much the same plans as reports are saying the Claude developer is looking at Samsung’s 2 nanometre manufacturing process for its own AI accelerator chips. The company reportedly wants AI data centres with a total capacity of 1GW and expects investments of around $50 billion. Around half of that spending would go into hardware such as custom AI chips, DRAM and NAND flash memory.
What Makes AI Different From Previous Software?
Vinod Jethwani, SEO strategist and AI search visibility expert at SearchTides, believes this has been building for quite a while rather than appearing overnight.
He said, “This is being seen as a sudden change whereas in reality the economy has been going on this path for some time. Demand for AI has grown far faster than traditional cloud applications, and companies building AI based products cannot always rely on others.”
Jethwani said AI asks far more from computers than traditional search ever did. He said, “AI search has redefined the equation. In the past, search primarily involved searching previously stored data. The Generative AI process generates a new response every time a question is asked, and it involves a lot of computing power in the background. Hardware makes a difference in performance, cost, and scalability.
“This is far from being just a matter of engineering. Training large AI models is expensive, and specialised chips will help organisations achieve greater predictability of expenses, better efficiency, and freedom to plan their products without restrictions due to hardware constraints.”
More from Tech
- Can AI-Powered Equipment Fix The Smart Home Gym Market?
- Is The Era Of Screenless AI Companions Finally Here?
- What Microsoft Doesn’t Want Users To Know About Browsing The Web On Windows PCs
- Your New Favourite Fashion Accessory Might Also Be A Spy Camera
- Why Are Tech Companies Updating Their Terms Of Service So Often?
- The Best Sleep Tech And Accessories For Surviving A Heatwave
- What Will An €80 Billion Investment Mean For The EU Tech Industry?
- The Real Bottleneck In Manufacturing Isn’t Machinery, It’s Data Input
Why Are Custom AI Chips Becoming So Valuable?
Juan Mathews Rebello Santos, cybersecurity researcher, ethical hacker and founder of BNVD.org, said cost is one of the biggest reasons companies want their own silicon.
He said, “The race to design custom AI chips is driven by three forces that off the shelf hardware cant solve. First is cost. Every token a company serves through an Nvidia GPU carries Nvidias margin on top. At BNVD we track inference costs across cloud providers and the markup on third party silicon is 40 to 60 percent versus custom TPUs or Trainium chips. At Google, Amazon, Microsoft and Meta scale, that difference is billions per year.
“Second is control over the security boundary. When you design your own silicon you define the trusted execution environment from the transistor up. Nvidias GPUs have had multiple vulnerabilities in their virtual memory manager and driver stack that allowed tenant escape in multi tenant AI workloads. Google and Amazon building their own chips means they can bake hardware level isolation for customer inference data directly into the architecture rather than relying on Nvidias patch cycle.
“The strategic layer is the real story though. AI chip design lets these companies own the full technology chain from silicon to API. When Google controls the TPU, the compiler, the framework, and the model, they can optimise across every layer without asking a vendor for permission. That vertical integration is the moat. If you are a cloud provider and you depend on Nvidia for chips, Nvidia is the one setting your roadmap and your margins. Designing your own chip is the only way to escape that dependency.”
Does Hardware Alone Win The AI Race?
AMD’s new Helios AI system shows that building chips is no longer enough on its own. The rack system brings together AMD’s GPUs, CPUs, networking technology and software into one package.
Forrest Norrod, head of AMD’s data centre business, told CNBC, “We’re very focused on providing the best total cost of ownership, the lowest cost per token, all in. And our customers are telling us that we’re achieving that.”
Counterpoint Research analyst Neil Shah believes software is equally important. He said AMD’s Helios chips are “on par” with Nvidia GPUs and CPUs, but “the secret sauce is in the software and optimisation.”
AI rewards those who can control every element of the system, from the silicon inside the server to the model answering your question. The old labels of hardware company and software company no longer stand…
Vladimir Beskorovainyi, Enterprise AI Architect and CTO commented as well, saying, “Custom silicon is what happens when your inference bill exceeds the cost of a chip design team. For Big Tech the dominant AI cost is no longer training, it is serving billions of requests, and the constraint inside that cost is not raw compute but memory capacity and bandwidth.
“I run production inference myself, and the operator’s truth is simple: the GPU runs out of memory long before it runs out of arithmetic. General-purpose GPUs make you pay for arithmetic you cannot feed.
“Designing your own chip is vertical integration against your own inference bill. It lets a company tune memory, interconnect and precision to its actual workloads instead of renting a one-size-fits-all card at premium margins. Add supply security in a market where GPU allocation is a boardroom topic, and the surprising thing is not that Google and its peers design chips. It is that anyone at their scale still does not.”
