Technical analysis comparing GPU versus ASIC economics for inference workloads and their diverging cost trajectories.
Google and Broadcom's TPU is rapidly closing the cost performance gap with Nvidia's GPU solutions for inference applications, according to an analysis by Tight Spreads examining the competitive dynamics between merchant solutions and custom ASIC offerings. The analysis quantifies these developments through an "inference cost curve" comparing solutions from different chip vendors and their trajectories. Our findings indicate a ~70% reduction in cost per token from TPU v6 to TPU v7, bringing absolute costs on par or slightly better than Nvidia's GB200 NVL72, while Trainium and AMD solutions have delivered only ~30% cost reduction and continue to lag on absolute cost metrics.
Nvidia maintains significant competitive advantages beyond raw compute performance. The company's "CUDA moat" and superior time to market positioning remain critical differentiators for enterprise customers, despite the narrowing gap in cost-per-token performance. This positioning is reinforced by Nvidia's substantial R&D spending, its strong networking capabilities through Mellanox, and recent innovations in memory technology with its context memory storage controller offering. These developments suggest Nvidia is well-positioned to maintain leadership as the accelerator market evolves, particularly given its rapid pace of innovation in the near term.
AMD and Amazon's Trainium currently lag behind Nvidia and Google on both generation-to-generation cost reductions and absolute inference costs. However, the competitive landscape is expected to shift in late 2026 when AMD's rack-level solutions become available. AMD claims its rack-level solution will be on par with Nvidia's VR200 for training and inference. Additionally, checks since AWS re:Invent suggest that Trainium 3 and Trainium 4 may offer stronger performance that rectifies some of Trainium 2's challenges, though these solutions warrant continued monitoring.
As raw compute performance approaches physical limits with accelerators already at reticle limits, further cost reductions will be driven by advancements in networking, memory, and packaging technologies rather than incremental compute improvements. Broadcom continues to lead in processor, accelerator, and ethernet networking solutions alongside its industry-leading SERDES capabilities, positioning the company well to benefit from these technological shifts. The analysis notes that beyond internal workload optimization, hyperscalers are expected to selectively adapt their internal offerings for use by external customers, as evidenced by customers such as Anthropic placing orders worth $21bn with Broadcom, with shipments expected in mid-2026.
The investment thesis favors Broadcom and Nvidia as the most sustainable beneficiaries of AI infrastructure capital expenditures, tied to the most critical technological advancement vectors. Sentiment toward AMD could shift if on-time OpenAI deployments drive fundamentals in the second half of 2026 and new hyperscaler customer wins materialize. The analysis explicitly acknowledges that its methodology, focused on accelerator compute performance, does not fully contemplate performance improvements from new networking, memory, or packaging innovations, nor software optimization uplifts that could influence the competitive positioning over time.