Broadcom vs Nvidia: Custom XPU Versus Merchant GPU in the Inference Build-Out
Broadcom tripled its AI revenue to $16.7 billion on capex of just $623 million—a 1% intensity versus Nvidia's 3%—exposing the structural margin gap that now defines inference procurement at every major hyperscaler.
Broadcom's most recent earnings revealed AI-related revenue tripled to $16.7 billion while its fiscal 2025 property, plant, and equipment spend reached only $623 million—a capex intensity of 1 percent against $63.89 billion in revenue. Nvidia's comparable figure for fiscal 2026 is $6.04 billion in capex on $215.94 billion in revenue, 3 percent intensity and 86.7 percent higher than the prior year. The divergence is structural: Broadcom is a fabless XPU designer that co-engineers custom accelerators with individual hyperscalers and collects design-IP rent at high gross margin; Nvidia sells a programmable merchant platform to the widest possible buyer pool. At current trajectory, analysts project Broadcom's AI-addressable revenue reaching $230 billion by 2028, driven by a roster of hyperscaler and frontier-model customers that internal management has described as the Genius accounts—OpenAI and Anthropic among them.
Broadcom's disclosed XPU customers include Google—whose TPU lineage traces through Broadcom silicon—and Meta, alongside at least two unnamed hyperscalers confirmed in the Q3 call, and the newly co-announced Jalapeño inference chip designed with OpenAI (reported by a single outlet and not yet confirmed by either company in official filings). The XPU development cycle runs 18 to 24 months per design, then moves directly into volume production at TSMC with guaranteed offtake from a single buyer—no channel inventory risk, no price discovery problem. Nvidia's merchant model faces the opposite dynamic: its Rubin-class accelerators and the inference rack architecture rolling out through the Groq partnership reach every addressable buyer simultaneously, from SB Energy's IPO-backed data centers to U.S. government diplomatic chip allocations deployed as foreign-policy instruments, but at the cost of maintaining a full-stack upgrade cycle that demands continuous reinvestment.
The margin arithmetic is more nuanced than the headline revenue gap suggests. Broadcom's $623 million in capex covers design tooling and test infrastructure; Nvidia's $6.04 billion reflects escalating investment in CoWoS advanced packaging, liquid-cooling systems, and the infrastructure-financing arm it is now building directly. Nvidia is repositioning from a chip vendor into an AI infrastructure investor: the reported $12.9 billion Hugging Face acquisition—cited by multiple outlets but unconfirmed by either company at this writing—and the $20 billion Groq rack deployment are both moves to capture recurring inference-infrastructure economics rather than one-time silicon margin. Broadcom captures margin at the design layer without volume-delivery risk; Nvidia is betting that owning the financing and the software ecosystem insulates revenue even as custom silicon progressively erodes the merchant GPU's share at the top of the hyperscaler stack.
On the cost-performance axis, directional evidence favors custom silicon at inference time. Analysis reported this week places Google's in-house TPU at up to 50 percent better inference performance per dollar than comparable Nvidia accelerators—a figure from a single analysis that should be read as indicative rather than audited, but consistent with the engineering rationale: a chip co-designed for a specific model topology eliminates the general-purpose overhead that makes the H100 and B200 programmable across workloads but costly at unit level for a single deployment. Microsoft's Maia 300 and Amazon's Trainium follow the same logic, suggesting the merchant GPU's inference premium is under systematic pressure across every major hyperscaler simultaneously.
Manufacturing exposure differs despite both companies being fabless. Broadcom's XPU designs are optimized to a specific customer's memory topology, typically avoiding the stacked HBM requirement that makes Nvidia's flagship GPUs expensive and packaging-constrained. Nvidia's CoWoS capacity demand at TSMC is a sustained bottleneck—one reason delivery lead times have remained stretched even as the 2025 demand spike normalizes. Broadcom can schedule TSMC runs to a single buyer's roadmap lock; Nvidia competes for the same packaging slots against AMD, Intel, and the hyperscaler XPU programs it is trying to pre-empt.
Three variables determine which model captures the inference build-out. First, whether the Jalapeño-OpenAI partnership—currently single-source and unconfirmed—translates into production volume that moves Broadcom's AI-revenue mix beyond its Google anchor. Second, whether Nvidia's Rubin deployments at sovereign and enterprise buyers establish a price floor for merchant GPU inference that keeps the custom XPU ROI unclear for mid-tier customers outside the top five hyperscalers. Third, whether Taiwan's proposed AI chip export restrictions disrupt Nvidia's supply routing without affecting the on-shore TSMC capacity that Broadcom's hyperscaler customers can reserve directly. Watch Broadcom's next earnings call for a confirmed count of XPU programs in production volume; watch Nvidia's capex trajectory for evidence it is pricing CoWoS packaging as a strategic moat rather than a pass-through cost; and watch for the first independently audited cost-per-token benchmark from a Google or Meta TPU deployment, which would convert the directional 50 percent advantage into a procurement-grade data point.