IBM's $240M Together AI Deal Brings NVIDIA B300 Inference to Enterprise Cloud
A $240 million multi-year agreement to build NVIDIA HGX B300 inference clusters on IBM Cloud marks Together AI's most significant enterprise validation to date, arriving six weeks after an $800 million Series C that valued the neocloud at $8.3 billion.
The announcement on August 12 that IBM and Together AI have signed a $240 million multi-year agreement to deploy large-scale NVIDIA HGX B300 inference clusters on IBM Cloud is the most significant enterprise validation Together AI has received to date. Under the terms reported, Together AI will supply the underlying GPU infrastructure and inference software for regulated enterprise AI workloads — the kind of compliance-constrained deployments where IBM's existing client relationships and hybrid-cloud credentials give it access that pure neoclouds struggle to reach independently. The pairing has a clear strategic logic: IBM gains access to state-of-the-art B300-generation inference capacity without absorbing the capital cost of owning GPU clusters outright, while Together AI acquires a distribution channel into financial services, healthcare, and government sectors it cannot easily build from scratch. NVIDIA, notably, sits at both ends of the arrangement — as a co-investor in Together AI's July round and as the hardware supplier at the center of the IBM agreement — a position that reflects its broader interest in cultivating inference demand across diverse channel partners, not only through its direct hyperscaler relationships.
Together AI's path to this agreement has accelerated sharply. The company completed an $800 million Series C in early July 2026, led by Aramco Ventures with NVIDIA and Vista Equity Partners as co-investors, at an $8.3 billion post-money valuation. The round coincided with the company disclosing that annual bookings had surpassed $1.15 billion — a figure that, if converted at typical enterprise retention rates, implies a reasonably measured revenue multiple relative to the valuation. By late June, the company had reported cumulative token throughput of 400 trillion, a volume that positions Together AI as a genuine scaled inference provider rather than a developer-facing API layer. The pace of these milestones — token volume, bookings, and a major institutional round, all compressed into a few months — reflects the accelerating enterprise shift toward dedicated inference infrastructure.
The IBM deal arrives as specialized inference clouds have built a credible competitive position against the large public clouds. Through the middle months of 2026, industry observers noted that Together AI, alongside Runpod and Nebius, had captured incremental share from AWS and Azure, in part because GPU supply constraints — particularly for newer Blackwell-generation hardware — left hyperscalers unable to fulfill enterprise AI demand fast enough. Enterprises and AI developers began treating neocloud alternatives as a practical default rather than a fallback option. Together AI reinforced this positioning in early July with the launch of Provisioned Throughput, a product that standardizes capacity reservation and pricing predictability for open-source model inference — a direct move toward enterprise procurement norms where unpredictable spot pricing creates budget and planning risk.
The open-source dimension matters to the IBM deal's underlying logic. Together AI's platform is built around open-weight models rather than proprietary APIs, and its enterprise pitch centers on portability, auditability, and cost compared with closed alternatives. IBM, which maintains its own open-source AI commitments through the Granite model family and the WatsonX platform, is a natural institutional partner for that framing. Multiple market observers cited in mid-2026 coverage described enterprises increasingly seeking open-model infrastructure as cost pressures, vendor-lock-in concerns, and the availability of high-performing open models converged. The IBM agreement may be one data point in that broader shift — though whether enterprise preference for open-source inference proves durable at scale, or reflects a cyclical hedge rather than a structural commitment, remains to be demonstrated.
The risks in Together AI's current position are real and warrant stating plainly. The neocloud sector's competitive intensity is rising: AWS, Google, and Microsoft have each expanded inference-specific offerings, and hyperscalers' structural advantage in absorbing capital expenditure at scale remains intact. The $1.15 billion in annual bookings is a reported order figure, not recognized revenue, and that distinction becomes increasingly material as the company approaches a scale where investor and acquisition-market scrutiny will focus on gross margin, cohort retention, and unit economics. The supply-side tailwind that has driven neocloud share gains — GPU scarcity redirecting enterprise demand — is not permanent; if major cloud providers normalize B-series GPU availability through 2027, the competitive advantage derived from that constraint may narrow significantly. And Together AI's open-source narrative, however well-supported in the current data, depends on frontier open-weight models continuing to close the quality and safety gap against proprietary alternatives — a race whose trajectory is genuinely uncertain.
Three concrete signals will determine whether Together AI's current momentum translates into durable scale. First, the conversion of reported $1.15 billion in annual bookings into recognized revenue over the next two to three quarters will reveal whether the customer base is sticky or transactional — the difference between a platform and a price-driven commodity. Second, additional enterprise partnership announcements beyond IBM, particularly from regulated-sector buyers, would confirm that the August deal is a repeatable commercial template rather than a one-off driven by headline visibility. Third, the pace of GPU supply normalization through 2027 will expose how much of Together AI's competitive position is structural and how much is borrowed from temporary hyperscaler capacity constraints. If and when supply tightness eases, the quality of Together AI's inference software, its platform differentiation, and the depth of its open-source ecosystem will need to sustain customer retention on their own merits.