IBM's $240 Million Bet on Together AI Marks a Structural Shift in Enterprise Inference
A multi-year IBM Cloud deal to deploy NVIDIA B300 clusters cements Together AI's pivot from open-source inference platform to enterprise infrastructure partner — and tests whether a neocloud with an $8.3 billion valuation can deliver at that scale.
When IBM and Together AI announced a $240 million multi-year agreement to deploy NVIDIA HGX B300 inference clusters on IBM Cloud — with both parties targeting Q1 2027 for completion — the deal gave concrete form to a trend that has been gathering momentum for more than a year: enterprise AI infrastructure is consolidating around specialist operators, and Together AI is emerging as one of the most legible examples of that realignment. The agreement, confirmed by multiple outlets including The Register and Analytics India Mag beginning August 12, is structured around regulated enterprise workloads — the compliance-heavy, mission-critical deployments that hyperscalers have historically dominated.
The terms reward examination. IBM gains a credible open-source inference partner and a differentiated GPU product line at a moment when its cloud franchise needs narrative momentum; Together AI gains IBM's enterprise distribution network and a large, committed revenue anchor. According to Fierce Network, Together AI expects the IBM-backed capacity to sell out before the service goes live — a claim that, if accurate, speaks to the depth of current demand, though it also reflects the company's incentive to project confidence ahead of a major launch.
To understand why IBM is writing a nine-figure check to a neocloud founded less than five years ago, it helps to trace Together AI's trajectory. The company built its early reputation as an inference-optimized platform for open-source large language models — Llama, Mistral, and their derivatives — at pricing designed to undercut hyperscaler on-demand rates. By late June 2026, Crypto Briefing reported that the platform's token volume had reached 400 trillion, reflecting accelerating demand for cost-effective, vendor-unconstrained model access. Together AI, RunPod, and Nebius each benefited as GPU availability tightened at AWS and Azure, with AI startups routing workloads to faster, less locked-in alternatives, as documented by multiple outlets through mid-2026.
The capital structure has tracked that demand curve upward. In early July 2026, Together AI closed an $800 million Series C at an $8.3 billion valuation, with Aramco Ventures leading and NVIDIA and Vista Equity Partners participating — confirmed by the company's own Business Wire release. The NVIDIA co-investment carries structural significance: it indicates the chip designer views specialized inference platforms built around its hardware as a legitimate channel rather than a competitive threat. At the time of the round's close, the company reported annual committed orders exceeding $1.15 billion. The following month, Together AI formalized a product designed to capture recurring revenue from that demand: Provisioned Throughput, which standardizes reserved-capacity pricing for open-model inference in a manner that mirrors the reserved-instance economics long familiar from traditional cloud providers.
The geographic footprint is expanding at a comparable pace. On August 14, Larsen & Toubro confirmed through its own press release a "mega" order to build India's largest NVIDIA B300 AI factory for Together AI in Chennai, configured with at least 10,000 GPUs; NDTV placed the contract value at approximately Rs 10,000 crore, or roughly $1.2 billion. That deployment, combined with the IBM Cloud agreement, sketches the outline of a multi-region inference network in which Together AI operates as tenant rather than reseller — owning the inference layer across cloud and co-location environments rather than arbitraging excess capacity. It is that positioning distinction which separates a durable platform thesis from a temporary supply play.
Several risks warrant equal weight. Together AI remains wholly dependent on NVIDIA supply at a time when B300 allocations are constrained across the industry; hardware slippage would push the Q1 2027 IBM target and strain a relationship partly premised on speed-to-capacity. The enterprise inference market is not without capable incumbents: AWS, Azure, and Google Cloud are investing in their own inference-optimized offerings, and the open-source cost advantage narrows as hyperscalers extend aggressive volume discounts. IBM's own cloud track record has been uneven, raising fair questions about whether its distribution channel is as powerful as the check it has written. And at an $8.3 billion valuation anchored to reported — not publicly audited — order commitments, Together AI carries execution risk that has not yet been stress-tested across a full market cycle. Three signals are worth tracking over the next six to twelve months: whether the IBM-backed cluster achieves its Q1 2027 go-live on schedule; whether the Chennai factory reaches operational status and at what utilization rate; and whether annual committed revenue grows beyond $1.15 billion or plateaus as hyperscaler competition intensifies.