Together AI Inference Exchange, September 2026: The neocloud's strategy to become enterprise AI's default infrastructure layer
The Equinix-NVIDIA-Together AI Inference Exchange, covering 260-plus colocation facilities, caps a summer in which the company signed a $240M IBM deal, a roughly $1.2–1.5B India factory build, and a Saudi national-AI partnership — pushing its annual-contract run-rate past $1.15 billion at an $8.3 billion valuation.
The September 3 announcement of the Equinix-NVIDIA-Together AI Inference Exchange is more consequential than its press-release format implies. For the first time, a neocloud built on open-source foundations is embedding its inference stack directly inside Equinix's colocation footprint — more than 260 facilities across roughly 70 metro markets — rather than asking enterprise customers to route sensitive workloads to a separate public cloud endpoint. The platform is described as offering secure, low-latency access to NVIDIA compute and more than 200 open-source models within physically proximate, customer-controlled colo environments, a configuration that directly addresses the two objections that most reliably slow enterprise AI adoption in regulated industries: data residency and round-trip latency.
The timing is not coincidental. Together AI closed an $800 million Series C in July 2026 at an $8.3 billion valuation, with Aramco Ventures leading and NVIDIA participating alongside Vista Equity Partners. The company reported an annual-contract-value run-rate exceeding $1.15 billion at the time of that close, and the months since have produced a cascade of named, corroborated transactions. IBM and Together AI announced a $240 million multi-year agreement in mid-August to deploy a large-scale NVIDIA HGX B300 inference cluster on IBM Cloud — reported target: Q1 2027 — and Together AI stated at announcement that it expected the capacity to be subscribed before launch. In India, Larsen and Toubro secured what the company publicly described as a mega order, reported by multiple outlets at roughly Rs 10,000 crore, or approximately $1.2 to $1.5 billion, to construct India's first 10,000-GPU unified AI cluster for Together AI in Chennai using NVIDIA B300 hardware. In Saudi Arabia, state-backed HUMAIN formalized a strategic partnership with Together AI on September 2 to anchor the kingdom's national AI infrastructure backbone, selecting Together AI software alongside AMD accelerators and Cisco networking.
The pattern across these engagements is deliberate. Together AI is not selling generic GPU rental capacity; it is positioning its inference software layer as the operating platform for politically and commercially strategic AI deployments — a sovereign Saudi project, an Indian manufacturing-scale cluster, an IBM-branded regulated-enterprise offering, and now Equinix's globally distributed colo network. The company has acknowledged looking abroad for additional data-processing capacity to supplement domestic compute constraints, and the international deal flow reflects that posture directly. The NVIDIA thread — investor, hardware supplier, and co-signatory on the Inference Exchange — runs through virtually every announced transaction, which reduces hardware-compatibility friction for customers but deepens a supplier dependency that will not go unnoticed by competitors or acquirers.
The commercial logic is reinforced by the market structure these deals are happening inside. Reporting from mid-2026 documented Together AI, alongside RunPod and Nebius, capitalizing on AWS GPU-supply tightness to win AI startups that had historically defaulted to hyperscaler infrastructure. The deeper structural argument — that enterprises are simultaneously shifting toward open-source models and finding hyperscaler responsiveness constrained by the same supply crunch — is precisely the window Together AI's Provisioned Throughput product, launched in July to standardize inference capacity reservations and pricing predictability, was designed to occupy. The Inference Exchange extends that logic into the physical layer: capacity pre-positioned in enterprise-adjacent colo, accessible without a hyperscaler intermediary.
The risks are real and worth naming. Capital intensity at this scale — sovereign infrastructure builds across two countries, a nine-figure IBM deployment, a global Equinix integration — requires sustained execution across multiple regulatory and geopolitical jurisdictions simultaneously. The HUMAIN partnership notably runs on AMD accelerators rather than NVIDIA silicon, introducing software-stack integration complexity that the company has not yet publicly addressed. Country-level exposure in Saudi Arabia and India adds a layer of risk that a company operating primarily from San Francisco would not have carried two years ago. The Inference Exchange model is differentiated today, but its terms depend on Equinix and NVIDIA remaining aligned partners; either could deepen its own inference offerings or renegotiate commercial arrangements as the market matures. And at an $8.3 billion valuation without disclosed profitability metrics, the pricing reflects a conviction about sustained outperformance against a competitive field that now includes every major hyperscaler with a freshly capitalized inference roadmap.
Three signals are worth watching as this thesis develops. First, whether the IBM B300 cluster comes online on its stated Q1 2027 schedule and whether Together AI confirms publicly that capacity was subscribed before launch — that would validate both execution discipline and enterprise demand depth. Second, the pace and disclosed metrics of Inference Exchange rollout across Equinix metros: named enterprise customers, latency benchmarks, and contract volumes would distinguish genuine traction from a distribution arrangement. Third, how Together AI manages the emerging AMD-versus-NVIDIA split in its international footprint — if its software layer proves genuinely hardware-agnostic, that broadens its addressable market and reduces single-vendor dependency; if it proves implicitly NVIDIA-native, the HUMAIN partnership represents an integration risk rather than a strategic hedge.