Analysis indicates that Nvidia NVSwitch will ultimately serve as the primary interconnect fabric for scale-up AI networks, functionally replacing traditional InfiniBand architectures.
In the early days of computing, two rival I/O architectures competed for dominance in PCs and servers. As the dot-com boom accelerated, industry players reached a consensus that established InfiniBand’s switched fabric as the standard I/O interconnect for computing.
However, the subsequent dot-com bust led vendors to abandon the standard in favor of Ethernet for point-to-point interconnects—such as those linking PCs and servers to the internet or internal networks—and for what is now termed scale-out networking to aggregate compute across multiple machines. Amid economic constraints, the industry also standardized on upgraded PCI-X and later PCI Express buses for peripheral connectivity.
InfiniBand survived as a high-performance, low-latency interconnect through Voltaire and Mellanox Technologies. Mellanox supplied switch ASICs and network interface cards (NICs), while Voltaire developed switches tailored for HPC modeling and simulation workloads. In November 2010, Mellanox acquired Voltaire for $218 million to consolidate its InfiniBand switching leadership. InfiniCon Systems, founded by former Unisys engineers, entered the market and eventually evolved into QLogic—the primary competitor to Mellanox in InfiniBand—before Intel acquired it for $125 million in January 2012. That QLogic engineering team was later spun out of Intel to form Cornelis Networks in September 2020. TopSpin, another InfiniBand switch vendor, was acquired by Cisco Systems in 2005. Meanwhile, Sun Microsystems, guided by co-founder Andy Bechtosheim, adopted InfiniBand as the clustering technology for its “Constellation” HPC clusters and developed proprietary InfiniBand ASICs and NICs. Oracle maintained this strategy after acquiring Sun Microsystems in April 2009 for $5.6 billion. CEO Larry Ellison further strengthened Oracle’s position by purchasing InfiniBand switch maker Xsigo Systems, before ultimately pivoting to Ethernet as its unified scale-up, scale-out, and scale-across networking platform.
Founded in July 2023, the Ultra Ethernet Consortium aims to combine InfiniBand’s low latency and high bandwidth, along with its quality-of-service and adaptive routing capabilities, with Ethernet’s superior scalability, multitenancy, and QoS features. The goal is to create a scale-out network that decisively outperforms InfiniBand, supporting AI clusters targeting one million endpoints. Ethernet is also penetrating Nvidia’s scale-up domain via NVSwitch interconnects, which link GPU memories into a single coherent, shared memory space for 72 accelerators. The architecture can be expanded to support 576 devices, albeit with additional latency hops.
For the record, I value both InfiniBand and NVSwitch. InfiniBand remained the undisputed low-latency leader for twenty-five years and retains a competitive edge today. This explains why I estimate Nvidia generated $25.74 billion over the trailing twelve fiscal months ending in July—a figure nearly identical to the $25.67 billion I attribute to combined sales of Ethernet and NVSwitch interconnects. Nevertheless, this likely represents peak InfiniBand adoption for large-scale clusters. Hyperscalers, cloud providers, and AI model developers are increasingly standardizing on upcoming Ultra Ethernet technologies. Concurrently, Ethernet ASIC vendors—including Broadcom, Cisco Systems, and Nvidia itself—are aggressively reducing port-hop and end-to-end latency within Ethernet fabrics.
Over a decade ago, I argued that Ethernet would move too quickly to displace InfiniBand. I expect InfiniBand to remain viable for the foreseeable future, primarily through Nvidia, which is essentially the sole major supplier today alongside certain custom HPC interconnects in China. These Chinese solutions are based on InfiniBand—licensed and sanctioned by the formerly independent Mellanox—but may have been modified sufficiently to no longer carry the InfiniBand designation. Still, Ethernet will likely relegate InfiniBand to a specialized niche, leveraging its scaling advantages and broad compatibility across campus, edge, and data center networks. Looking ahead, Nvidia’s own sales will be dominated by NVSwitch for scale-up and Spectrum-X Ethernet for scale-out. This shift is driven by customer demand rather than vendor promotion. Enterprises and AI model builders alike prefer Ethernet—even Meta Platforms, which previously championed InfiniBand in earlier AI cluster deployments.
The longstanding networking adage holds that Ethernet always prevails, and it consistently does. Its success stems from adopting proven innovations from competing fabrics and adapting them for broad enterprise deployment, as well as for specialized HPC and AI workloads. This dynamic applies equally to scale-out networks that cluster systems and connect users, as it does to scale-up networks that interconnect GPUs and XPUs within AI clusters. Fast Ethernet is already emerging as a transport layer for alternative memory-sharing architectures.
This includes UALink, which challenged NVLink and NVSwitch upon its founding in May 2024 by AMD, Broadcom, Cisco Systems, Google, Hewlett Packard Enterprise, Intel, Meta Platforms, and Microsoft. Notably, AMD utilizes Broadcom Tomahawk 6 Ultra Ethernet switches to implement the UALink protocol for scale-up networking within its “Helios” racks. UALink operates as a memory-coherent protocol for linking XPUs and GPUs, though it can equally facilitate memory sharing among CPUs and DPUs. Similarly, the ESUN protocol, championed by Meta Platforms and Microsoft, has attracted support from AMD, Arista Networks, Arm, Cisco, Hewlett Packard Enterprise, Marvell, Nvidia, OpenAI, and Oracle. Broadcom participates in the ESUN/SUE-T initiative—where SUE-T manages load balancing and higher-level functions within the Ethernet scale-up stack—but has withdrawn from UALink and, to date, shows no interest in the NVLink/NVSwitch ecosystem. However, this stance could shift if Broadcom’s CPU and XPU clients in its semiconductor design services division choose to adopt NVSwitch as their rack-scale fabric and deploy Nvidia-style server racks housing both GPUs and XPUs.
Predicting the ultimate outcome remains difficult. Recently, Nvidia purchased $3.5 billion in convertible bonds issued by Taiwanese system-on-chip (SoC) manufacturer MediaTek. Nvidia co-founder and CEO Jensen Huang described MediaTek on Bloomberg as the “biggest and best maker of SoCs in the world”—a claim AMD and Intel would rightly contest. Under the agreement, MediaTek gains access to the NVLink Fusion stack, enabling third-party chip designers to integrate NVLink ports and participate in the NVSwitch memory fabric. Importantly, Nvidia does not sell NVSwitch ASICs or chassis independently. Deployment requires either a Grace or Vera CPU, or a Blackwell or Rubin GPU, and presupposes procurement of Nvidia racks equipped with NVSwitch as the scale-up interconnect.
Eventually, Nvidia may welcome sales of NVSwitch networking to customers deploying NVLink ports across CPUs, GPUs, and XPUs. Such a shift would undoubtedly require corresponding software enablement. To date, however, Nvidia’s official position prohibits it.
In the long term, several developments will shape scale-up networking in AI systems. Currently, Nvidia commands approximately 95 percent of GPU revenue and roughly 75 percent of combined GPU/XPU revenue, granting it a dominant share of scale-up networking via NVLink and NVSwitch. This market position is expected to decline only marginally through calendar year 2027.
Nevertheless, hyperscalers, cloud providers, AI developers, and large enterprises strongly prefer open, widely adopted standards that ensure hardware interoperability. NVLink Fusion appears designed to disrupt UALink and ESUN/SUE-T development cycles, potentially delaying industry consolidation for one or two generations.
Tech giants possess the capital to establish proprietary standards for their own hardware, yet historical precedent suggests they will ultimately converge on common protocols. This mirrors the 2014 shift toward 25 Gbps signaling for 100 Gbps and faster Ethernet, which occurred after Google, Microsoft, and Arista Networks grew impatient with the IEEE’s reluctance to abandon 10 Gbps signaling across ten lanes. Faced with significant industry pressure, the IEEE ceded control to the consortium.
UALink can compete directly with NVLink and NVSwitch on a port-for-port basis and, in theory, scale nearly twice as far. The UALink 1.0 specification, released in April 2025, supports up to 1,024 XPUs or GPUs within a single switching tier. By contrast, achieving memory sharing across 576 GPUs requires a two-tier NVSwitch architecture.
There is a chance—some might even say a probability—that open, Ethernet-based scale-up fabrics will ultimately prevail, fundamentally reshaping the competitive landscape for AI interconnects.