AMD plans deeper integration of its EPYC CPUs with Radeon GPUs for data center applications.
AMD is pursuing tighter integration between its EPYC processors and Radeon Instinct GPUs, according to Scott Aylor, corporate vice president and general manager of AMD's Datacenter Solutions. "I think you'll see a greater and greater coupling in terms of timing, capability [and] workload affinity that I think will be more the norm," Aylor said. This integration strategy extends across multiple domains, including high-performance computing, machine learning, artificial intelligence, and visualization. The approach represents a significant strategic shift for AMD following a pivotal 2019, when the chipmaker launched its first 7-nanometer CPUs for desktops and servers and completed its re-entry into the data center market with its new EPYC Rome processors.
Previously, AMD launched new versions of EPYC and Radeon products on separate timelines—the Radeon Instinct MI50 and MI60 GPUs released in November 2018, while second-generation EPYC Rome released in August 2019. The company plans to synchronize these releases more closely going forward, with both hardware and software improvements including continued development on AMD's Infinity Fabric and PCIe connectivity as well as the ROCm software platform for GPU-accelerated computing. Aylor explained AMD's vision for this software integration: "When you hear the guys at Oak Ridge talk about Frontier, they have a vision that we're working to enable: you decide where you want to target the work, CPU- or GPU-based, rather than I've got to do a bunch of bespoke stuff to make the GPU work, I've got to do a bunch of bespoke stuff to make the CPU work."
The 1.5-exaflop Frontier supercomputer that AMD is developing with Cray for the U.S. Department of Energy's Oak Ridge National Laboratory serves as the north star for this integration strategy. The supercomputer, set to go online in 2021, is expected to become the world's fastest supercomputer and will use EPYC processors based on a future iteration of AMD's Zen architecture along with new Radeon GPUs and a custom Infinity Fabric interconnect. "You can think about how those have been co-architected to create a one-and-a-half exaflop system," Aylor said. "You don't get to that mountaintop in one step."
Early results demonstrate the benefits of AMD's integration approach. According to NAMD 2.13 benchmark results, a server configuration consisting of two EPYC 7742 processors and eight Radeon Instinct MI50 GPUs is 26 percent faster than a server with two Intel Xeon Platinum 8280 processors, two PCIe 3.0 switches and eight Nvidia Tesla V100 GPUs. The performance advantage stems largely from PCIe 4.0 support in AMD's latest processors, which enables the AMD system to achieve throughput speeds of 512 GB/s for CPU-to-GPU communication compared to 64 GB/s in the Intel system. As Ogi Brkic, corporate vice president and general manager of AMD's Datacenter GPU business, noted, "It's all about feeding the beast." Additionally, AMD's configuration eliminates the need for PCIe switches, reducing system costs while decreasing latency and maintaining persistent performance.
While AMD emphasizes openness and compatibility with rival products from Intel and Nvidia, the company is prioritizing what it calls "A plus A"—delivering a better integrated experience through AMD GPU and AMD CPU combinations. Dominic Daninger, vice president of engineering at Nor-Tech, a Burnsville, Minnesota-based HPC system integrator, observed that AMD's CPU-GPU integration work addresses throughput bottlenecks in servers as processors increasingly outpace data transfer speeds. "Intel sees the same thing with their GPU products: this ability to move data between [the CPU and GPU] at higher bandwidths is going to be important for both camps," Daninger said.