Saturday, August 8, 2026
DarkSubscribe
AI Infrastructure · News & Analysis
HomePower & EnergyReport
Power & Energy · Report

Deep learning acceleration drives unprecedented increases in datacenter power density requiring infrastructure redesign

Creates power and cooling bottleneck necessitating facility upgrades, novel thermal solutions, and grid capacity expansions
Trade pressSlicast · March 27, 2017 · Global · Source: datacenterknowledge.com
importance 80

Rob Ober, chief platform architect at Nvidia's Accelerated Computing Group, oversees the design of Tesla, the most powerful GPU on the market for Machine Learning. Building computer systems for Artificial Intelligence is, as Ober notes, "a really hard engineering problem to build a system for GPUs for deep learning training. It's really, really hard. Even the big guys like Facebook and Microsoft struggled." GPUs, or Graphics Processing Units, originally designed for graphics, have become essential processors for Deep Learning, the type of AI that Google uses to serve targeted ads and Amazon Alexa uses for instantaneous answers to voice queries.

Deep learning involves two categories of computing workloads: training and inference. Training teaches a deep neural network a new capability from existing data—for example, teaching a neural net to recognize dogs in photos by repeatedly viewing tagged images. Inference is where a neural net applies its knowledge to new data, such as recognizing a dog in an image it hasn't seen before. Training is especially difficult in the data center because it requires extremely dense clusters of GPUs, with interconnected servers containing up to eight GPUs per server. One such cabinet can easily require 30kW or more—power density most data centers outside of the supercomputer realm aren't designed to support. About 20 such cabinets require as much power as the Dallas Cowboys jumbotron at the AT&T stadium, the world's largest 1080p video display, which contains 30 million lightbulbs.

The engineering challenges extend beyond raw power consumption. "With deep learning training you typically want to make as dense a compute pool as possible, and that becomes incredibly power-dense, and that's a real challenge," Ober explained. Another critical problem is controlling voltage, as GPU computing produces significant power transients—sudden spikes in voltage that are "difficult to deal with." Networking also presents a major obstacle. "Depending on where your training data comes from it can be an incredible load on the data center network," Ober said. "You can be creating a real intense hot spot." According to Ober, power density and networking are "probably the two biggest design challenges in data center systems for deep learning."

Hyper-scale operators like Facebook and Microsoft typically address the power density challenge by spreading deep learning clusters over many racks, though some "dabble" in liquid cooling or liquid-assist solutions. Liquid cooling delivers chilled water directly to chips on the motherboard (a common supercomputer approach), while liquid-assist cooling brings chilled water to a heat exchanger that cools air pushed through servers. Specialized high-density data center providers, which lack the luxury of hundreds of thousands of square feet of space that hyperscalers possess, have increasingly adopted liquid-assist cooling and seen recent spikes in demand.

This surge in demand reflects broader enterprise interest in machine learning. "Right now the GPU-enabled workloads are the ones where we're seeing the largest amount of growth, and it's definitely the enterprise sector," said Chris Orlando, co-founder of high-density data center provider ScaleMatrix. "The enterprise data center is not equipped for this." Orlando noted that ScaleMatrix experienced hockey stick-shaped growth with the knee around the middle of last year. Beyond general machine learning, demand is also driven by computing for life sciences and genomics—with the J. Craig Venter Institute among the largest customers at ScaleMatrix's San Diego data center—as well as geospacial research and big data analytics.

Read the original
Deep learning acceleration drives… · Slicast