TechRadar examines the grid, thermal, and interconnect challenges limiting data center growth.
The AI boom is quickly reaching the physical limits of grid capacity and thermal dissipation. The electric grid is stretched as never before, causing longer connection waits for AI data centers, and conventional air and liquid cooling systems struggle to meet the demands of rapidly growing facilities.
As they continue to grow at a rapid pace, AI data centers face a mix of challenges: computing power is outpacing the data transfer capabilities of GPU interconnects, volatile load swings are stretching grid power factor correction (PFC) systems, and rack power density is overloading traditional cooling systems.
Below, we break down the major thermal and electrical bottlenecks slowing the AI boom and how companies are innovating their way out.
**The thermal wall (rack-level friction)**
Standard server racks generate a lot of heat, and high-performing GPU clusters generate even more, pushing thermal limits further than ever. Modern GPU racks exceed 50 to 100 kW of power consumption per cabinet, making traditional air and liquid cooling methods insufficient.
Standard cooling systems keep clusters operating between 65 and 85 degrees, but 90 to 100 degrees triggers performance throttling to prevent permanent GPU damage. Because GPU racks draw massive electrical energy, which converts into heat, preventing them from reaching throttling limits has become harder than ever.
Traditional air cooling systems were not built for rack densities surpassing 30 to 40 kW, which AI clusters easily exceed, with 100 kW or more becoming the norm. Maintaining airflow at high densities to cool racks of 100 kW or more requires running power-hungry fans at extreme speeds. This saps electrical energy and worsens the Power Usage Effectiveness (PUE) of data centers. To cool mega AI clusters with air alone, operators would need "hurricane-grade" airflow, which is impractical. That leaves liquid cooling.
Liquid cooling relies on fluid mechanics: cold plates use microchannels to force turbulent flows directly over chips to keep them cool. Keeping enough fluid flowing through these microchannels requires massive pumping power, an increasingly significant constraint for data center operators. To mitigate this, operators are pivoting to direct-to-chip cold plates mounted directly on the processor, rather than over it, which keeps chips cool while consuming less pumping power.
Phase-change mechanics are also helping GPU clusters cool more efficiently. Here, clusters use the latent heat of liquid vaporization or solid-liquid transitions to absorb the massive heat spikes from high-density GPU racks. Specialized dielectric fluid comes into direct contact with the hot GPUs and boils into vapor, instantly pulling heat away. The vapor then rises to a condenser, where it converts back into liquid droplets that drip back into the pool to repeat the process. This process, known as two-phase immersion cooling, reduces cooling energy by up to 40%.
Some data center operators are also pivoting to closed-loop liquid cooling systems, in which recirculated fluid cools the GPUs with little to no evaporation loss, eliminating the need for frequent top-ups.
For example, Microsoft, a top-three data center operator by capacity, has implemented a closed-loop system in newer data centers. Water is filled once and circulated to servers to absorb heat, moved to air-chilled coolers to cool down, and then recirculated to absorb more heat without requiring fresh supplies. The system is costly and will take time to roll out at scale, but it demonstrates how companies are innovating against thermal bottlenecks.
**The interconnect and board-level limits**
GPU clusters rely on electrical copper or optical interconnects to share data, memory, and workloads. Electrical copper interconnects have long been the standard, but heavy GPU workloads have pushed them to their limit.
Electrical signals lose strength over long distances (greater than two meters), and usable reach halves every time bandwidth requirements double. Massive GPU clusters have made communication distances longer than ever. Closely packed copper interconnects also generate electromagnetic interference that disrupts other signals, requiring complex equalization systems that drain more power.
Power is also lost across copper traces on GPU circuit boards due to resistance. The more current that passes through copper, the more heat it generates, requiring more cooling energy, which is itself already a bottleneck.
With electrical interconnects reaching their absolute limits, data center owners have little choice but to pivot to optical interconnects, which convert data into light pulses that travel through fiber optic cables with minimal loss over long distances. These cables are far thinner than copper bundles, creating more airflow paths in dense, heat-generating GPU racks.
Companies are even moving from pluggable fiber optic cables to silicon photonics, or co-packaged optics (CPO), in which optical circuits are embedded directly on GPUs to transmit light over long distances. For now, optical interconnects are much costlier than electrical copper, in some GPU setups up to 7 times more. However, there is little choice but to absorb this cost in the short term, because copper has reached its physical limits. As more companies adopt optical interconnects and invest in mass production, costs will likely fall.
**The grid and substation crisis**
Finally, grid connection is the primary challenge affecting AI data centers. Modern electric grids were not designed to handle the rapidly increasing loads of AI data centers, with 30 GW of new capacity added worldwide from 2021 to 2025 alone. These rapid additions have stretched the supply chain for grid components and introduced engineering challenges that grid operators must work around.
For instance, high-voltage transformers, which ensure safe voltage delivery to GPU clusters, are in short supply. Wait times have spiked from 6 to 12 months as of 2020 to two to four years currently, far longer than expensive mega GPU clusters can wait to be powered.
Existing power factor correction systems on grid networks have been strained by massive load draws from new AI data centers. GPU clusters shift power demand within seconds: one second of low inference demand draws stable power from the grid, and the next second draws excessive power because of usage spikes. These fluctuations create complex harmonics that overheat transformers and affect the grid, requiring grid operators to invest heavily in mitigation systems.
AI clusters are scaling to gigawatt levels, pushing voltage step-down switchgear to its most tolerable limits. Larger transformers, needed to serve these gigawatt-scale clusters, lower electrical impedance, which can force potentially high short-circuit fault currents down a line. If a short circuit occurs, magnetic fields exert powerful mechanical forces on switchgear busbars, causing them to bend or rip free from their supports. The result is structural failure, which grid operators and isolated data center power networks cannot allow.
Companies have turned to various solutions to counter these electrical bottlenecks. Chief among them is a major shift toward solid-state transformers (SSTs), which use high-frequency power electronics and semiconductors to step down voltage rather than relying on magnetic fields like conventional transformers.
High-frequency electronics enable solid-state transformers to actively regulate voltage and filter harmonics, resulting in fewer fluctuations. They are 97% to 99% efficient, compared with 95% to 97% for conventional transformers, and that 1 to 2 percentage point difference adds up to a great deal in gigawatt-level data centers. They are also much smaller and lighter than conventional transformers, enabling faster logistics and deployment.
However, SSTs introduce their own challenges. The internal high-frequency components and chips that power them are expensive, often five times or more per kVA than low-frequency alternatives. They also require multiple power conversion stages (AC-DC and DC-AC/DC-DC), and tiny losses compound during these conversions, making peak efficiency difficult to achieve. Internal components generate heat and age faster, requiring more active maintenance than conventional transformers.
Cost is also an issue, as SSTs are currently much more expensive to build and maintain than magnetic transformers. With more investment in mass production, however, costs will likely fall over time, allowing AI GPU clusters to adopt them faster. Siemens, the world's second-largest transformer manufacturer, recently partnered with fellow German tech firm Rein