Tuesday, September 29, 2026
AI 인프라 · 뉴스 & 분석
홈 › 데이터센터 › 리포트
데이터센터 · 리포트

Across hyperscale data centers, energy delivery and power distribution have become the primary performance bottleneck, surpassing GPU deployment as the limiting constraint.

Power scarcity forces infrastructure investment toward grid upgrades, cooling systems, and power distribution networks rather than GPU procurement, reshaping AI buildout economics.
업계 전문지Slicast · 2026년 9월 28일 20:00 UTC · 글로벌 · 출처: HPCwire
중요도 68

Across hyperscale data centers, the limiting factor is no longer how many accelerators can be deployed, but how much energy can be delivered, distributed, and used effectively. In many environments, the ceiling isn't defined by chip density but by the power the system and facility can sustain.

Global data center electricity consumption is projected to reach 945 terawatt-hours by 2030—roughly doubling from today's levels—driven largely by AI workloads, according to the International Energy Agency. As a result, system design is increasingly shaped not by theoretical compute capacity, but by how efficiently available power can be translated into usable performance.

AI performance at this scale is no longer determined by individual components. It's defined by system behavior. Thousands of compute elements operate concurrently, sharing data across complex interconnect and memory hierarchies. What matters is not just how fast any one engine can run, but how efficiently the system as a whole moves, prioritizes, and delivers data under load. The challenge is shifting from adding compute to orchestrating it—and at the center of that shift is the interconnect.

Data centers represent the most extreme form of AI deployment. Systems are designed to maximize throughput across massive clusters of heterogeneous compute—CPUs, GPUs, NPUs, and specialized AI accelerators—all operating concurrently on shared datasets. Unlike edge systems, which are constrained by fixed power and thermal budgets, data centers operate at the limits of available energy. The question is not just how much compute can be deployed, but how effectively available power can be converted into useful work. Power spent on inefficient data movement reduces the energy available for compute, and as AI systems scale, inefficient data movement increasingly consumes that budget. Interconnect efficiency has become central to the design of data center AI systems—moving more data per unit of energy directly translates into higher effective system performance.

Industry conversation often focuses on off-chip bandwidth, such as high-bandwidth memory (HBM), or on scale-up and memory-pooling implementations using CXL and PCIe, which remain critical for training workloads where memory capacity and external bandwidth are key constraints. Yet in many modern architectures, bottlenecks are increasingly distributed across both on-chip and off-chip data movement, requiring coordinated optimization across the entire system.

As AI accelerators grow in complexity, they integrate increasing numbers of compute engines—multi-core NPUs, systolic arrays, and vector units—all generating and consuming data simultaneously. On-chip interconnects must now handle high-bandwidth streaming data flows, latency-sensitive control traffic, and coherency-driven memory access patterns. Cache hierarchies are becoming more critical; local caches can deliver significantly higher effective bandwidth than external memory when data reuse is high, but they also introduce additional traffic patterns that must be coordinated efficiently. Even the fastest external memory systems cannot fully compensate for an internal fabric that cannot keep up. AI performance is increasingly determined by how well data is moved, reused, and prioritized within the system, not just by how quickly it can be fetched from outside.

Modern AI systems are inherently heterogeneous. GPUs and NPUs drive high-throughput, streaming workloads on very wide multi-threaded data streams; CPUs introduce bursty, latency-sensitive traffic; accelerators may require tightly synchronized, coherency-aware communication. Supporting this diversity requires interconnect architectures that can manage multiple traffic classes simultaneously, enforce prioritization, and maintain predictable latency under load. Quality of service (QoS), traffic isolation, and coherence domain management become critical for managing contention and prioritization. Maintaining data consistency across shared-memory domains introduces synchronization overhead that must be carefully managed to avoid performance degradation. Without intelligent arbitration and scheduling, contention quickly erodes the benefits of additional compute. The interconnect governs system behavior under real workloads, shaping not just data movement but overall performance and efficiency.

To continue scaling AI compute, the industry is rapidly adopting chiplet-based architectures. By partitioning large SoCs into multiple dies, designers can improve yield, reduce cost, and scale systems beyond reticle limits. But rather than eliminate data movement challenges, this architectural shift redistributes them. Within a chiplet, localized communication can be optimized through shorter wires, higher bandwidth, and lower latency, while memory and compute can be placed closer together. Communication between chiplets, however, introduces new constraints. Compared to on-chip interconnects, die-to-die interfaces are typically more constrained in bandwidth density, latency, and power efficiency, and can quickly become bottlenecks under AI workloads demanding continuous, high-volume data exchange. Chiplets increase the need for cross-die QoS and prioritization, efficient scheduling of inter-die traffic, and careful partitioning of workloads to minimize unnecessary data movement. They both solve and create problems—enabling scalable system design while making system-level data movement architecture even more critical.

As AI systems grow in scale and complexity, the traditional approach to interconnect design is breaking down. Handcrafted interconnects are effective in tightly scoped designs but do not scale efficiently to systems with hundreds of compute elements, multiple memory hierarchies, and complex traffic interactions. The challenge begins at the level of architectural intent. Manually specifying how data should move through a system is time-consuming, error-prone, and difficult to optimize. The shift is toward system-level modeling and exploration, automated interconnect generation, and more abstract, programmable ways to describe architecture. The goal is not simply to build interconnects faster, but to preserve architectural intent from concept through implementation and ensure that performance, power, and scalability targets are met predictably.

Historically, interconnect design has often been treated as a secondary concern, addressed after compute and memory decisions are made. In modern AI systems, the interconnect is a primary determinant of system performance, efficiency, and scalability. The next phase of AI scaling in the data center is being defined by how efficiently data can be delivered to compute within the limits of available power. This means prioritizing data movement early in the design process, treating interconnect as a first-class architectural element, and adopting solutions that combine high-performance fabrics with system-level automation. Performance is no longer determined only by the number of accelerators in a rack, but by how consistently those accelerators can be kept busy. That requires systems that can prioritize, balance, and deliver data predictably under load. The next generation of AI infrastructure will be defined not just by more compute, but by systems designed from the outset to move data as efficiently as they process it.

*Andy Nightingale is Vice President of Product Management and Marketing at Arteris, with over 40 years of high-tech industry experience. A Chartered Member of the British Computer Society and Chartered Institute of Marketing, he spent 23 years at Arm leading a product marketing team specializing in system IP products including network interconnects, memory and interrupt controllers, and system MMUs. At Arteris, he oversees the Magillem system-on-chip deployment tooling and FlexNoC and Ncore network-on-chip products.*

원문 보기
Across hyperscale data centers, energy… · Slicast