Huawei is deploying its OceanStor KV cache storage architecture specifically engineered for hyper-scale AI data centers.
Huawei has announced the OceanStor M900, a scale-out, all-flash storage cluster that provides an up to 64 PB KV cache base tier for its Atlas 960 SuperPods—rack-scale AI accelerators comparable to Nvidia’s SuperPOD designs.
Nvidia’s AI Data Platform reference design integrates its GPUs, BlueField NICs, Spectrum-X switches, and AI software with storage systems that support GPUDirect, RDMA, and KV cache extensions through its Dynamo software and CMX scheme. This architecture employs a multi-tier design: GPU high-bandwidth memory (HBM) serves as the top and fastest tier, followed by associated x86 server DRAM as level 2, server-local SSDs as level 3, and intelligent BlueField-4 NIC-connected NVMe SSDs housed in dedicated flash storage servers as level 3.5. The L3.5 tier enables a single network hop to a GPU’s HBM, delivering microsecond-class latency.
Huawei is now challenging Nvidia’s SuperPOD architecture with its own SuperPods tailored for hyper-scale AI data centers. These systems are engineered to handle 10-trillion-parameter models and context windows exceeding one million tokens, meaning a single accelerator’s HBM cannot hold all required key-value tokens and necessitates a robust caching scheme.
According to the company, a single Atlas 960E SuperPod can scale to 4,096 NPUs, delivering 8 EFLOPS of FP8 compute performance, up to 1 petabyte of HBM capacity, and a 256 TB unified memory pool. By deploying 5,500 Hi-ONE units utilizing UnifiedBus near-packaged optics (NPO), Huawei reduces power consumption by more than 550 kilowatts compared to the 48,000 traditional 800G optical modules previously required to interconnect all NPUs. This optical approach also doubles the system’s fault-free operating time, achieving 99.8 percent system availability.
The SuperPod employs a multi-tier KV caching strategy, with the M900 supplying a petabyte-scale KV cache for the L3.5 layer. It delivers up to 4 PB of shared L3.5 KV cache capacity and 40 TB/sec of aggregate bandwidth per cluster via optical networking, providing terabytes of KV cache capacity per NPU. Huawei claims that in typical AI programming scenarios, this architecture doubles the inference cluster’s token throughput and halves the time to first token (TTFT).
The stated 40 TB/sec bandwidth represents a 1.5x improvement over peer solutions, though Huawei did not name specific competitors. Industry analysis suggests this generically refers to DDN, Everpure, IBM Storage Scale, MinIO, and VAST Data systems, which typically offer cluster-scale bandwidth in the 10–25 TB/sec range.
The M900 features an integrated architecture that consolidates the CPU, network controller, and NAND controller into a single unit. This design gives SuperPod NPUs a direct, one-hop connection to the SSDs, reducing access latency from milliseconds to approximately 60 microseconds.
David Wang, Deputy Chairman of Huawei’s Board and Rotating Chairman, stated: “OceanStor M900 also uses hybrid media and an optimized retention algorithm, extending SSD read/write lifespan by 16-fold. This ensures a higher KV cache hit rate alongside long-term stability and reliability from the ground up.”
In greater detail, Huawei notes that the M900 incorporates KV-aware adaptive storage technology, which predicts the expected lifetime and value of each KV cache data segment. Based on these predictions, the system schedules and places data across different media tiers, including on-chip memory, DRAM, and SSDs. This optimized placement and retention strategy allows the SSDs to sustain up to 24 drive writes per day (DWPD)—a notably high figure—claiming a 16x increase in SSD endurance and supporting three years of stable operation with significantly fewer drive replacements.
Huawei has not yet released official datasheets or technical backgrounders for the M900. Consequently, details regarding node rack unit dimensions, controller specifications, drive count and capacity, cluster node counts, and other architectural specifics remain undisclosed.
Huawei’s Ascend Neural Processing Units are designed from the ground up specifically for AI workloads, unlike GPUs that originated as graphics processors. They utilize Huawei’s Da Vinci architecture, featuring Cube (matrix) and Vector cores, and support low-precision formats such as FP8 and FP4 to accelerate large-model inference and training. When scaled into massive systems via Huawei’s UnifiedBus all-optical interconnect, thousands to hundreds of thousands of NPUs operate as a single logical machine. Recent and upcoming products, including the Atlas 350 (powered by the Ascend 950PR) and the forthcoming Atlas 960 SuperPods, are positioned as viable alternatives to Nvidia GPUs within the Chinese market.