Fujitsu presented its upcoming Monaka CPU at Hot Chips 2026, describing it as a highly efficient Arm-based processor optimized for data center and AI workloads.
At Hot Chips 2026, Fujitsu presented detailed specifications for Monaka, its next-generation server CPU and successor to the initial Arm-based A64FX. While the A64FX was primarily deployed in the Fugaku supercomputer, Fujitsu is targeting broader data center applications with Monaka. The company positions the chip for modern AI and agentic workloads, which increasingly demand more from CPUs beyond traditional GPU coordination. Recognizing that AI adoption is currently constrained by power limitations, Fujitsu emphasizes Monaka’s ultra-low-voltage operation and high energy efficiency as critical solutions, alongside architectural adjustments tailored to sovereign AI requirements.
Monaka is built on the Armv9.3-A architecture and utilizes a chiplet-based design assembled through 3D stacking and die-to-die hybrid bonding. The core die sits atop an SRAM die, which rests on a silicon interposer. The core chiplet is fabricated on TSMC’s N2P process node, accounting for less than 30 percent of the total silicon area, while the SRAM and I/O dies are manufactured on TSMC’s N5 node. This multi-node approach leverages GAAFET technology for density and performance gains where needed, while accelerating time-to-market. The processor integrates 144 CPU cores per chip, with scalability up to 2P per node, and features 12 channels of DDR5 memory.
Unlike most Arm designs that utilize 128-bit SVE2 SIMDs, Monaka employs a wider 256-bit SVE2 execution unit—a configuration positioned between standard 128-bit implementations and the 512-bit SVE2 found in the A64FX. Each core contains dual 256-bit SVE2 units paired with dual 256-bit load/store units. Fujitsu considers this width optimal for data center workloads, balancing throughput with practical implementation constraints.
Power consumption scales with capacitance, frequency, and the square of voltage, making voltage reduction the most effective path to efficiency. Fujitsu has implemented a non-standard design methodology requiring specialized tools to operate Monaka at approximately 30 percent lower voltages than comparable designs, though exact figures remain undisclosed. Low dropout regulators (LDOs) received particular attention, as analog circuits do not scale efficiently with advanced nodes. Additional power-saving measures include a floating-point register cache optimized for programs with high temporal locality, such as GEMM workloads; selective bypassing of the predicate register during read operations; and masking narrower-than-256-bit vectors as zeros to reduce SIMD power draw. For performance, Fujitsu prioritized full utilization of the 256-bit SIMD through features like a combined gather instruction, particularly beneficial for HPC applications.
Monaka supports three NUMA configurations: single-node (144 cores), four-node (36 cores per node), and eight-node (18 cores per node). The single-node layout targets large-memory applications, while the eight-node configuration maximizes throughput. Each NUMA node can be divided into two LLC regions, and the chip supports up to 63 MPAMs independent of the NUMA topology. Confidential computing capabilities are integrated via the Arm Confidential Compute Architecture, necessitating coordinated hardware, firmware, and software stack development.
Fujitsu is building Monaka’s software ecosystem around extensive open-source foundations, supplemented by components from NVIDIA’s open-source stack to ensure comprehensive compatibility. The processor will launch in two SKUs, both featuring 144 cores. The high-performance variant operates at 500 watts with a 2.9 GHz base frequency, delivering an estimated 50 percent performance increase over the high-efficiency SKU, which runs at 350 watts and a 2.1 GHz base frequency.
Monaka is scheduled for release in 2027. Concurrently, Fujitsu is developing Monaka-X, a follow-on processor fabricated on a 1.4 nm process node. Monaka-X will incorporate NVLink Fusion support to enhance connectivity with NVIDIA and other NVLink-compatible accelerators.