NVIDIA details its upcoming Vera server CPU architecture at Hot Chips 2026, featuring 88 cores per chip designed to support next-generation rack-scale deployments.
For the third and final presentation in the morning’s opening CPU session at Hot Chips 2026, NVIDIA took the stage to unveil Vera, its next-generation server CPU. Built around the new Olympus CPU core architecture, Vera is an in-house Arm-based server processor featuring 88 cores. It serves as a critical component of NVIDIA’s upcoming Vera Rubin AI systems and represents the company’s most ambitious effort to expand its footprint in the server CPU market, targeting a larger share with a specialized yet highly capable processor. The presentation follows the release of the Vera/Olympus whitepaper last month.
At a high level, Vera is engineered as a modern, high-performance processor that prioritizes instructions per cycle (IPC) over sheer core count. The chip integrates eight 128-bit LPDDR5X memory controllers and features an NVLink-C2C interface designed to pair seamlessly with NVIDIA’s Rubin GPUs. The interface also supports connections to other NVLink-C2C implementations or a second Vera processor for a dual-socket (2P) pure CPU configuration. As the successor to NVIDIA’s Grace CPUs—which powered both the Grace Hopper and Grace Blackwell generations—Vera delivers significant advancements across every metric. NVIDIA frames its development around the reality that artificial intelligence represents the most complex computational workload to date. Rather than a single query, AI operates as a full workflow spanning multiple tools and computation types, requiring optimization across the entire stack rather than a single use case.
Vera Rubin was developed through extreme hardware and software co-design, encompassing architectures such as NVL72, LPX3, and Bluefield racks. Over the past year, NVIDIA has frequently illustrated the performance frontier—the trade-off between token throughput and user interactivity. While high batching maximizes token generation, it degrades response times. Vera Rubin is engineered to push this frontier outward, elevating the optimal balance between throughput and interactivity. At higher interactivity levels (measured in tokens per user per second), Vera Rubin delivers up to 30 times the total throughput of Grace Blackwell, though results remain dependent on positioning along the performance curve; Grace Blackwell experienced performance collapse at an earlier threshold. For deployments requiring capabilities beyond standalone Vera Rubin configurations, NVIDIA now offers Groq LPX3 racks. These systems leverage Groq’s dedicated decode acceleration hardware to extend the performance curve further, enabling very high interactivity rates without degradation. NVIDIA reports that long-context decode on Groq LPX3 is four times faster than current public services.
The design philosophy behind Vera reflects a fundamental shift in server CPU economics. Traditional server designs historically favored renting out numerous weaker cores, but AI workflows demand higher single-threaded performance. Consequently, NVIDIA prioritized a high-IPC architecture optimized for latency-critical operations. In agentic systems, where autonomous agents transition rapidly between tasks, minimizing completion time directly accelerates subsequent operations—a scenario that inherently benefits high-IPC designs like Vera. NVIDIA highlighted agentic workloads as a compelling use case, noting that headless browsers running within agentic frameworks exhibit distinct performance profiles compared to traditional interactive browsers. Overall, NVIDIA positions Vera as differentiated from conventional server CPUs through five key innovations: high IPC, power efficiency, massive data movement capacity, security and confidential computing capabilities, and comprehensive I/O bandwidth. Deterministic performance under heavy load remains a central promotional focus.
The Olympus CPU core, first detailed in NVIDIA’s recent whitepaper, incorporates substantial on-die resources. Each core features massive buffers, a 10-wide decode front-end, and a correspondingly large execution backend to consume decoded instructions. While optimized for agentic workloads, Olympus is not exclusively agentic-only. NVIDIA implemented statically partitioned spatial multithreading rather than dynamic resource sharing across threads. This approach delivers performance parity with traditional simultaneous multithreading (SMT) while significantly reducing susceptibility to thread contention and resource noise. In benchmark testing using a curated subset of SPEC CPU 2026 mapped to representative agentic and data processing workloads, NVIDIA reported performance gains of approximately 1.8x for agentic proxies and 1.5x for data processing proxies, explicitly noting these figures represent proxy metrics rather than exhaustive benchmarks.
Memory subsystem selection aligned with strict energy efficiency constraints, making LPDDR5X the definitive choice. Due to reticle size limitations, NVIDIA could not fabricate Vera on a single monolithic die. Instead, the processor utilizes a multi-chip module approach comprising six dies: a monolithic compute core die paired with separate chiplets handling memory and I/O functions. All components reside on a single interposer. The system I/O die provides 96 lanes of PCIe/CXL connectivity, while NVLink-C2C facilitates high-speed links to Rubin GPUs or secondary Vera processors in dual-socket configurations.
Confidential computing represents another foundational priority for Vera and its accompanying Rubin architecture. NVIDIA engineered a full-stack confidential compute solution spanning the entire rack, ensuring tamper-proof operation when deployed alongside compatible software platforms such as NVIDIA’s DOCA ecosystem. Achieving this required intensive co-design across the CPU, GPU, and network interface controller. Within this architecture, the GPU validates incoming requests from the CPU, leveraging per-core encryption keys to maintain security integrity. To support integration with the most critical data center assets, Vera incorporates 96 lanes of PCIe and CXL, utilizing these lanes to connect directly to NVIDIA ConnectX network interface cards.
Vera will be available as part of the NVL72 platform and as standalone Vera servers. NVIDIA and its partner ecosystem will offer both rackscale systems and individual server units. The presentation concluded with acknowledgment of the engineering team’s achievements, particularly the successful transition to an in-house designed CPU core. This architectural shift enabled NVIDIA to realize their exact performance targets while demonstrating expanded capabilities in custom silicon design.