NVIDIA is open sourcing its cuFile storage APIs and launching the Storage-Next initiative with over 40 storage vendors t
At this week's Future of Memory and Storage conference, NVIDIA unveiled storage advancements demonstrating that AI infrastructure depends as much on storage efficiency as on computing power itself. Surging AI demands are driving the need for massive datasets and context windows that exceed system memory capacity, while AI agents generate thousands of concurrent storage requests.
Storage systems serving these requests must continuously encrypt, compress, verify and reconstruct data. When thousands of agents access storage simultaneously, these critical data services become significant bottlenecks. NVIDIA's Vera CPU, part of the Vera BlueField-4 STX, delivers up to 3.21x higher throughput than x86 CPUs in two-stage compression and encryption pipelines, allowing storage platforms to absorb AI data more efficiently with significantly less compute infrastructure.
With accelerated computing, storage becomes an active part of the data path rather than a passive repository. This fundamentally changes the economics of where data belongs. Forty years ago, the tradeoff between accessing data in memory versus on storage drives was measured in minutes. Today, paired with AI storage solutions, that same tradeoff plays out in microseconds.
NVIDIA announced it is open sourcing its cuFile application programming interfaces and the vertical storage software stack underneath them, enabling GPUs rather than just CPUs to read from and write to storage directly. Using hundreds of thousands of GPU threads and fast high-bandwidth memory, cuFile enables secure data access from storage in microseconds. This represents how the industry is unifying a security-first storage stack based on Linux best practices, providing interoperability between GPUs and data.
NVIDIA also launched Storage-Next, an initiative bringing together over 40 leading storage and flash vendors including DDN, KIOXIA and Micron to align on GPU-driven storage behavior and create interoperable industry standards. The initiative is grounded in accelerated data access for large AI datasets and introduces SCADA, or scaled accelerated data access, a framework that lets massively parallel GPUs pull only necessary data directly from storage into their own high-speed memory. For example, DDN is integrating SCADA with Infinia, its software-defined, AI-native data intelligence platform built to eliminate storage bottlenecks at scale.
According to Sven Oehme, chief technology officer at DDN, "AI success will be defined not by how much infrastructure organizations own, but by how productively they use it. Our collaboration with NVIDIA is helping create a more direct, efficient connection between GPUs and data — keeping accelerated computing resources productive, speeding time to insight and enabling customers to achieve stronger business and financial returns from their AI investments."
These advancements build on NVIDIA's Vera BlueField-4 STX, a modular rack-scale foundation powered by the NVIDIA Vera Rubin platform and NVIDIA Spectrum-X Ethernet networking. The platform uses the unified NVIDIA DOCA security stack to enable continuous policy enforcement in the AI data path. NVIDIA CMX Context Memory Storage provides an AI-native context tier for long-context, multi-turn agentic AI inference.
Direct storage access offers speed but requires careful implementation to prevent security vulnerabilities. NVIDIA SCADA achieves scaled direct access through a two-part approach where user-facing application components that need raw speed remain outside the trusted computing base, while a separate privileged component configures protected access between the user application and approved storage at setup, adhering to standard Linux security protocols. These advancements in fast, massively parallel, efficient, secure storage infrastructure enable applications and AI factories to produce more useful, accurate and grounded intelligence at scale.