ScaleFlux는 GPU 메모리 압력을 완화하기 위해 Nvidia CMX 및 KV 캐시 오프로드를 지원하는 AI 최적화 SSD 플랫폼을 선보입니다.
MILPITAS, Calif., July 30, 2026 — ScaleFlux, a leader in advanced memory controller and storage optimization technology, today announced an AI-optimized SSD platform designed to meet the storage requirements of NVIDIA CMX and other KV-cache-intensive AI inference infrastructure. The platform integrates Context-Insight SSD for workload analysis and optimization, delivers 7 to more than 10 effective drive writes per day (DWPD) for KV cache workloads, and supports more than 200 Flexible Data Placement (FDP) write streams per drive for fine-grained, lifecycle-aware data placement.
The platform addresses three interconnected challenges that arise as AI inference systems adopt SSDs as a shared context tier beyond GPU HBM and host memory: understanding real-time workload behavior, segregating KV cache blocks with differing lifecycles, and sustaining intensive write workloads without triggering excessive SSD capacity inflation or replacement costs.
Long-context inference, shared-prefix reuse, agentic applications, and idle-session retention are rapidly expanding the volume of reusable KV cache and other runtime state that AI systems must maintain. While SSDs offer a cost-effective shared context tier, this function imposes new demands on the drive. KV blocks may be written frequently, retained for varying durations, reactivated after idle periods, and invalidated asynchronously across numerous sessions, workers, or tenants.
ScaleFlux’s high-endurance architecture is engineered to deliver 7 to more than 10 effective DWPD over five years for KV cache workloads, contingent upon workload characteristics, FDP utilization, and device configuration. Higher effective endurance reduces the amount of SSD capacity infrastructure operators must provision solely to absorb write traffic. This allows more installed capacity to store useful KV cache and AI runtime state rather than functioning primarily as endurance overhead, directly mitigating the endurance-driven capacity cost—often termed the “endurance tax”—associated with high-churn AI inference runtime-state workloads.
The ScaleFlux Context-Insight SSD capabilities are designed to complement the recently announced NVIDIA CMX Context Memory Storage Platform, offering AI factories a high-performance, high-efficiency storage system optimized for KV cache. CMX provides a shared, pod-level context tier for high-speed KV-cache access and reuse, while ScaleFlux addresses the endurance, data-placement, and write-amplification requirements of the underlying SSD tier.
Support for more than 200 FDP write streams per drive enables system software to partition KV data based on expected lifecycle, session or tenant ownership, shared-prefix classification, reuse behavior, or other software-defined categories. Grouping data with similar lifecycles together reduces garbage-collection movement, lowers write amplification, limits cross-class interference, and further enhances effective endurance. In preliminary controlled testing, ScaleFlux recorded more than a twofold reduction in write amplification using lifecycle-aware FDP placement compared to a baseline configuration. Actual performance will vary based on workload characteristics, lifecycle classification, software integration, and device configuration.
Context-Insight SSD augments these endurance and placement capabilities by demonstrating how software-level KV cache policies impact actual SSD behavior. The platform monitors latency, queue depth, throughput, request-size distribution, data age, write-to-first-read timing, read reuse, NAND write volume, garbage-collection movement, and write amplification. In standalone SSD mode, Context-Insight initiates workload characterization without requiring modifications to the upper software stack. When integrated with software, it correlates high-fidelity SSD telemetry with metadata such as session ID, worker or tenant ID, shared-prefix ID, KV-block ownership, lifecycle state, and key-to-block mapping. This ownership-aware analysis helps pinpoint which sessions, prefixes, tenants, or lifecycle classes are driving SSD latency, endurance consumption, and write amplification.
Together, these capabilities provide ScaleFlux and its partners a practical pathway from measurement to production optimization. Context-Insight identifies and quantifies storage-efficiency opportunities, while scalable FDP placement and 7–10+ effective DWPD supply the mechanisms required to translate those insights into reduced write amplification, enhanced endurance, and improved AI infrastructure economics.
“AI inference infrastructure needs SSDs that provide more than additional capacity,” said Hao Zhong, CEO and Co-Founder of ScaleFlux. “Infrastructure teams need to understand how KV workloads affect the drive, separate data according to lifecycle, and sustain high write rates without deploying excess capacity simply to dilute writes. ScaleFlux brings workload intelligence, scalable FDP placement, and 7-10+ effective DWPD together in one AI-optimized SSD platform.”
ScaleFlux is also developing a trace-driven simulator that models KV cache movement across GPU HBM, host memory, and SSD tiers. The simulator generates replayable SSD traces to evaluate placement, eviction, and lifecycle-grouping policies under controlled conditions.
“As AI inference systems extend KV cache beyond GPU Memory and DRAM, understanding the behavior and requirements of the SSD tier becomes increasingly important,” said Jason Hardy, Vice President of Storage Technology at NVIDIA. “Our engagement with ScaleFlux is helping characterize how KV cache offload affects storage requirements for latency, endurance, and write amplification, contributing to the broader storage ecosystem around NVIDIA CMX.”
ScaleFlux will showcase the platform at FMS, highlighting Context-Insight workload analysis, KV metadata correlation, support for more than 200 FDP write streams per drive, lifecycle-aware placement, write-amplification reduction, and high-endurance operation tailored for write-intensive KV cache workloads.
The platform aligns with ScaleFlux’s broader strategy of enabling SSDs to manage increasingly valuable AI runtime state through workload intelligence, software-assisted lifecycle placement, and controller-level endurance optimization.
About ScaleFlux
ScaleFlux is a semiconductor solutions company delivering advanced storage and memory technologies designed to transform data infrastructure in a scalable and sustainable manner. Founded in 2014 and headquartered in Milpitas, California, ScaleFlux develops innovative storage and memory controller technologies for AI, cloud computing, data center, enterprise, and edge applications. For more information, visit www.scaleflux.com.