Friday, September 11, 2026
AI 인프라 · 뉴스 & 분석
헤드라인리포트
헤드라인 · 리포트

d-Matrix와 NVIDIA는 AI 추론을 위한 NVLink Fusion 랙 시스템 개발을 계획하고 있으며, 하드웨어 배포 표준화를 위한 다년간의 제품 로드맵으로 뒷받침됩니다.

표준화된 추론 랙 아키텍처는 GPU 클라우드 배포를 간소화하고 대규모 추론 워크로드에 대한 하드웨어 활용률을 향상시킬 것입니다.
업계 전문지Slicast · September 11, 2026 · 글로벌 · 출처: HPCwire
중요도 85

d-Matrix today announced a collaboration with NVIDIA centered on an NVLink Fusion-enabled rack-level system for ultra-low latency AI inference, featuring d-Matrix's next-generation inference XPUs alongside NVIDIA's latest rack reference architecture. The multi-year product roadmap will integrate d-Matrix's Raptor XPUs directly with NVIDIA Vera CPUs, NVLink switches, BlueField-4 DPUs, ConnectX-9 SuperNICs, and Spectrum-X Ethernet networking.

As an NVIDIA NVLink Fusion partner, d-Matrix will collaborate closely with NVIDIA on the integration, leveraging the mature MGX ecosystem and supply chain for modular, cable-free trays. d-Matrix is also partnering with Astera Labs, a connectivity solution leader within the NVLink Fusion ecosystem, to deliver custom solutions ensuring high-throughput, seamless data flow throughout the system.

NVLink Fusion provides d-Matrix a mature, high-bandwidth, low-latency scale-up foundation for connecting XPUs to NVIDIA rack-scale infrastructure. The first engagement point will be d-Matrix Raptor XPUs plugging into the NVIDIA MGX rack, delivering higher performance, greater deployment flexibility, and a scalable architecture for expanding Raptor-based inference clusters as demand grows.

"This collaboration with NVIDIA is a defining moment on our journey to infinite inference, accessible to all," said Sid Sheth, founder and CEO at d-Matrix. "Being integrated into NVIDIA's latest MGX rack-scale infrastructure with NVLink Fusion means our customers can deploy our inference XPUs alongside the broadly available NVIDIA AI factory platform. That's the future d-Matrix has been building toward—ultra-low latency, energy-efficient inference XPUs and GPUs working together, at rack scale, to deliver premium AI experiences."

"NVLink Fusion enables partners to integrate custom silicon with NVIDIA's deep ecosystem of NVLink, advanced packaging, rack-scale systems and networking technologies," said Jensen Huang, founder and CEO of NVIDIA. "With NVIDIA AI infrastructure deployed across cloud and on-premises data centers worldwide, NVLink Fusion gives partners like d-Matrix a path to integrate seamlessly with NVIDIA compute platforms — expanding accelerator choice for customers building the next generation of AI factories."

"Purpose-built connectivity is what turns innovative compute into high-performing AI factories," said Jitendra Mohan, CEO of Astera Labs. "Our partnership with d-Matrix and NVIDIA brings this vision to life within the NVLink Fusion ecosystem, delivering high-throughput for low latency AI inference."

As agentic AI workloads have spiked inference demand, AI service providers increasingly seek mixed-architecture systems to deliver the inference economics their customers require. The d-Matrix MGX rack system targets latency-sensitive applications such as AI coding assistants, real-time chatbots, and voice agents where interactivity is paramount and customers are willing to pay a premium for speed. Built on NVIDIA's proven MGX ecosystem and supply chain, the system extends a unified rack architecture giving AI factories the flexibility to deploy the right compute for each workload.

Using heterogeneous disaggregation, operators can split workloads between d-Matrix Raptor XPUs and NVIDIA Vera Rubin, optimizing each inference phase. In popular disaggregated applications like AI coding, GPUs handle the compute-intensive prefill phase while d-Matrix inference XPUs accelerate the latency-sensitive decode phase.

A follow-on to the d-Matrix Corsair XPU platform currently in production, Raptor extends d-Matrix's memory-centric architecture through a first-of-its-kind 3D DRAM stacking approach that brings a DRAM memory chip and an SRAM compute chip together in a single "two-story" package. Technical details were recently published by IEEE and previewed by d-Matrix co-founder and CTO Sudeep Bhoja at the 2026 Hot Chips conference.

d-Matrix designed Raptor—expected to tape-out before year-end—from the ground up for integration with NVIDIA NVLink Fusion and the NVIDIA MGX rack-scale ecosystem. The platform is being actively evaluated at AI hyperscalers and frontier labs for its unique memory stacking solution and is backed by more than 100 patents.

Initial availability of d-Matrix Raptor XPUs integrated into the NVIDIA MGX rack is expected in Q4 2027. Demonstrations will be available at AI Infra Summit at the d-Matrix booth, or for more information, contact d-Matrix at www.d-matrix.ai/contact-sales.

원문 보기