Friday, August 28, 2026
DarkSubscribe
AI Infrastructure · News & Analysis
HomeChips & HardwareReport
Chips & Hardware · Report

d-Matrix claims its new chip delivers 20x bandwidth density over NVIDIA Rubin.

Introduces a direct performance challenge to NVIDIA’s upcoming interconnect standards, potentially reshaping accelerator packaging economics.
Trade pressSlicast · August 25, 2026 · US · Source: Google News
importance 80

In AI inference workloads, two distinct phases dominate: prefill and decode. The decode phase is overwhelmingly the more time-consuming portion because it is strictly memory-bandwidth bound. Regardless of available compute power—which benefits prefill significantly—decode performance remains constrained by memory throughput. To address this bottleneck, companies like Groq (whose technology was recently acquired by NVIDIA) and d-Matrix have developed chips featuring massive SRAM arrays to deliver exceptional memory bandwidth. d-Matrix now introduces a new approach that tackles the bandwidth challenge from a fundamentally different angle: beneath the compute layer.

The new accelerator, dubbed Raptor, leverages 3D DRAM technology. Much like AMD’s second-generation 3D V-Cache, which places an SRAM cache die beneath Ryzen 9000 CPU cores to optimize thermal performance, d-Matrix stacks its memory directly below the logic die. However, instead of SRAM, Raptor utilizes up to four layers of DRAM to maximize capacity. While this configuration offers lower density than High Bandwidth Memory (HBM), d-Matrix claims it achieves approximately 20 times the bandwidth density over NVIDIA’s upcoming Rubin architecture, alongside a 13.5-fold reduction in power consumption per gigabyte transferred.

These metrics directly address critical limitations facing modern AI processors as HBM approaches its physical boundaries. Because HBM stacks sit adjacent to the compute die on a silicon interposer, bandwidth is inherently constrained by the finite number of I/O pins that can be placed along the chip’s perimeter. Furthermore, routing data horizontally across the interposer requires dedicated interface drivers (PHYs) that consume substantial power, often accounting for hundreds of watts in high-end packages. Raptor circumvents these constraints by employing 36-micron face-to-face bonding to fuse a TSMC N4 logic die directly atop the DRAM stack. This design eliminates the PHY entirely, allowing electrons to travel vertically over mere micrometers across the die’s entire footprint. The result is an I/O energy reduction to approximately 0.37 pJ/bit while delivering SRAM-class bandwidth.

Although Raptor’s 32 GB per card capacity is lower than typical HBM configurations, this trade-off aligns with its intended purpose: AI inference rather than training massive frontier models. Positioned between ultra-expensive, low-capacity SRAM blocks and high-density HBM stacks, Raptor is designed for horizontal scaling. d-Matrix plans to deploy the chips in rack-scale clusters, such as 72-card setups interconnected via high-speed fabrics. In this configuration, aggregate memory bandwidth scales linearly into the petabytes-per-second range, enabling seamless parameter distribution across nodes and supporting multi-million-token context windows without bottlenecks during the decode phase.

d-Matrix’s performance projections are highly ambitious. Beyond architectural slides, detailed specifications remain limited. The company asserts that Raptor can deliver approximately 1,000 tokens per second per user while serving a “3T-class model at 1M context.” Supporting benchmarks indicate that a full rack configuration achieved 2,121 tokens per second per user across 32 concurrent users running GLM 5.2, and 785 tokens per second per user with 32 users executing the Kimi K3 model. The recent Hot Chips presentation primarily highlighted d-Matrix’s engineering solutions for stacking DRAM beneath logic—a concept that, while not novel, presents significant manufacturing hurdles. As of now, these figures represent projections rather than independently verified results, and no live product demonstration has been conducted. Nevertheless, d-Matrix has already shipped commercial products, distinguishing it from early-stage ventures making unverified claims. Further details and third-party benchmarks will be essential to validate whether Raptor can meet its stated targets.

Read the original
d-Matrix claims its new chip delivers 20x… · Slicast