D-Matrix is developing the Raptor 3D-DRAM accelerator, scheduled for presentation at Hot Chips 2026. This accelerator is designed for generative inference, a compute-intensive task in AI. The increasing size of model weights and the KV cache, which scales with context length and batch size, create significant capacity and bandwidth challenges for current hardware architectures.
Current memory technologies face limitations in meeting the demands of large AI models. SRAM offers high bandwidth but has limited capacity and high power consumption at scale. HBM provides better capacity but struggles with bandwidth, reaching practical ceilings around 20 TB/s for HBM4 packages due to pin speed, I/O width, and package beachfront constraints. High HBM bandwidth also incurs substantial power costs.
D-Matrix's solution involves stacking compute logic directly on top of DRAM dies. This 3D stacking aims to overcome the bandwidth and capacity limitations of traditional architectures. The company states that a 1-Hi logic-on-top stack operating at no more than 0.5 W/mm2 can be liquid-cooled, keeping DRAM temperatures below 100 C, despite the thermal and power delivery challenges inherent in stacking.
The 3D-DRAM approach offers improved energy efficiency compared to HBM. D-Matrix indicates that vertical 3D I/O consumes approximately 0.3 to 0.4 pJ per bit, which is about 10 times lower than HBM's 2.5 to 5 pJ range. This efficiency gain is attributed to the PHY-less, millimeter-scale path of 3D I/O, contrasting with the centimeter-scale interposer route used by HBM.
✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →
One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.
One email a day. Unsubscribe in one click, any time.
Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.
▶ Play today's briefNew every morning, and the back catalogue is archived by date.
D-Matrix will present its Raptor 3D-DRAM accelerator for generative inference at Hot Chips 2026. This technology addresses the growing capacity and bandwidth demands of large language models by stacking compute directly on DRAM dies, offering a potential solution for efficient AI inference.