d-Matrix presented its Raptor AI accelerator at Hot Chips 2026, positioning it as the first 3D DRAM accelerator for generative inference. The design integrates a TSMC 4nm compute die directly onto a custom-designed DRAM die using a face-to-face bond at a 36-micron pitch. This configuration delivers 100 TB/s of bandwidth from 32GB per card.
The vertical interface of the Raptor accelerator demonstrates an energy cost of 0.37 pJ/bit, which is significantly lower than the approximately 2.4 pJ/bit required for data movement into an HBM4 base die. An accompanying ISCA 2026 paper, co-authored with the University of British Columbia, projects that Raptor could achieve around 4.7 times higher throughput per card compared to designs utilizing HBM. These figures are based on early silicon and d-Matrix's own projections.
Raptor's architecture inverts the typical 3D stacking arrangement by placing the logic die on top and the DRAM underneath. This allows a cold plate to sit directly on the compute silicon, with the DRAM die serving as an interposer for PCIe and die-to-die signals via TSVs. This design choice enables higher bandwidth for the same power consumption. The one-high stack has a power density of approximately 0.5W per square millimeter, manageable with liquid cooling.
The custom DRAM is engineered for a junction temperature of 105°C, where retention time decreases from 32ms to 4ms, necessitating an eight-fold increase in refresh frequency. To mitigate the performance impact of more frequent refreshes, d-Matrix reduced each microbank to 1,366 rows and about 5.33MB. This design ensures that a full refresh sweep consumes only 1.37% of the overall bandwidth. The die also incorporates 840 banks per chiplet, with 72 (around 9%) designated as spares, managed by a two-level mux chain for fault tolerance.
✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →
One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.
One email a day. Unsubscribe in one click, any time.
Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.
▶ Play today's briefNew every morning, and the back catalogue is archived by date.
d-Matrix introduced Raptor, an AI accelerator featuring a TSMC 4nm compute die bonded directly onto a custom DRAM die, achieving 100 TB/s bandwidth per card. This design inverts traditional 3D stacking to improve energy efficiency and throughput for generative AI inference, potentially offering higher performance than HBM-based solutions.