Nearly every month now, a clear trend has emerged in AI inference accelerators: multiple companies have begun signaling on their roadmaps that they will stack DRAM directly with compute.
Qualcomm unveiled HBC (High Bandwidth Compute). Shortly after, at Hot Chips, Cerebras said the CS-6 will stack DRAM dies on top of its compute die. Samsung has been aggressively promoting its zHBM technology and calling it the endgame. There are rumors that Groq’s LP40 is also considering stacked DRAM. And at Hot Chips, d-Matrix showed its second-generation accelerator, Raptor, which also uses stacked DRAM—and announced a partnership with NVIDIA.
In this piece, I’ll explain why this trend makes sense. I’ll go into three specific reasons:
3D DRAM enables the highest-performance architecture
3D DRAM can drive energy per bit below 0.1 pJ/bit, freeing more of the power budget for compute and networking
Unlike HBM, access to 3D DRAM can be made deterministic, like SRAM
3D DRAM enables the best, highest-performance architecture
The highlight of this year’s NVIDIA GTC was the fireside chat between two giants of computing—NVIDIA Chief Scientist Bill Dally and former Google DeepMind Chief Scientist Jeff Dean. Their contributions to modern computing are enormous. Among the many topics they covered, one idea kept coming back: the most energy-efficient, highest-performance way to do computation is to put data right next to the tensor engines and move it as little as possible.
Read more