High-bandwidth memory (HBM) has been a key technology driving the AI revolution. Even though HBM is more expensive than other types of memory, chip designers have packed more and more of it into AI accelerators. While customers design next-generation HBM, they have also increased capacity per XPU by adding more memory cubes, raising stack density, and increasing stack height. As a result, HBM has taken an ever-larger share of total DRAM wafer capacity—and that has ultimately produced the severe DRAM shortage we face today.
For both technical and supply-chain reasons, we are now on the cusp of a shift in this trend. Next-generation accelerators will standardize on 8-high stacks, versus the current 12-high standard. NVIDIA has already chosen this approach for Rubin Ultra: HBM capacity per GPU falls from 288 GB on standard Rubin and B300 to 192 GB. SemiAnalysis first reported this change. The supply chain is preparing for 8-high stacks to become the new standard—less than a year after the industry still expected a move to 16-high stacks or even taller.
Yet, as we have argued in our memory model, we believe this trend will continue. For inference workloads where memory bandwidth is critical, 4-high HBM offers the best bandwidth cost-performance and thus the lowest cost per chip. Capacity matters up to a threshold; beyond that, incremental HBM capacity yields diminishing returns, while those extra bits still carry the same bill-of-materials cost. That is exactly what hardware teams at the major labs want to achieve in their ASIC programs starting with next-generation HBM4. Just as hardware designers work to maximize capacity per watt under DC power constraints, 4-high HBM is an effective way to maximize chips per HBM wafer—another scarce resource.