HBM solves AI's core bottleneck—the memory wall—by stacking DRAM on a wide silicon-interposer bus beside the GPU, lifting bandwidth from 256 GB/s (HBM2) to 2 TB/s (HBM4) and capacity from 8GB to 64GB, though supply stays concentrated in three makers led by SK Hynix's 58% share.
The Memory Wall: Why Does AI Computing Need HBM?
The memory wall describes a mismatch in which GPU compute units process data far faster than memory can supply it, and since large language model inference repeatedly reads weights back from memory during each pass, memory bandwidth—not raw compute—becomes the practical limit on inference speedCITE:E1.
HBM's 3D Stacked Architecture: Breaking Traditional Memory Design Limits
High Bandwidth Memory (HBM, 高頻寬記憶體) stacks DRAM dies vertically on top of a silicon interposer wired with an extremely wide data bus, a 3D packaging approach that trades chip footprint for bandwidth far beyond what traditional DDR or GDDR memory can deliverCITE:E2. Because this stacked package sits directly next to the GPU, the physical distance data must travel is shortened, compounding the bandwidth gain already achieved by the wide-bus architectureCITE:E2.
HBM vs. DDR/GDDR: How Bandwidth Differences Reshape AI Accelerators
A single DDR5 DIMM channel used in conventional servers delivers roughly tens of gigabytes per second, and consumer GPUs running GDDR6/7 reach roughly hundreds of gigabytes per second, while a single AI accelerator carrying multiple HBM stacks reaches a combined bandwidth in the terabytes-per-second rangeCITE:E7. This gap is the reason HBM exists as a distinct standard purpose-built for AI and high-performance computing workloads rather than as a general-purpose graphics or server memory formatCITE:E7.
A Decade, Eightfold: HBM's Leap From HBM2 to HBM4 in Performance and Capacity
Per-stack bandwidth rose from 256 GB/s under the 2016 HBM2 standard to 2 TB/s under the 2025 HBM4 standard—a nearly eightfold increase over a decade—while maximum stack capacity climbed in parallel from 8-Hi/8GB to 16-Hi/64GBCITE:E3CITE:E4.
| Generation | Standard Year | Bandwidth per Stack | Max Stack Height / Capacity |
|---|
| HBM2 | 2016 | 256 GB/s | 8-Hi / 8GB |
| HBM2E | — | 307 GB/s | — |
| HBM3 | 2022 | 819 GB/s | 16-Hi / 64GB |
| HBM3E | 2024 (mass production) | ~1.2 TB/s | — |
| HBM4 | 2025 | 2 TB/s | 16-Hi / 64GB |
The combination matters because bandwidth and capacity scaled togetherCITE:E3CITE:E4: a wider pipe alone would not help if an accelerator still could not hold a large model's weights close to the compute die, and the jump from 8GB to 64GB per stack lets a single AI accelerator keep far more of a model's weights in nearby memoryCITE:E4.
The HBM Supply Bottleneck: Manufacturing Constraints and Market Concentration Risk
HBM production is manufacturing-constrained because vertical stacking requires through-silicon via (TSV) technology to connect each DRAM layer, a process with yield and packaging complexity far higher than planar DRAM, and this difficulty is why capacity expansion has not kept pace with the surge in AI demandCITE:E6. Supply is also concentrated: in 2026 Q1, SK Hynix held 58% of the HBM market, down from 69% in the same period a year earlier, while Samsung Electronics and Micron each held roughly 21% and tied for second placeCITE:E5.
| Supplier | 2026 Q1 Share | Year-Earlier Share |
|---|
| SK Hynix | 58% | 69% |
| Samsung Electronics | ~21% | — |
| Micron | ~21% | — |
All three suppliers' HBM4 products have already been certified by NvidiaCITE:E5.
What This Means
The same reason HBM exists—closing the gap between GPU compute speed and memory bandwidth described by the memory wallCITE:E1—is what makes its stacked, wide-bus design necessary rather than optionalCITE:E2, and the resulting terabyte-per-second scale is precisely what separates it from DDR and GDDR memoryCITE:E7. That same stacking, however, is also what makes HBM hard to produce: the TSV process behind a decade of nearly eightfold bandwidth and capacity gainsCITE:E3CITE:E4 is the identical constraint now limiting how fast three suppliers—one of them losing share—can expand output to match AI demandCITE:E5CITE:E6.