Memory: How It Works and Why It Sets the Cost of Compute
The AI systems of today are incredibly powerful - they can write code better than most software engineers, analyze documents, generate images, and can reason through complex problems to reach coherent solutions.
Frontier models wait on memory more than they wait on arithmetic. Prefill and decode keep weights and a growing key-value cache in high-bandwidth memory, so HBM capacity and bandwidth set how much context, occupancy, and therefore usable compute a GPU can deliver. The essay’s thesis is that memory, not transistors, is now the price-setting input. HBM cannot expand smoothly: stacking yield, CoWoS packaging, and a three-vendor allocation market reprice at contract resets rather than on a continuous spot curve. Because memory already dominates the bill of materials on leading accelerators, a discrete HBM step change reprices whole clusters at once. The practical question is whether buyers and lenders can treat HBM as a scarce, hedgeable commodity instead of a background component cost.
Read the full essay on Substack