Memory Is The Lynchpin Of The IT Industry, And Doubly So For AI

AI & Compute

DDR5 Prices Up 13X Since ChatGPT Launch as Memory Becomes AI's Bottleneck

DDR5 server memory costs 9X–13X more per GB than in November 2022, and DDR4 has doubled or tripled as AI demand and fleet refreshes collide with tightening DRAM supply.

By
Sophie Lindqvist
Filed
Channel
AI & Compute
Read
4 min read

DDR5 server memory now costs 9X to 13X more per gigabyte than it did in November 2022, when OpenAI's ChatGPT ignited the generative AI boom and inverted the economics of the IT industry. DDR4 street prices have doubled or tripled over the same four years, and a 30 TB TLC enterprise SSD runs 6X to 7X higher, according to street-price tracking summarized in the source analysis.

The backdrop is a structural shift. For five decades, compute — first CPUs, then GPUs — was the focal point of system design and budgets. Memory, storage, and I/O were secondary, and organizations routinely skimped on them to maximize CPU or GPU utilization. Machine learning, generative AI, and now agentic AI have ended that hierarchy. Memory has eclipsed compute as the control point across most of the IT sector, and the speed of the reversal is what disorients industry veterans.

What do the revenue numbers show?

Gartner's latest chip revenue breakdown, which counts flash as memory, shows the overall semiconductor market nearly doubling between 2025 and 2026 — a Moore's Law-style revenue curve. Forecasts show growth decelerating between 2026 and 2027, an outcome the source calls inevitable since the GenAI boom began: at some point competition kicks in, and customers eventually have enough capacity.

Working from Gartner's stated 2026 growth rates for the two categories, the source estimates the DRAM-versus-NAND split and assumes other memory types are worth roughly $4 billion a year, growing at twice global GDP. Layering in Micron's latest total addressable market forecast for HBM shows that implied revenues for other DRAM — DDR4, DDR5, LPDDR4, LPDDR5 — will dwarf HBM sales, probably for the foreseeable future.

Why are DRAM supplies tightening?

Three demand streams are colliding. GenAI itself is a memory and flash hog. A massive fleet refresh is underway, replacing servers that are five, six, or seven years old. And AI-driven designs chase near-maximum core density to save space and power, pulling high-capacity DRAM and flash with them.

Meanwhile, memory makers are shifting conventional DRAM wafer capacity into HBM stacked memory for AI accelerators, chasing higher revenues — though not necessarily higher profits, since finished HBM cost per GB is only up 1.6X since the GenAI boom started, while yield losses mount with each generation. Traditional IT — transaction processing, data warehousing, analytics, web infrastructure — now competes directly with the tech titans for the same DRAM and flash supply.

Disk is no refuge. A 30 TB nearline drive costs about 2.5X more than four years ago, and hyperscalers and cloud builders have contracted the vast majority of supply from the three remaining makers: Seagate, Western Digital, and Toshiba.

What does HBM actually cost?

Only Samsung, SK Hynix, and Micron make HBM, and there are no spot prices — the market is dedicated to Nvidia, AMD, and the XPU makers. The pricing record, pieced together from rumors and reported figures:

  • HBM2 (2017): ~$20/GB at debut, settling to ~$8/GB by 2019
  • HBM2E: ~$10/GB
  • HBM3: ~$12/GB
  • HBM3E: ~$13.50/GB
  • HBM4: rumored ~$16/GB, with Nvidia's contract price said to be $560–$600 per 36 GB stack

Expectations that HBM4E will cost twice as much per GB as HBM4 meet skepticism given that trend line. The source can see $20/GB, but calls $36/GB tough to fathom. The prevailing idea is that HBM makers will raise HBM4E prices in 2027 to bring HBM profitability in line with standard DRAM.

These figures are not what end users pay. HBM likely represents half the cost of an Nvidia or AMD GPU to users, and closer to 65 percent of the cost of an XPU accelerator from hyperscalers, cloud builders, or AI model builders — a share that now includes OpenAI's "Jalapeno" accelerator.

How are GPU makers responding?

With supply constrained, designers are making trade-offs. Some are staying with more stacks of HBM3E instead of moving to HBM4, keeping bandwidth for inference decode while sacrificing capacity. AMD is holding at 488 GB for its "Altair" MI455X accelerator. Nvidia is rumored to be reconsidering "Rubin Ultra" plans — originally four reticle-limited chiplets with up to 1 TB of HBM4E — and is said to be weighing as little as 192 GB on some future Rubin designs.

A likelier path is architectural. HBM4 and HBM4E enable customizable base dies, and Nvidia's NVHBM moves the memory controller off the GPU chiplet into the 3D stack. As Nvidia explains: "By integrating the memory controller into the 3D HBM stack instead of the XPU, NVHBM delivers up to 30 percent greater memory bandwidth and 15 percent lower HBM power consumption, and frees up to 25 percent more area on XPU compute die compared with standard HBM4E." Others will pursue custom base dies, some embedding math functions in the manner of processor-in-memory, an approach with a decade of demos and no market traction.

The source's own bet: Rubin Ultra uses the freed area across two chiplets to deliver 512 GB of HBM4E — or 384 GB at the same bandwidth — down from the original spec of 1,024 GB and 58 TB/sec aggregate. Every GPU and XPU maker faces the same calculus: maximum HBM capacity and bandwidth on leading-edge technology, at a much lower price per GB.

Original: gartner.com

Share this article:

More from Sophie Lindqvist

Sophie Lindqvist

Show full bio

News editor covering business strategy at Chip Dispatch.

164 articles

Related articles

« Previous articleNext article »