FMS 2026 | Part 3 of 4
How memory and storage turn constrained infrastructure into more useful AI work
Part 1 framed the memory wall in the agentic AI era. Part 2 laid out the hierarchy: an optimized five-tier stack for reasoning, a four-tier stack for execution, and similar memory and storage technologies doing different jobs under different pressures. The same placement logic now has to extend beyond KV cache to the rest of the infrastructure: how do we make every constrained resource in the data center produce more useful AI work? That is where the portfolio matters. Built tier by tier, mapped to both stacks, each product solves a different problem. Together, they improve the productive utilization of the whole system.
The point is not that Micron has a product in each category. It is that the categories now depend on each other. Agentic AI will not be won by adding stranded compute to already-constrained data centers. It will be won by increasing productive utilization: getting more useful work out of the power, space and silicon we already have. HBM keeps the hottest context moving. SOCAMM2 and RDIMM keep larger working sets close enough to matter. CXL-pooled memory puts idle capacity back to work. SSDs preserve state. High-density data-lake storage keeps long-term context available. This is not a collection of parts. It is the architecture agentic AI will require.
The leap starts closest to the XPU
Inference is often bandwidth-bound. Even the fastest XPU, meaning the full class of AI accelerators including GPUs, TPUs and custom AI silicon, can sit idle waiting on data, and adding compute alone does not fix that. Bandwidth does. That is why HBM matters so much to the token economy: it keeps the decode pipeline fed when the active context is hottest and closest to the accelerator. In memory-bound inference workloads, higher HBM bandwidth can deliver token-throughput gains that outpace the bandwidth increase itself. The payoff comes because the same XPU becomes more productive. We are in high-volume production of 36-gigabyte, 12-high HBM4 for exactly that reason, delivering greater than 2.8 terabytes per second of bandwidth per stack and more than 20% better power efficiency versus our prior generation. Future generations go further. Custom HBM can also move selected logic closer to the memory stack, reducing pressure on the XPU and creating more room for cost, power or system-level optimization. Memory used to sit next to the XPU and wait its turn. Increasingly, it will participate more directly in the work. That is the leap.
The threshold nobody sees coming
For years, the industry treated infrastructure as a compute-first conversation: how fast is the processor? That is no longer the right question. The real question is whether the system keeps the right data in the right place. In AI, almost enough memory is the most expensive configuration you can build. Miss the threshold and performance does not decline gracefully. It collapses into spilling, eviction and recomputation. SOCAMM2 exists to solve exactly that: a high-capacity tier sitting just below HBM in the same system as the XPU. The execution stack needs the same kind of capacity at the same tier, just feeding a different processor. RDIMM does that job for the CPU, just as SOCAMM2 does for the XPU. Same tier, same idea, different chip on either side.
Fit the box to the agent
Even generous memory has a ceiling once it is locked inside one server. Whatever sits idle there stays idle there, no matter how badly the server next to it needs it. The goal is not to shrink the agent to fit the system. It is to expand the system to fit the agent. Pool memory across the rack dynamically, wherever it is needed, when it is needed: active memory, not cold overflow, tuned for the workload. CXL-based pooled memory is built on that same premise. Rack-scale systems can expose up to 40 terabytes of DDR5 in a single 4U chassis and scale beyond 160 terabytes across a rack for memory-intensive AI, HPC and database workloads. In Micron and ecosystem demonstrations, disaggregated memory has shown large gains in tokens per second while reducing XPU resource consumption for constrained workloads. That is productive utilization in its purest form: not more XPUs, better-fed XPUs. The execution stack pools the same fabric for the same reason, so CPU memory requirements do not have to be trapped inside a single box either.
Built for the workloads you cannot see yet
Nothing beats HBM and main memory for raw speed. Everyone has more data than fits there, though, and that data must live somewhere for longer than a single session. The next tier down handles that job: persistent, still fast, provisioned ahead of the workloads that will need it.
We sampled the Micron 9650, our PCIe® Gen6 SSD, more than a year before high-volume production. That head start gave the ecosystem time to prepare and helps explain why this ramp can move faster than past generational transitions. The product was timed for the next wave of agentic AI platforms, not for a catch-up cycle after those platforms arrived. For us, leadership means putting products into customers’ hands early enough for the ecosystem to prepare. The latest XPU platforms need storage that can keep pace with their data-movement demands, and PCIe Gen6 storage is designed for exactly that role.
The efficiency case is just as direct. A system that previously needed eight PCIe Gen5 drives for a target read-performance envelope may be able to reach that class of performance with four Micron 9650 drives, because the Micron 9650 delivers up to twice the read performance of Gen5 drives and significantly higher performance per watt. That means fewer drives, less power, less space and more useful work inside the same envelope. The Micron 9650 also supports liquid cooling in the E1.S form factor and is available in E1.S and E3.S form factors. It is designed for the thermal reality of a modern AI data center while still supporting air-cooled environments that need to refresh existing infrastructure with the latest hardware. The Micron 7600, our PCIe Gen5 SSD, rounds out the portfolio for workloads that do not need Gen6 yet.
The execution stack leans on the same drives for a different reason: letting agents pause and resume instead of starting over. Every restart wastes time, compute and context. Persistent local state turns that wasted motion back into productive work.
The data-lake foundation is being rebuilt
At the bottom of the hierarchy sits the data lake, where agentic AI stores and recalls everything, warming cold data into insight far more often than the past required. The reasoning stack draws on it for long-term context. The execution stack draws on the same foundation to ground every action in real enterprise data. Consolidating data onto fewer, denser drives is the right move because it lets the data center do more inside the same footprint. As capacity density rises and storage bandwidth requirements increase, storage architectures also need stronger resiliency, including dual-port configurations that support fast path failover. Hard drives do not scale with this era: performance per terabyte falls exactly when the workload demands more of it. Our 245-terabyte Micron 6600 ION drives are today’s answer: the world’s highest-capacity commercially available SSD, enabling more than 176 petabytes per rack. We are building Cloud Storage Optimized SSDs for this next wave of data growth: more bandwidth per watt, more capacity per rack and more useful AI work from a smaller physical footprint.
The wall comes down tier by tier
Five tiers on the reasoning side. Four on the execution side. Similar memory and storage technologies are doing different jobs under different pressures. That is what building for every tier means inside the data center. Breaking down the memory wall is about turning constrained infrastructure into more useful AI output: more tokens, more users, more state preserved and more long-term context available, all inside the power and space reality customers face. The foundation underneath the wall must be built before the wall can come down. Part 4 takes the same question outside our walls, to where AI runs once it leaves them.