Weekly signal for the people building and buying AI infrastructure: what moved, why it matters, what's next.
The only audited benchmark in AI storage published a new round this week, and the interesting part is not the top number. It is the list of companies that did not show up. Meanwhile Google put a figure on the memory wall that should stop any server buyer cold, Micron is roughly doubling its HBM output, and Kioxia wants to sell you NAND that pretends to be DRAM. Six stories, two dives, one term.
A running index of a few figures we track issue over issue, so the trend is visible, not just the snapshot. Two new component-level series start here, since spot pricing is where the memory squeeze shows up first.
| Issue | Date | Metric | Value |
|---|---|---|---|
| Issue 03 | Aug 30, 2026 | 30TB enterprise SSD list price | $22,600 |
| Issue 04 | Sep 4, 2026 | 30TB enterprise SSD list price | $22,600 (flat) |
| Issue 05 | Sep 1, 2026 | DDR4 1Gx8 3200MT/s spot price | $44.54 (+2.08% WoW) |
| Issue 05 | Aug 31, 2026 | 512Gb TLC NAND wafer spot price | $20.71 (-0.90% WoW) |
Six things that happened this week, and why I'd pay attention to each one.
Nineteen organizations submitted 143 results, eleven of them first-timers. Everpure's FlashBlade//EXA topped checkpoint write at 877.5 GiB/s across 30 data nodes and posted 1,623.4 GiB/s of KV cache read on a 51-host configuration. Azure Managed Lustre became the first hyperscale cloud service ever submitted, at 642.2 GiB/s of checkpoint write. DDN, Dell, Huawei, IBM, NetApp, VAST Data, WEKA, Hammerspace, and Lightbits all sat this round out.
Nikhil Cherian, senior director of supply chain infrastructure at Google Cloud, said mixture-of-experts and multimodal architectures are moving AI from compute-bound to memory-bound, with memory now the majority of server hardware cost. Google's answer is a split TPU line: the 8i for low-latency inference, the 8t for very large-scale training.
Micron's HBM capacity sat at roughly 40K to 50K wafers per month last year. The expansion is aimed squarely at 12-Hi HBM4 for Nvidia's Vera Rubin accelerator, with that product's share of Micron output projected to rise from 20% to 30% early this year to around half by December.
Pairing high-speed NAND with a CXL controller cuts read latency to under a tenth of conventional NAND, according to Kioxia. Swapping part of a system's DRAM for those modules can double memory capacity and lift performance about 30%. A related SSD with improved read and write performance is slated to begin sampling in 2028.
Dell posted $47B in Q2 FY27 revenue, up 58%, with storage at $4.9B (up 26%) and a record $95B AI server backlog. NetApp cleared $2B for the first time in a quarter, up 30%, with all-flash array revenue up 47% to $1.31B and 350 AI and data lake deals closed. HPE hit a record $12.2B, up 34%, and disclosed a $3.5B hyperscaler inferencing server deal signed after quarter close.
Shipped LTO capacity fell to 160.3 EB in 2025 from the 176.5 EB record set in 2024, which the LTO group attributes partly to buyers waiting on LTO-10. Then the first quarter of 2026 came in 57% above Q1 2025. Blocks & Files runs the back-of-envelope math to roughly 250 EB for the full year, which would be a record by a wide margin, with the obvious caveat that one quarter times four is not a forecast.
A running snapshot, updated whenever a vendor discloses something new, not re-explained from scratch every week.
HBM5 targeted for ~2028 mass production at 2x HBM4E performance on a 2nm base die. Longer-term zHBM stacks DRAM directly on the processor, after 2029.
Roadmap disclosedHybrid bonding pushed from HBM4E to HBM5 after hitting a 775-micron packaging ceiling. Racing for 16-Hi HBM4 delivery to Nvidia by Q4 2026.
DelayedReportedly adding up to 60K HBM wafers per month toward roughly 100K by year end, with 12-Hi HBM4 for Vera Rubin heading toward half of output. Still racing for 16-Hi HBM4 by Q4 2026.
Capacity rampThe Storage and Memory Signal reaches people who actually buy and operate AI storage and memory infrastructure, AI DevOps engineers, infra leads, and the executives they report to. If that's your buyer, this slot is available.
The v3.0 round rewrote what the benchmark measures. Who submitted, and who didn't, says as much as the throughput tables.
MLCommons published MLPerf Storage v3.0 on September 1, and the release is a structural rewrite rather than a refresh. The old suite measured training data delivery with Unet3D, Cosmoflow, and Resnet50 against simulated A100 and H100 accelerators. Version 3.0 keeps training (Unet3D and RetinaNet, now simulated on B200 and MI355 class accelerators) and adds three workloads that reflect where storage pressure actually accumulates: checkpointing across Llama 3-style models at 8B, 70B, 405B, and 1.25T parameters, a vector database test on a one-million-vector Milvus index at 1,536 dimensions, and a KV cache test measuring how many concurrent long-context conversations a system can keep fed.
Start with the headline result, because it is a genuine milestone for one vendor. Everpure, the company formerly known as Pure Storage, made its first MLPerf Storage appearance and scaled FlashBlade//EXA nearly linearly with node count: 327.6 GiB/s of checkpoint write at 10 data nodes, 484.1 at 15, 655.6 at 20, and 877.5 at 30. Everpure has been making large throughput claims for //EXA for about a year. This is the first time an audited third-party result has backed the shape of those claims, and linear scaling is a harder thing to fake than a peak number.
The Azure result is the one that changes the competitive map. Azure Managed Lustre is the first hyperscale cloud storage service ever submitted to this benchmark, and it landed second in checkpointing at 642.2 GiB/s. Two rounds ago the assumption was that cloud file services were fine for convenience and hopeless for frontier-scale training I/O. A managed service posting numbers inside the same order of magnitude as purpose-built arrays undercuts that assumption in a way marketing slides could not. Nvidia's 20 AIStore submissions running S3-layer workloads on OCI, AWS, and GCP make a related point: roughly one sixth of this round used the new object-storage access layer, topping out at 133.3 GiB/s of checkpoint write. Object storage is slower than the parallel file systems here, but it is now on the same chart, and MLCommons co-chair Curtis Anderson said plainly that object-based systems may become the preferred option as AI contexts scale into the trillions.
Then there is the roster. The 3D U-Net training leaderboard is led by YanRongTech, TuringData, UBIX, Suzhou Zishan Longlin, and FarmGPU, names most Western infrastructure buyers have never evaluated. HPE was the only established enterprise array vendor to submit training results at all, with the K3000 at 345.8 GiB/s. MLCommons itself noted that traditional NAS is thin in this round and that legacy enterprise incumbents are mostly missing. There are innocent explanations, and vendors are legitimately selective about which workloads they submit. Even so: a benchmark that now measures checkpointing and KV cache is measuring exactly the thing every AI storage vendor claims to be best at, and the vendors with the largest installed bases and the loudest AI messaging mostly declined to post a number.
One caveat to carry into any vendor conversation this quarter. Because the simulated accelerator moved from H100 to B200, the per-accelerator bandwidth bar went up and v3.0 results cannot be compared to v2.0. If a vendor quotes you a v2.0 placement, that is a different benchmark. Ask which round.
StorageReview put meters on Micron's 6600 ION against nearline disk. The interesting number is watts, not IOPS.
The SSD-versus-HDD debate used to end at dollars per terabyte, and disk won. StorageReview's September 3 analysis of the 245TB Micron 6600 ION makes the case that the tiebreaker has moved. In their test, one 6600 ION replaced eight Seagate Exos M 30TB drives in RAID 5 in the same Dell R5715 chassis, with the HDD backplane and RAID controller removed entirely. The flash configuration writing sequentially at full tilt drew 170.2W. The HDD configuration sitting idle drew 173.5W.
That last figure is the one worth repeating in a capacity planning meeting. In a power-capped building, every watt handed to storage is a watt unavailable to revenue-generating compute, and StorageReview's arithmetic says the flash swap across roughly three racks of drives liberates enough power to operate an entire NVL72 rack. The IEA projects data center electricity consumption more than doubling to around 945 TWh by 2030. Facilities absorbing that growth are not budget-constrained in the traditional sense; they are constrained by what the utility will deliver and what the cooling loop can carry.
Be honest about the boundaries, though. Nearline HDD still wins decisively on acquisition cost per terabyte, and for genuinely cold archives that advantage is not close, which is exactly why tape shipments are climbing in this same issue. The 6600 ION case holds for read-heavy bulk storage at scale inside a fixed power envelope. That is a large and growing slice of AI infrastructure, but it is a slice, not the whole plate.
One piece of storage or memory vocabulary, explained properly, every week.
In plain English: periodically dumping the entire state of a training run to storage, so a crash costs you minutes instead of weeks.
A large training job holds an enormous amount of mutable state: model weights, optimizer moments, learning-rate schedules, data loader position, random number generator seeds. All of it lives in accelerator memory and system memory across thousands of GPUs. If a node fails, a network partition happens, or the job is preempted, that state evaporates. Checkpointing is the practice of writing it all to durable storage at regular intervals so the run can resume from the last saved point rather than from zero.
The catch is that checkpointing is a stop-the-world event. Classically, every rank pauses, flushes its shard of state, waits for the slowest writer, and only then resumes computing. That means the entire cluster sits idle for the duration of the write. At frontier scale the state being written is measured in terabytes, and the intervals are frequent because the failure rate across tens of thousands of components is not small. This is why MLCommons made checkpointing a first-class MLPerf Storage workload in v3.0 rather than a footnote to training: it is a verified bottleneck, and every second of flush time is a second of idle accelerators you are still paying for.
Look at what the benchmark actually asks for and the reason storage vendors care becomes obvious. The v3.0 checkpoint test runs 10 writes followed by 10 reads at four model sizes up to 1.25 trillion parameters, and scores on minimizing duration. Everpure's 877.5 GiB/s write result corresponds to a 17.74 second checkpoint. Multiply the difference between a 17 second checkpoint and a 90 second one by the number of checkpoints in a multi-week run, then multiply that by the hourly cost of the cluster, and you have the actual dollar value of a fast storage tier. That is a far more persuasive number than raw sequential throughput, which is presumably why the vendors who can post it did, and an interesting question for the ones who did not.
Reads matter too, and they are the half people forget. Checkpoints are not just written; they are read back on restart, read for evaluation, and read when a fine-tuning job forks from a base run. The benchmark scores both directions for exactly that reason.
AI Infra Summit 2026, Santa Clara. About as close as this space gets to a dedicated conference. Expect vendors to time announcements around it, and expect at least one of the MLPerf absentees to explain themselves on stage.
OCP Global Summit 2026, San Jose McEnery Convention Center. The Open Compute Project's flagship event, themed "Scaling Innovation for the AI Era." Rack-scale power, cooling, and open storage and memory designs land here, which makes it the natural venue for anything following this issue's power-versus-density thread.
Q3 earnings season continues for memory and storage vendors (Samsung, SK hynix, Micron, Western Digital, Seagate). With Micron reportedly doubling HBM wafer starts, capex guidance and any comment on conventional DRAM allocation are the lines to read first.
SC26, Chicago. The HPC and storage world's biggest annual gathering. Parallel file system and interconnect announcements tend to cluster here, and it is the first real venue where the MLPerf v3.0 results will be argued over in person.