Issue 05 · Week of Sep 1–6, 2026Weekly Storage & memory infrastructure for AI

The Storage and Memory Signal

Weekly signal for the people building and buying AI infrastructure: what moved, why it matters, what's next.

By AiInfraEnergy.com Editors

The only audited benchmark in AI storage published a new round this week, and the interesting part is not the top number. It is the list of companies that did not show up. Meanwhile Google put a figure on the memory wall that should stop any server buyer cold, Micron is roughly doubling its HBM output, and Kioxia wants to sell you NAND that pretends to be DRAM. Six stories, two dives, one term.

01: The pulse

This week in four numbers

877GiB/s Top audited checkpoint write bandwidth in MLPerf Storage v3.0, from Everpure's FlashBlade//EXA at 30 data nodes
75%+ Share of an AI server's hardware bill of materials now taken by high-performance memory, per Google Cloud
100K wpm Micron's reported HBM wafer-per-month capacity target by year end, up from roughly 40K to 50K a year ago
57% Year-on-year rise in LTO tape capacity shipped in Q1 2026, after a 9.2% decline across full-year 2025
02: The index

Tracked over time

A running index of a few figures we track issue over issue, so the trend is visible, not just the snapshot. Two new component-level series start here, since spot pricing is where the memory squeeze shows up first.

IssueDateMetricValue
Issue 03Aug 30, 202630TB enterprise SSD list price$22,600
Issue 04Sep 4, 202630TB enterprise SSD list price$22,600 (flat)
Issue 05Sep 1, 2026DDR4 1Gx8 3200MT/s spot price$44.54 (+2.08% WoW)
Issue 05Aug 31, 2026512Gb TLC NAND wafer spot price$20.71 (-0.90% WoW)
03: Signal

What moved, and why it matters

Six things that happened this week, and why I'd pay attention to each one.

MLPerf Storage v3.0 lands with checkpointing, KV cache, and vector database workloads MLCommons · StorageReview · Blocks & Files

Nineteen organizations submitted 143 results, eleven of them first-timers. Everpure's FlashBlade//EXA topped checkpoint write at 877.5 GiB/s across 30 data nodes and posted 1,623.4 GiB/s of KV cache read on a 51-host configuration. Azure Managed Lustre became the first hyperscale cloud service ever submitted, at 642.2 GiB/s of checkpoint write. DDN, Dell, Huawei, IBM, NetApp, VAST Data, WEKA, Hammerspace, and Lightbits all sat this round out.

Why it mattersThe benchmark finally measures what AI storage actually spends its day doing, and the incumbents who dominate the enterprise install base chose not to be measured doing it. That absence is a data point.
Google: high-performance memory is now more than 75% of an AI server's bill of materials Google Cloud at SEMICON Taiwan · Commercial Times via TrendForce

Nikhil Cherian, senior director of supply chain infrastructure at Google Cloud, said mixture-of-experts and multimodal architectures are moving AI from compute-bound to memory-bound, with memory now the majority of server hardware cost. Google's answer is a split TPU line: the 8i for low-latency inference, the 8t for very large-scale training.

TranslationWhen the buyer with arguably the best cost visibility in the industry says three quarters of the box is memory, "GPU shortage" stops being a useful frame for budgeting. You are buying DRAM with an accelerator attached.
Micron reportedly adding up to 60K HBM wafers per month, targeting ~100K by year end ETNews via TrendForce

Micron's HBM capacity sat at roughly 40K to 50K wafers per month last year. The expansion is aimed squarely at 12-Hi HBM4 for Nvidia's Vera Rubin accelerator, with that product's share of Micron output projected to rise from 20% to 30% early this year to around half by December.

Worth notingRoughly doubling HBM wafer starts in a year is not incremental. It also removes that wafer capacity from conventional DRAM, which is precisely the mechanism that keeps DDR5 contract pricing stubborn even when spot demand softens.
Kioxia plans fast NAND on a CXL module as a partial DRAM substitute Nikkei via Blocks & Files

Pairing high-speed NAND with a CXL controller cuts read latency to under a tenth of conventional NAND, according to Kioxia. Swapping part of a system's DRAM for those modules can double memory capacity and lift performance about 30%. A related SSD with improved read and write performance is slated to begin sampling in 2028.

Why it mattersEveryone has been treating CXL as a way to pool more DRAM. Kioxia is proposing to use it to buy less. Given where DRAM pricing sits, that is a pitch with an obvious audience, though 2028 sampling means the relief is years out.
Storage earnings: Dell, NetApp, and HPE all beat, all credit AI Blocks & Files · StorageNewsletter

Dell posted $47B in Q2 FY27 revenue, up 58%, with storage at $4.9B (up 26%) and a record $95B AI server backlog. NetApp cleared $2B for the first time in a quarter, up 30%, with all-flash array revenue up 47% to $1.31B and 350 AI and data lake deals closed. HPE hit a record $12.2B, up 34%, and disclosed a $3.5B hyperscaler inferencing server deal signed after quarter close.

TranslationThe AI money is finally landing in traditional storage line items, not just in GPU servers. Note Dell's own framing though: memory shortages are pushing server prices up, so some of this growth is inflation wearing a demand costume.
Tape has a very strange year: down 9.2% in 2025, up 57% in Q1 2026 LTO Program TPCs · Blocks & Files · StorageReview

Shipped LTO capacity fell to 160.3 EB in 2025 from the 176.5 EB record set in 2024, which the LTO group attributes partly to buyers waiting on LTO-10. Then the first quarter of 2026 came in 57% above Q1 2025. Blocks & Files runs the back-of-envelope math to roughly 250 EB for the full year, which would be a record by a wide margin, with the obvious caveat that one quarter times four is not a forecast.

Worth watchingWith flash at multiples of last year's price and power the binding constraint in most facilities, the cheapest cold tier in the building suddenly looks strategic again. Tape's revival is a direct consequence of the flash squeeze, not a separate story.
04: Vendor tracker

Where the big three stand on HBM

A running snapshot, updated whenever a vendor discloses something new, not re-explained from scratch every week.

SamsungUpdated Issue 04

HBM5 targeted for ~2028 mass production at 2x HBM4E performance on a 2nm base die. Longer-term zHBM stacks DRAM directly on the processor, after 2029.

Roadmap disclosed
SK hynixUpdated Issue 03

Hybrid bonding pushed from HBM4E to HBM5 after hitting a 775-micron packaging ceiling. Racing for 16-Hi HBM4 delivery to Nvidia by Q4 2026.

Delayed
MicronUpdated Issue 05

Reportedly adding up to 60K HBM wafers per month toward roughly 100K by year end, with 12-Hi HBM4 for Vera Rubin heading toward half of output. Still racing for 16-Hi HBM4 by Q4 2026.

Capacity ramp
05: Deep cut

MLPerf Storage grew up, and half the industry declined to be graded

The v3.0 round rewrote what the benchmark measures. Who submitted, and who didn't, says as much as the throughput tables.

MLCommons published MLPerf Storage v3.0 on September 1, and the release is a structural rewrite rather than a refresh. The old suite measured training data delivery with Unet3D, Cosmoflow, and Resnet50 against simulated A100 and H100 accelerators. Version 3.0 keeps training (Unet3D and RetinaNet, now simulated on B200 and MI355 class accelerators) and adds three workloads that reflect where storage pressure actually accumulates: checkpointing across Llama 3-style models at 8B, 70B, 405B, and 1.25T parameters, a vector database test on a one-million-vector Milvus index at 1,536 dimensions, and a KV cache test measuring how many concurrent long-context conversations a system can keep fed.

PublishedSeptember 1, 2026, by MLCommons
Submissions143 results from 19 organizations, 11 of them first-time submitters
Top checkpoint writeEverpure FlashBlade//EXA, 877.5 GiB/s write and 833.0 GiB/s read at 30 data nodes
Cloud firstAzure Managed Lustre, 642.2 GiB/s checkpoint write from a 4,096 TiB deployment with 128 clients
Top training readYanRongTech F9000X at 543.9 GiB/s feeding 99 simulated B200s
Not submittingDDN, Dell, Huawei, IBM, NetApp, VAST Data, WEKA, Hammerspace, Lightbits
Comparabilityv3.0 numbers cannot be compared to v2.0; the simulated accelerator moved from H100 to B200

Start with the headline result, because it is a genuine milestone for one vendor. Everpure, the company formerly known as Pure Storage, made its first MLPerf Storage appearance and scaled FlashBlade//EXA nearly linearly with node count: 327.6 GiB/s of checkpoint write at 10 data nodes, 484.1 at 15, 655.6 at 20, and 877.5 at 30. Everpure has been making large throughput claims for //EXA for about a year. This is the first time an audited third-party result has backed the shape of those claims, and linear scaling is a harder thing to fake than a peak number.

The Azure result is the one that changes the competitive map. Azure Managed Lustre is the first hyperscale cloud storage service ever submitted to this benchmark, and it landed second in checkpointing at 642.2 GiB/s. Two rounds ago the assumption was that cloud file services were fine for convenience and hopeless for frontier-scale training I/O. A managed service posting numbers inside the same order of magnitude as purpose-built arrays undercuts that assumption in a way marketing slides could not. Nvidia's 20 AIStore submissions running S3-layer workloads on OCI, AWS, and GCP make a related point: roughly one sixth of this round used the new object-storage access layer, topping out at 133.3 GiB/s of checkpoint write. Object storage is slower than the parallel file systems here, but it is now on the same chart, and MLCommons co-chair Curtis Anderson said plainly that object-based systems may become the preferred option as AI contexts scale into the trillions.

Then there is the roster. The 3D U-Net training leaderboard is led by YanRongTech, TuringData, UBIX, Suzhou Zishan Longlin, and FarmGPU, names most Western infrastructure buyers have never evaluated. HPE was the only established enterprise array vendor to submit training results at all, with the K3000 at 345.8 GiB/s. MLCommons itself noted that traditional NAS is thin in this round and that legacy enterprise incumbents are mostly missing. There are innocent explanations, and vendors are legitimately selective about which workloads they submit. Even so: a benchmark that now measures checkpointing and KV cache is measuring exactly the thing every AI storage vendor claims to be best at, and the vendors with the largest installed bases and the loudest AI messaging mostly declined to post a number.

One caveat to carry into any vendor conversation this quarter. Because the simulated accelerator moved from H100 to B200, the per-accelerator bandwidth bar went up and v3.0 results cannot be compared to v2.0. If a vendor quotes you a v2.0 placement, that is a different benchmark. Ask which round.

Worth watchingWhether the absent incumbents show up in v4.0. If DDN, VAST, and WEKA submit next round, this was a scheduling problem. If they sit out again while newcomers and Azure keep posting, buyers should start asking why.
06: Second cut

A 245TB SSD is now a power argument, not a capacity argument

StorageReview put meters on Micron's 6600 ION against nearline disk. The interesting number is watts, not IOPS.

The SSD-versus-HDD debate used to end at dollars per terabyte, and disk won. StorageReview's September 3 analysis of the 245TB Micron 6600 ION makes the case that the tiebreaker has moved. In their test, one 6600 ION replaced eight Seagate Exos M 30TB drives in RAID 5 in the same Dell R5715 chassis, with the HDD backplane and RAID controller removed entirely. The flash configuration writing sequentially at full tilt drew 170.2W. The HDD configuration sitting idle drew 173.5W.

Swap ratioOne 245TB 6600 ION for eight 30TB nearline HDDs in RAID 5
Power170.2W flash under sequential write vs 173.5W disk at idle; ~55W saved per unit blended
Efficiency72.7 MB/s per watt vs 10.4 for the HDD array, roughly 7x on sequential reads
Rack mathAn exabyte in 6 racks instead of 22, freeing ~320 sq ft of white space per exabyte
The punchline2,182 drives, under three racks, frees the 120kW that runs a full GB200 NVL72

That last figure is the one worth repeating in a capacity planning meeting. In a power-capped building, every watt handed to storage is a watt unavailable to revenue-generating compute, and StorageReview's arithmetic says the flash swap across roughly three racks of drives liberates enough power to operate an entire NVL72 rack. The IEA projects data center electricity consumption more than doubling to around 945 TWh by 2030. Facilities absorbing that growth are not budget-constrained in the traditional sense; they are constrained by what the utility will deliver and what the cooling loop can carry.

Be honest about the boundaries, though. Nearline HDD still wins decisively on acquisition cost per terabyte, and for genuinely cold archives that advantage is not close, which is exactly why tape shipments are climbing in this same issue. The 6600 ION case holds for read-heavy bulk storage at scale inside a fixed power envelope. That is a large and growing slice of AI infrastructure, but it is a slice, not the whole plate.

Why it's hereTwo of this week's stories point the same direction: storage decisions are increasingly settled by the facility's power budget rather than the procurement spreadsheet. Density is now a compute lever.
07: Uplevel highlight

This week's term: checkpointing

One piece of storage or memory vocabulary, explained properly, every week.

Checkpointing

In plain English: periodically dumping the entire state of a training run to storage, so a crash costs you minutes instead of weeks.

A large training job holds an enormous amount of mutable state: model weights, optimizer moments, learning-rate schedules, data loader position, random number generator seeds. All of it lives in accelerator memory and system memory across thousands of GPUs. If a node fails, a network partition happens, or the job is preempted, that state evaporates. Checkpointing is the practice of writing it all to durable storage at regular intervals so the run can resume from the last saved point rather than from zero.

The catch is that checkpointing is a stop-the-world event. Classically, every rank pauses, flushes its shard of state, waits for the slowest writer, and only then resumes computing. That means the entire cluster sits idle for the duration of the write. At frontier scale the state being written is measured in terabytes, and the intervals are frequent because the failure rate across tens of thousands of components is not small. This is why MLCommons made checkpointing a first-class MLPerf Storage workload in v3.0 rather than a footnote to training: it is a verified bottleneck, and every second of flush time is a second of idle accelerators you are still paying for.

Look at what the benchmark actually asks for and the reason storage vendors care becomes obvious. The v3.0 checkpoint test runs 10 writes followed by 10 reads at four model sizes up to 1.25 trillion parameters, and scores on minimizing duration. Everpure's 877.5 GiB/s write result corresponds to a 17.74 second checkpoint. Multiply the difference between a 17 second checkpoint and a 90 second one by the number of checkpoints in a multi-week run, then multiply that by the hourly cost of the cluster, and you have the actual dollar value of a fast storage tier. That is a far more persuasive number than raw sequential throughput, which is presumably why the vendors who can post it did, and an interesting question for the ones who did not.

Reads matter too, and they are the half people forget. Checkpoints are not just written; they are read back on restart, read for evaluation, and read when a fine-tuning job forks from a base run. The benchmark scores both directions for exactly that reason.

Why it's in this newsletterCheckpointing is the workload that turned AI storage from a data-delivery problem into a write-bandwidth problem. It is the headline event of this issue's Deep Cut, and it is why "how fast can you ingest" stopped being the only question worth asking a storage vendor.
08: On the radar

What's coming up

SEP 15–17

AI Infra Summit 2026, Santa Clara. About as close as this space gets to a dedicated conference. Expect vendors to time announcements around it, and expect at least one of the MLPerf absentees to explain themselves on stage.

OCT 12–15

OCP Global Summit 2026, San Jose McEnery Convention Center. The Open Compute Project's flagship event, themed "Scaling Innovation for the AI Era." Rack-scale power, cooling, and open storage and memory designs land here, which makes it the natural venue for anything following this issue's power-versus-density thread.

SEP–OCT

Q3 earnings season continues for memory and storage vendors (Samsung, SK hynix, Micron, Western Digital, Seagate). With Micron reportedly doubling HBM wafer starts, capex guidance and any comment on conventional DRAM allocation are the lines to read first.

NOV 15–20

SC26, Chicago. The HPC and storage world's biggest annual gathering. Parallel file system and interconnect announcements tend to cluster here, and it is the first real venue where the MLPerf v3.0 results will be argued over in person.