跳到主要內容

Can QLC Flash Memory Challenge HBM? SanDisk and Kioxia Unveil 9th-Gen QLC — and an AI Memory Repricing

One-sentence takeaway: AI memory demand is spreading from HBM to NAND flash. SanDisk and Kioxia's 9th-generation 2Tb QLC, paired with the new HBF architecture, could rewrite NAND's role in AI — and repric the entire AI hardware supply chain.
9th-gen 2Tb QLC 3D flash: 6-plane architecture, 4.8Gb/s NAND interface

One announcement, a 5.73% single-day surge

On August 13, SanDisk shares jumped 5.73% in a single day, ignited by the company's joint unveiling with Kioxia of 9th-generation 2Tb QLC 3D flash memory built for AI infrastructure. This is no routine product refresh — industry analysts point to a repricing of the AI memory market worth up to $94 billion, and QLC, once dismissed as the "slow, cheap" storage technology, is now at the center of it.

Why QLC was dismissed — and why it's back

QLC stores 4 bits per cell (TLC stores 3), delivering higher density and lower unit cost at the expense of speed and endurance. That made it a natural fit for consumer SSDs and cold storage — but the AI era has changed the game. Large language model weights and KV caches routinely run into the hundreds of GB or even TBs, and data centers need a memory tier that is large, fast to read, and cheap. QLC's weaknesses have become its strengths.

The 9th-generation technology uses a CBA (CMOS Bonded Array) architecture, manufacturing the CMOS circuitry and memory array separately and bonding them together, so each can evolve independently. Compared with the 8th-gen 2Tb QLC, the new parts switch to a 6-plane architecture for higher read/write bandwidth, improve power efficiency, and lift the NAND interface speed to 4.8Gb/s — a 33% jump that closes much of the performance gap with TLC.

SanDisk projects KV cache will account for 35% of AI data center NAND workloads by 2030

HBF: the architecture taking aim at HBM

Even more notable than the QLC itself is HBF (High Bandwidth Flash), which SanDisk unveiled alongside it. HBF delivers HBM-like read bandwidth at 8–16x the capacity, targeting memory-heavy AI inference workloads. SanDisk and SK hynix have already published a first open HBF specification: up to 512GB per HBF stack, roughly 4TB per accelerator with 8 stacks, and bandwidth tiers from 0.4TB/s to 3TB/s.

HBF architecture: HBM-class read bandwidth, 8–16x capacity, aimed at AI inference memory bottlenecks

SanDisk's CTO demoed the implications at its investor day: a high-end GPU normally paired with 192GB of HBM can carry roughly 4TB of near-compute memory with HBF. Running a simulated agentic coding workload on a ~490B-parameter model, the HBM system needed at least 8 GPUs just to fit the model — while the HBF system delivered the same output with 4 GPUs. This is the "GPU memory tax" in action: many buyers add GPUs not because they lack compute, but because the model won't fit.

The logic behind the $94 billion repricing

On the demand side, SanDisk projects KV cache will account for 35% of AI data center NAND workloads by 2030 — the longer AI agents live and the longer their context windows, the bigger the KV cache grows. On the supply side, SanDisk and Kioxia used just 13% of industry NAND capex between 2021 and 2025 while producing 29% of total NAND output — a 223% output-to-capex ratio that far outpaces peers. As AI memory demand spills over from HBM into NAND, that efficiency record positions SanDisk as one of the biggest beneficiaries of the repricing.

SanDisk and Kioxia used just 13% of industry NAND capex (2021–2025) to produce 29% of NAND output

Risks remain: HBF needs new system designs and ecosystem support, and QLC's write endurance still depends on smarter controller management. But the direction is clear — HBM is no longer the only star in AI memory. The flash era is just beginning.

FAQ

Q1: What's the difference between QLC and TLC, and why does AI infrastructure need QLC now?

TLC stores 3 bits per cell, QLC stores 4. QLC offers higher density and lower cost per bit, but with slower speeds and lower endurance. Surging demand for model weights and KV cache capacity makes "large and cheap" storage attractive, and new architectures close the bandwidth gap.

Q2: What triggered SanDisk's 5.73% surge on August 13?

SanDisk and Kioxia jointly unveiled 9th-gen 2Tb QLC 3D flash and, on the same day, laid out the HBF architecture and the AI memory repricing story. The market read it as SanDisk transforming from a legacy NAND maker into a key AI memory supplier.

Q3: Can HBF replace HBM?

Not entirely, at least in the near term. SanDisk envisions several deployment modes: replacing some HBM stacks, sitting alongside HBM as high-capacity read-optimized memory, or using HBM as cache while HBF holds model weights and KV cache. The two are complementary.

Q4: What is a KV cache and why does it consume so much memory?

A KV cache is the buffer that stores the keys and values of already-processed tokens so an LLM doesn't recompute them during inference. Longer conversations and longer-lived agentic AI mean bigger KV caches — a primary driver of AI inference memory demand.

Q5: How should investors evaluate memory cycle stocks?

Price-to-earnings alone is misleading. Look at inventory cycles, bit shipment growth, capex discipline, and output-to-capex efficiency. SanDisk's 223% output-to-capex ratio is a concrete example of an "efficiency moat."

留言