Huawei says OceanStor M900 extends SuperPoD KV cache into a shared SSD backed tier with up to 64 PB per cluster. Terabytes of cache available per NPU refers to access to pooled, tiered storage—not terabytes of new on chip memory.
Published byEdited with GPT-6 SolImages generated with GPT Image 2
Research answer

Create a landscape editorial hero image for this Studio Global article: How does Huawei’s OceanStor M900 AI memory storage address the KV Cache memory wall in hyperscale, long-context and agentic inference SuperP. Article summary: Huawei positions OceanStor M900 as a shared, SSD-backed **KV-cache tier** for SuperPoD inference, not as a replacement for the NPU’s fast memory. Its aim is to keep and reuse more context across long-running and agentic . Topic tags: general, general web. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fake numbers, clic
Long-context and agentic inference can generate more reusable key-value (KV) cache than a SuperPoD can keep close to its processors. Huawei’s OceanStor M900 addresses that capacity constraint by adding a shared, SSD-backed cache tier beneath on-chip memory and DRAM. The goal is to retain and reuse more context without treating SSDs as a replacement for the NPU’s fastest memory. 8
10
12
Huawei says its Lingqu, or UnifiedBus, interconnect pools KV cache across a SuperPoD and supports tiering between on-chip memory, DRAM and SSDs. It claims up to 64 PB of capacity per cluster and says the KV-cache capacity available to each NPU rises from gigabytes to terabytes. That per-NPU figure describes access to shared, tiered capacity—not an increase in physical on-chip memory. 8
10
This makes M900 an inference-focused context-storage system rather than another compute chip. It complements NPUs and depends on the interconnect for pooled access; it does not replace either. Huawei describes it as designed for agent-heavy, longer-context workloads, while industry reporting identifies Atlas 960 SuperPoDs as a target setting. 12
6
Huawei attributes a one-hop NPU-to-SSD path to a design integrating CPU, network-controller and NAND-controller functions. It says the path avoids CPU forwarding and protocol conversion, cutting access latency by 90% to 60 microseconds. Huawei also claims 40 TB/s of aggregate cluster bandwidth, or 1.5 times its stated comparison solution; that is a cluster-wide figure, not bandwidth guaranteed to each NPU. 8
For a “typical AI programming” workload, Huawei reports double the inference cluster’s token throughput and half the time to first token. Those figures should not be read as a universal speedup for every agentic or long-context request. 8
14
M900’s KV-Aware scheduling predicts cache-data lifetime and manages placement across storage media, according to Huawei. The company claims support for up to 24 drive writes per day, 16-fold longer SSD-media life and three years of stable operation. It argues that fewer replacements and less maintenance can lower inference costs, but the cited material does not establish a measured cost-per-token saving. 8
The central distinction is between more cache capacity within reach and proven faster, cheaper inference at scale. Independent tests would need to establish usable capacity and cache-hit rates on real workloads; latency and sustained bandwidth under contention; token gains against a specified baseline; and SSD wear and total cost over time. The available third-party coverage reports Huawei’s performance figures rather than independently validating them. 6
14
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
Huawei says OceanStor M900 extends SuperPoD KV cache into a shared SSD backed tier with up to 64 PB per cluster.
Huawei says OceanStor M900 extends SuperPoD KV cache into a shared SSD backed tier with up to 64 PB per cluster. Terabytes of cache available per NPU refers to access to pooled, tiered storage—not terabytes of new on chip memory.