Huawei’s OceanStor M900 extends SuperPoD KV caching from chip memory and DRAM onto shared SSD storage, with a claimed 64 PB of capacity per cluster. Huawei reports roughly 60 microsecond access latency and 40 TB/s of aggregate cluster bandwidth; neither figure describes the speed of every cache lookup.
Published byEdited with GPT-6 SolImages generated with GPT Image 2
Research answer

Create a landscape editorial hero image for this Studio Global article: What is Huawei’s OceanStor M900 Context Memory Storage, introduced at HUAWEI CONNECT 2026 for AI inference SuperPoDs, and how does its Unifi. Article summary: Huawei’s OceanStor M900 is a storage system for SuperPoD AI inference, introduced at HUAWEI CONNECT 2026. It uses UnifiedBus to make the key–value (KV) cache—a reusable record of attention calculations—available across a. Topic tags: general, general web. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fake numbers, clic
Introduced at HUAWEI CONNECT 2026, OceanStor M900 is Huawei’s context-memory storage system for SuperPoD AI inference. Its job is to keep more reusable inference context available across accelerators as prompts, conversations and agent workflows grow. 1
9
A key–value (KV) cache holds information computed while a model processes a prompt. Reusing a suitable cached prefix can spare an inference system from processing that prefix again. Long contexts and repeated agent interactions make that cache more valuable—but also harder to fit entirely in memory close to each neural processing unit (NPU). Huawei positions the M900 as a way to expand and share that cache across a SuperPoD. 1
3
6
The M900 uses Huawei’s UnifiedBus network to pool KV-cache data across on-chip memory, DRAM and SSDs. Huawei says a single cluster can provide 64 PB of KV-cache capacity, moving the amount accessible to an NPU from gigabytes into the terabyte range. That is shared cluster storage, not 64 PB of memory inside an NPU. More retained context could improve reuse when requests have cacheable prefixes, but the benefit depends on what the workload actually reuses. 3
6
17
To make the SSD tier practical for inference, Huawei describes a one-hop path between a SuperPoD NPU and SSDs. Its design integrates CPU, network-controller and NAND-controller functions to avoid some conventional protocol-conversion and CPU-forwarding steps. Huawei reports access latency of roughly 60 microseconds and 40 TB/s of aggregate bandwidth per cluster. Those are stated architecture and system figures, not a guarantee for every request. 3
6
28
32
If a request can reuse cached work, it may reach its first generated token sooner; less repeated computation could also free capacity for more tokens overall. Huawei claims that, in a typical AI coding scenario, the architecture doubled inference-cluster token throughput and halved time to first token. Those workload-specific claims should not be read as universal performance gains. 28
33
Huawei also describes KV-aware placement that manages how long cached data stays on different storage media. The company cites SSD endurance of up to 24 drive writes per day and a 16-fold lifespan improvement, with a three-year stability target. The intended economic case is less recomputation and fewer storage replacements, lowering inference cost per token. The provided evidence does not independently establish those performance, lifetime or cost outcomes in production SuperPoDs. 6
8
33
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
Huawei’s OceanStor M900 extends SuperPoD KV caching from chip memory and DRAM onto shared SSD storage, with a claimed 64 PB of capacity per cluster.
Huawei’s OceanStor M900 extends SuperPoD KV caching from chip memory and DRAM onto shared SSD storage, with a claimed 64 PB of capacity per cluster. Huawei reports roughly 60 microsecond access latency and 40 TB/s of aggregate cluster bandwidth; neither figure describes the speed of every cache lookup.