At AICC2026, Inspur announced a 128 chip MetaBrain SD200 Ultra that it says can run 2.8 trillion parameter Kimi K3 at under 5.85 ms per generated token. The companion HC2000 targets token production capacity rather than the SD200 Ultra’s single user latency; its tenfold output claim also needs a disclosed baseline.
Published byEdited with GPT-6 SolImages generated with GPT Image 2
Research answer

Create a landscape editorial hero image for this Studio Global article: What did Inspur Information announce about the MetaBrain SD200 Ultra at AICC2026, including its intended workloads, 128-chip and memory conf. Article summary: Inspur Information announced the MetaBrain SD200 Ultra at AICC2026 as a tightly coupled supernode for large frontier models and latency-sensitive AI-agent workloads. It also introduced the MetaBrain HC2000 as a separate,. Topic tags: general, general web, user generated. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fa
Inspur Information unveiled the MetaBrain SD200 Ultra at AICC2026 as a supernode for very large models and latency-sensitive AI-agent workloads. Alongside it, the company introduced the MetaBrain HC2000 for a different goal: producing more inference tokens across matched workloads. The headline performance numbers for both systems come from Inspur and should be evaluated as claims, not independent benchmarks. 3
18
20
The SD200 Ultra uses a 3D Hyper Mesh architecture to tightly couple 128 domestically made AI accelerators. Inspur specifies 8 TB of unified-address accelerator memory and 64 TB of system memory. It says a single machine can run the 2.8-trillion-parameter Kimi K3 model and support frontier models of up to 10 trillion parameters. The larger figure describes claimed model capacity; it does not establish inference speed at that size. 3
17
For Kimi K3, Inspur reports less than 5.85 milliseconds per generated token, which it equates to about 170 tokens per second for one user and calls five times an industry average. The sources do not provide enough test detail to treat that comparison as a like-for-like benchmark. Per-token generation latency also differs from first-token latency or the time a user waits for a complete response. 1
3
Inspur says global unified addressing and memory-semantic communication reduce data movement between chips. Its symmetric-memory approach lets an accelerator directly access another accelerator’s memory; the company reports communication latency as low as 0.69 microseconds and a 3.5-fold reduction in AllReduce communication time. Those are communication measurements, not application-level response times. 2
17
For Kimi K3 specifically, Inspur says it fuses operations associated with KDA, Gated MLA and MoE into larger compute-and-communication operators. It claims this optimization cuts the operator count by about tenfold and improves inference performance by more than threefold. The improvement should be read as a workload-specific vendor claim, not a speedup guaranteed for other models. 17
19
The HC2000 is Inspur’s capacity-inference offering. It combines a liquid-cooled rack, DirectCom 2.0 architecture and MCIS inference software; Inspur claims ten times the token output for the same investment. That is a throughput-and-economics pitch, not a claim that HC2000 matches the SD200 Ultra’s reported single-user token latency. The comparison needs a defined baseline, workload and cost boundary. 18
20
Published accounts describe the SD200 Ultra’s accelerators as domestic but do not identify their supplier. Before purchasing either system, buyers should request the chip identity and software-support details, then run reproducible tests on their own models. Those tests should specify precision, context length and concurrency, and measure first-token latency, sustained token throughput, power use and total cost. Without those conditions, the launch figures cannot reliably predict performance in a buyer’s deployment. 1
3
18
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
At AICC2026, Inspur announced a 128 chip MetaBrain SD200 Ultra that it says can run 2.8 trillion parameter Kimi K3 at under 5.85 ms per generated token.
At AICC2026, Inspur announced a 128 chip MetaBrain SD200 Ultra that it says can run 2.8 trillion parameter Kimi K3 at under 5.85 ms per generated token. The companion HC2000 targets token production capacity rather than the SD200 Ultra’s single user latency; its tenfold output claim also needs a disclosed baseline.