On September 30, 2026, DeepSeek released Ascend versions of six infrastructure components. The reported inference setup used an Ascend 950 UBL128 system, EP32 and a 128K context; figures were reported with PoC firmware and excluded serving and framework overhead.
Published byEdited with GPT-6 LunaImages generated with GPT Image 2
Research answer

Create a landscape editorial hero image for this Studio Global article: What did DeepSeek open-source for Huawei’s Ascend platform on September 30, 2026, how do its Ascend versions of TileLang, DeepGEMM, TileKern. Article summary: DeepSeek open-sourced an Ascend counterpart to its NVIDIA-oriented software stack on September 30, 2026. The releases establish corresponding programming, compute, attention, selection and communication components, but t. Topic tags: general, general web, user generated, news. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts w
DeepSeek’s September 30, 2026 release added Ascend versions of six components spanning kernel programming, compute, attention, data selection and distributed communication. The software provides building blocks for running AI workloads on Huawei hardware, but the available evidence does not show a like-for-like performance match with NVIDIA—or confirm that DeepSeek completed full-scale V4 training on Ascend. 2
5
The stack pairs components with roles in DeepSeek’s NVIDIA-oriented software, while using Ascend-specific implementations where needed. 2
5
That alignment is useful for developers, but API or component correspondence is not proof that the two platforms have identical internals, full feature parity or comparable speed.
A reported DeepEP test on Ascend hardware measured 375 GB/s for dispatch and 347 GB/s for combine. Those figures describe Ascend communication performance; they are not a head-to-head comparison with NVIDIA. 19
47
For inference, a reported DeepSeek-V4.1-Flash test used an Ascend 950 UBL128 system, an EP32 deployment strategy, a 128K context and offline inference. At a 5 ms time per output token (TPOT) setting, it reported 2,469 output tokens per second per card; at 10 ms TPOT, it reported 5,102 tokens per second per card. The reported figures excluded serving-scheduling and framework load-balancing overhead. 47
These are configuration-specific results, not a general throughput guarantee. The DeepEP repository describes its measurements as using a manually configured proof-of-concept hardware development package; another report says the figures used PoC firmware and pointed to a commercial release planned for around October 15. That means the firmware and setup matter when interpreting or trying to reproduce the numbers. 12
19
The supplied reporting does not provide enough detail to independently reproduce every inference setting, and it describes no large-scale independent third-party validation. The results therefore should not be treated as a verified comparison against NVIDIA systems. 47
Huawei and DeepSeek described a jointly defined Ascend 950 SuperPoD Flex / UBL128 design. Reports give the scale-up configuration as 128 cards with a 3.2 Tbps single-stage network, and describe a two-stage scale-out design intended to extend to 256,000 cards. That is a stated design capability, not evidence that a system of that size ran a DeepSeek training job. 10
47
49
Huawei also said joint work would be shared through the CANN community, with reported coverage including model deployment, large-scale training, long-context KV-cache pooling and agentic reinforcement learning. 14
The release establishes a broader Ascend software stack, and the reported inference run shows a specific offline configuration. But neither component correspondence nor an advertised scale-out design confirms a full-scale V4 training run on Ascend. The cited benchmark figures also do not show NVIDIA-equivalent performance: they use different hardware, and the available reports do not provide a matched cross-platform test. 2
5
19
47
The clearest reading is that DeepSeek has published meaningful Ascend software and Huawei has described a system architecture and selected test results. The extent to which those results translate to other workloads—or compare with NVIDIA in a controlled test—remains unresolved.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
On September 30, 2026, DeepSeek released Ascend versions of six infrastructure components.
On September 30, 2026, DeepSeek released Ascend versions of six infrastructure components. The reported inference setup used an Ascend 950 UBL128 system, EP32 and a 128K context; figures were reported with PoC firmware and excluded serving and framework overhead.