China's top AI developers — including DeepSeek, Kimi K3, and Alibaba — still train their most advanced models on Nvidia chips, held in place by CUDA software lock in [3][5]. The core barrier is Nvidia's CUDA ecosystem — 15+ years of mature libraries, debugging tools, and global developer community — versus Huawei's...
Research answer

Create a landscape editorial hero image for this Studio Global article: Why do China's leading AI developers still rely on Nvidia chips to train their most advanced models, and what specific barriers — including. Article summary: China's leading AI developers — including those behind DeepSeek, Kimi K3, and Alibaba's models — remain heavily reliant on Nvidia chips for training their most advanced models, despite government pressure and export rest. Topic tags: general, education, news, general web. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fa
China's leading artificial intelligence developers — including the teams behind DeepSeek, Kimi K3, and Alibaba's Qwen models — remain heavily reliant on Nvidia chips for training their most advanced large language models, despite intense government pressure and US export restrictions . The core barrier is not hardware performance but software ecosystem lock-in: Nvidia's CUDA platform has become the de facto standard for AI development, and migrating to Huawei's Ascend chips with the CANN toolkit imposes prohibitively high engineering costs
.
Nvidia's CUDA (Compute Unified Device Architecture) enjoys a 15+ year head start with mature libraries (cuDNN, TensorRT), debugging tools, and a global developer community . Almost all major AI frameworks — PyTorch, TensorFlow, JAX — are optimized for CUDA out of the box
.
By contrast, Huawei's CANN (Compute Architecture for Neural Networks) toolkit, even after being open-sourced in August 2025, remains far less mature in terms of operator coverage, optimization, and community tooling . Developers report that many CUDA-optimized kernels have no direct CANN equivalent, requiring manual rewriting
.
This self-reinforcing cycle is difficult to break: the larger Nvidia's developer community and the richer its library of pre-optimized code, the harder it becomes for alternatives to gain traction .
An AI researcher at a Shanghai-affiliated institute estimates that switching training from Nvidia CUDA to Huawei Ascend would add at least 50% in both time and costs . Industry sources confirm this figure as a realistic floor
.
This includes:
Perhaps the most telling example of the software gap comes from DeepSeek, one of China's most prominent AI labs. DeepSeek reportedly dropped plans to train its R2 model on Huawei's Ascend NPUs, citing unresolved performance and toolchain issues . This is a concrete case of a top-tier lab concluding that the domestic alternative was not ready for their most demanding workloads.
While Huawei's Ascend 910C NPU has improved significantly, developers report it still underperforms Nvidia's Hopper-generation GPUs on large-scale distributed training tasks . The Ascend 910C, built on SMIC's enhanced 7nm process, delivers about one-third of the BF16 throughput of Nvidia's B200 . Huawei's upcoming Ascend 950 is expected to narrow this gap further, with Morgan Stanley data showing the performance gap between Chinese chips and Nvidia's export-compliant hardware has narrowed significantly
.
Paradoxically, US export restrictions have not fully cut off Nvidia supply to China. The US approved the H200 for sale to China, and Beijing subsequently blocked its acceptance — but Chinese labs already hold massive stockpiles of Nvidia chips . ByteDance, Alibaba, and Tencent collectively spent an estimated $16 billion to stockpile roughly 1.3 to 1.6 million H20 units in anticipation of a ban
.
This installed base creates powerful inertia: as long as existing Nvidia hardware continues to work, the high switching cost delays the move to domestic alternatives .
China is not standing still. In response to these barriers:
Progress is real: domestic chip training share reportedly grew from 8% in early 2025 to 42% by April 2026 . But much of this adoption appears concentrated in less critical workloads and inference tasks rather than frontier model training
.
Huawei's open-sourcing of CANN is a meaningful step toward building the developer ecosystem that CUDA enjoys . But most analysts agree the software ecosystem gap — not raw hardware specs — will take years to close
. In the short term, China's top AI labs will continue training on Nvidia chips while experimenting with hybrid domestic deployments for less critical workloads, and while Beijing's $295 billion grid plan aims to accelerate the transition
.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
China's top AI developers — including DeepSeek, Kimi K3, and Alibaba — still train their most advanced models on Nvidia chips, held in place by CUDA software lock in [3][5].
China's top AI developers — including DeepSeek, Kimi K3, and Alibaba — still train their most advanced models on Nvidia chips, held in place by CUDA software lock in [3][5]. The core barrier is Nvidia's CUDA ecosystem — 15+ years of mature libraries, debugging tools, and global developer community — versus Huawei's less mature CANN toolkit, which requires extensive code rewriting [1][2][6].
China's chipmakers aim to triple domestic AI chip output in 2026, and Beijing is drafting a $295 billion plan to link data centers into a single AI grid by 2028 [2][8].