This is why a surge in total tokens can understate the shortage of premium inference. The market may be producing more tokens overall while experiencing a sharper shortage of the hardware needed to produce reliable results for complex tasks. Investment in high-quality token capacity and systems that combine heterogeneous computing resources reflects that distinction.
China’s average daily token calls exceeded 140 trillion in March 2026, compared with about 100 billion at the beginning of 2024—a rise of more than 1,000 times according to figures reported from China’s National Data Administration.
The increase signals a change in how AI is being used. Models are moving beyond experimentation and isolated prompts into assistants, enterprise software and agents that perform real-world workflows. Inference is consequently becoming a major driver of computing consumption rather than a secondary step after training.
Adding high-end processors is difficult when export controls restrict access to advanced Nvidia products, supply is limited and hardware ecosystems are expensive to replace. China’s AI developers also face substantial transition costs because existing models, tools and workflows are deeply tied to established software platforms.
Software optimization cannot create new silicon, but it can increase the amount of useful work extracted from each available processor. The practical approaches include:
These measures are a bridge between rising demand and constrained supply. They are not a complete substitute for frontier hardware, especially when quality-sensitive workloads remain dependent on a narrower set of processors.
The shortage is harder to solve because inference prices are under pressure. Jefferies analysts estimated that Chinese models cost, on average, about one-sixth as much per token as offerings from major U.S. companies.
Low prices can accelerate adoption, but they also limit how much revenue is available to fund scarce premium compute. At the same time, coding assistants and agentic products can consume more resources than simple chat. Chinese AI companies therefore face a difficult combination: demand is expanding rapidly, customers expect inexpensive access, and the most capable inference capacity cannot be expanded or replaced easily.
That tension has already appeared in service availability. A report on the market described major Chinese model developers and cloud providers throttling services, rationing access or raising prices as demand exceeded available computing resources, particularly for resource-intensive coding assistants.
China’s AI infrastructure challenge is not simply a shortage of chips. It is a shortage of the right combination of advanced processors, compatible software and economical capacity for high-value inference.
Domestic hardware can absorb more of the workload as software improves and deployments are designed around its strengths. But the current evidence points to a transitional strategy: optimize every layer of inference, use domestic chips where they are effective and preserve scarce Nvidia capacity for the hardest jobs. That can extend existing resources, but sustained growth in agent-based AI will ultimately depend on whether China can expand both its hardware supply and the software ecosystem that makes different processors practical to use.