Silicon Data’s blended LLM token price index reached a record low $0.97 per million tokens on September 1, 2026—less than half its earlier summer high—as low cost Chinese models reshaped the mix of inference demand. DeepSeek V4 Flash’s listed $0.14 per million input token rate illustrates the new low cost floor, whi...
Research answer

Create a landscape editorial hero image for this Studio Global article: How has the Chinese AI-lab price war—marked by June 2026 international API cuts of 50–99% by DeepSeek, Alibaba Cloud, ByteDance, Moonshot AI. Article summary: China’s price war appears to have reset the marginal market price of LLM inference rather than reduced demand. Aggressive low-cost Chinese APIs pulled customers toward cheaper models and altered the usage mix, helping dr. Topic tags: general, general web, user generated, education, news. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermark
Chinese AI labs’ aggressive API pricing has helped reset expectations for what routine LLM inference should cost. Silicon Data’s LLM Token Expenditure Index reached $0.97 per million tokens on September 1, 2026, a record low and less than half its earlier-summer peak. The move does not mean AI demand has collapsed: the index is a usage-weighted measure of the effective price paid per million tokens, rather than a count of all tokens consumed or all AI dollars spent. 2
4
Silicon Data describes SDLLMTK as a benchmark for blended LLM inference token prices. It combines provider prices with observed consumption across a defined model universe, producing an effective market price per million tokens. 2
4
That distinction matters. If users route a larger share of workloads to lower-cost models, the index can fall even when total token volume—and the total amount organizations spend on AI—rises. In other words, it captures a changing price and workload mix, not aggregate market revenue.
The price war supplied unusually inexpensive alternatives for workloads that do not require the most expensive frontier model. DeepSeek V4-Flash, for example, was listed at $0.14 per million input tokens and $0.28 per million output tokens. 49
DeepSeek also cut V4-Pro pricing by 75% in May 2026, a step that reporting framed as an escalation of competition among Chinese model providers. 58 Other pricing moves and low-cost offerings across the Chinese ecosystem increased the pressure on providers competing for high-volume inference workloads.
4
52
The evidence supports a strong association between this competitive environment and falling blended token prices. It does not, however, isolate a precise causal share for any one lab, model, or announced cut. Model efficiency improvements, capacity expansion, and customers’ own routing decisions also affect the index.
The apparent contradiction disappears with a simple relationship:
total token outlay = effective price per token × token volume
A lower unit price makes more use cases economical. Companies can run more automated workflows, allow longer prompts and contexts, generate more code, or operate more agentic systems. If token volume grows faster than effective price falls, overall spending rises.
That pattern is consistent with reports that per-token prices have declined more than 90% since 2023 while spending on LLM usage has roughly doubled since late 2025. 41
43 The underlying mechanism is expansion of use, not a failure of the price cuts: cheaper inference broadens the set of tasks worth automating.
Reported daily AI token consumption in China exceeded 140 trillion by March 2026, up from 100 trillion at the end of 2025, according to government projections cited by Reuters. 18
Enterprise and financial-sector disclosures point in the same direction. China Merchants Bank reported that its average daily token throughput increased by more than 78% year over year, while separate reporting put the bank’s daily consumption at about 33 billion tokens as of the end of May. 19
29 Reports also said some Chinese banks recorded much larger increases in average daily token use.
20
These figures should not be read as a direct measure of API revenue or as proof that every workload is economically productive. They do show that lower-cost inference is being converted into much higher usage at scale.
For AI buyers, falling inference prices are beneficial. For model vendors, they create a harder business equation: revenue per token may fall while spending on chips, data centers, training, and model development remains substantial.
The central question is therefore no longer whether customers will consume cheap tokens. The evidence suggests they will. The question is whether higher volumes, premium offerings, and enterprise products can generate durable revenue and measurable productivity gains quickly enough to support continued infrastructure investment. Reporting on the index’s decline has highlighted that risk to provider pricing power. 4
36
The most useful metrics are not any single model’s list price. Watch the relationship among:
The key takeaway is straightforward: China’s low-cost API competition has helped make LLM inference dramatically cheaper and pulled blended market pricing below $1 per million tokens. But lower prices are encouraging more use, so token consumption and aggregate AI expenditure can keep climbing at the same time. 2
4
41
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
Silicon Data’s blended LLM token price index reached a record low $0.97 per million tokens on September 1, 2026—less than half its earlier summer high—as low cost Chinese models reshaped the mix of inference demand.
Silicon Data’s blended LLM token price index reached a record low $0.97 per million tokens on September 1, 2026—less than half its earlier summer high—as low cost Chinese models reshaped the mix of inference demand. DeepSeek V4 Flash’s listed $0.14 per million input token rate illustrates the new low cost floor, while reports of more than 140 trillion daily tokens used in China show how falling unit prices can unlock vastly more...
Cheaper tokens can expand, rather than shrink, AI budgets when organizations deploy more workflows, agents, and generated code workloads faster than per token prices fall.