Silicon Data’s index reached 97 cents per million tokens on August 31, 2026, a record low. DeepSeek’s permanent 75% V4 Pro cut and Xiaomi’s cuts of up to 99% helped turn Chinese model competition into a broader inference price war.
Research answer

Create a landscape editorial hero image for this Studio Global article: What caused the Silicon Data LLM Token Expenditure Index to fall below $1 per million tokens to a record low, how did Chinese AI laboratorie. Article summary: The index fell below $1 because the effective market price of inference was pushed down by cheaper open-weight Chinese models, aggressive API discounting, lower production costs, and competitive cuts by frontier provider. Topic tags: general, general web, user generated, news. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts w
The Silicon Data LLM Token Expenditure Index fell below $1 per million tokens because the market shifted toward cheaper models while providers repeatedly cut API prices. Its broad-market reading reached 97 cents on August 31, 2026—the lowest since the index was created and less than half its earlier-summer peak. 4
That headline needs a careful reading. The index is a usage-weighted measure of what the market pays for one million inference tokens, not a tally of total AI spending or token consumption. A falling index therefore shows that intelligence is becoming cheaper on a unit basis; it does not, by itself, show that demand is falling. 2
7
8
Several forces pushed the effective price downward at the same time:
The reported path was broadly downward: the index was above $2 per million tokens in early June, reached roughly $1.16 to $1.18 in early August, and later crossed below $1. 1
38
4 Exact interim readings are less certain because public snapshots and intermediary reports do not always match; one public reading cited for July 27 was $1.43.
1
2
The most visible trigger was DeepSeek’s decision in May to make a 75% reduction to its V4-Pro API pricing permanent. The company said the new rates would remain at one quarter of the original level, with prices varying by usage type. 33
Other Chinese providers followed. Xiaomi announced API cuts of up to 99% for its MiMo-V2.5 models, while reports described simultaneous reductions from several major Chinese labs, including ByteDance, Tencent, MiniMax, and Alibaba. 39
46
This strategy changes the competitive question. Instead of asking only which model is most capable, buyers can compare capability per dollar, latency, context handling, reliability, and deployment flexibility. When models become close enough for many routine tasks, a large price difference can redirect usage quickly.
The price war also reached a turning point. In August, DeepSeek raised prices for newer flagship models and introduced peak and off-peak rates, with off-peak pricing set at half the peak rate. That move suggests that providers cannot treat unlimited demand at the lowest possible price as a sustainable business model; they also need to manage GPU utilization and recover infrastructure costs. 35
The apparent contradiction is simple:
Total spending = price per token × number of tokens used.
If the price falls by 90% but usage expands by more than 90%, total spending can still increase. Lower prices make workloads that were previously too expensive more attractive, including automated customer service, coding, document processing, long-context applications, and agentic workflows.
Some analyses report that token prices have fallen by more than 90% since 2023 while total AI spending has roughly doubled. That interpretation is consistent with a market in which cheaper inference is expanding the number of economically viable applications, rather than destroying demand. 19
22
China’s broader token economy offers a demand signal. The National Data Administration reported daily token consumption above 140 trillion in March, more than 40% above the 100 trillion recorded at the end of the previous year. 26
32 China’s major telecom operators have also been investing in AI-computing capacity; China Telecom reported a 95% first-half increase in intelligent-computing revenue, while China Unicom’s computing-infrastructure capital expenditure rose by more than 80% year over year.
17
18
These figures do not prove that every model provider is profitable. They show instead how lower unit prices can support much larger volumes of infrastructure and usage.
The bullish interpretation is that falling inference prices are the beginning of mass adoption. More affordable intelligence could drive enough token volume to support continued spending on cloud capacity, data centers, chips, networking, and electricity.
The skeptical interpretation focuses on the timing of returns. Reports citing Allianz Research put the gap between AI capital expenditure and sales growth at about 46%, compared with roughly 32% during the 2001 telecom excess cycle. 20
22
23
That comparison is a warning, not proof that AI is repeating the telecom bubble. Infrastructure may eventually support much larger workloads, but providers and infrastructure buyers still face three risks:
A new low in the Silicon Data index is best understood as a signal about pricing power and willingness to pay, not as a standalone measure of AI demand. It tells readers that the market is paying less for each unit of inference, partly because usage is moving toward cheaper models and providers are competing aggressively on price. 7
8
14
The key question is what happens next: whether falling prices unlock enough new workloads to expand total revenue, or whether the price war outruns the growth in usage and leaves the industry with expensive infrastructure but weak returns. The supplied evidence supports the first possibility as a demand story and the second as a material financial risk; it does not settle the outcome.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
Silicon Data’s index reached 97 cents per million tokens on August 31, 2026, a record low.
Silicon Data’s index reached 97 cents per million tokens on August 31, 2026, a record low. DeepSeek’s permanent 75% V4 Pro cut and Xiaomi’s cuts of up to 99% helped turn Chinese model competition into a broader inference price war.
Cheaper inference can expand usage faster than prices fall, but the reported 46% gap between AI capital expenditure and sales growth shows that monetization and returns remain major risks.