Global OpenRouter traffic reached 113 trillion tokens in August 24–30, up 21.11% week over week; Chinese models accounted for 55.16 trillion versus 17.07 trillion for U.S. The top three models were reportedly all Chinese: Ox Alpha—later identified as Z.ai’s GLM 5.3 Flash—led with 15.7 trillion tokens, followed by De...
Research answer

Create a landscape editorial hero image for this Studio Global article: What did the National Business Daily’s analysis of OpenRouter data for August 24–30, 2026 reveal about global AI model usage—including the 1. Article summary: National Business Daily’s reading of OpenRouter traffic for August 24–30 portrayed a sharp shift in global inference demand: 113 trillion tokens for the week, up 21.11%, with Chinese models dominating both country-level . Topic tags: general, general web, user generated, news. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts w
A National Business Daily analysis of OpenRouter data for August 24–30, 2026, described a decisive week for China’s AI model ecosystem. Global traffic reached 113 trillion tokens, up 21.11% from the previous week, while Chinese models reportedly led U.S. models in inference volume for an 18th consecutive week. 13
The headline is especially notable because the leading model was not initially marketed under a recognizable company name. The anonymous Ox Alpha—also called “Niu Lai” by some Chinese developers—was later identified by Z.ai, also known as Zhipu AI, as GLM-5.3-Flash. 5
13
The analysis attributed 55.16 trillion tokens to Chinese models, compared with 17.07 trillion for U.S. models. It also described the week as the first time the top three models were all Chinese-developed systems. 13
That distinction matters, but OpenRouter traffic should be read as a measure of usage flowing through a global model-routing platform—not as a direct census of all AI use worldwide. The figures show where traffic on that layer was concentrated during the measured week; they do not, by themselves, explain the users’ locations or every workload behind the tokens.
Ox Alpha reportedly ranked first with 15.7 trillion tokens. DeepSeek-V4-Flash followed with 12.3 trillion, while Xiaomi MiMo-V2.5 remained in third place for a second consecutive week at 9.14 trillion tokens. 13
The model first appeared on OpenRouter on August 20 as an unnamed, free-access listing. Reports described a context window of roughly one million tokens and support for text, image and video input. Z.ai confirmed the model’s identity on August 26, then released open weights under an MIT license. 2
5
7
The launch strategy helped turn a model mystery into a large-scale real-world test. Developers could try the system without the normal price barrier, while its long context and multimodal input made it relevant to coding and extended agent workflows. OpenRouter later described Ox Alpha as its biggest model launch, with more than 20 trillion tokens processed in six days. 15
After the free preview, reported standard pricing for GLM-5.3-Flash was $0.15 per million input tokens and $0.50 per million output tokens. A launch promotion cut those rates to $0.075 and $0.25 through September 9, according to Z.ai and OpenRouter posts. 1
15
The precise comparison with GLM-5.3—described in the original reporting as roughly one-tenth the price, or one-twentieth during the promotion—was not independently established by the supplied evidence. The broader point is better supported: inexpensive access, a free trial and high-capacity routing can produce rapid token growth even before a model has built a long public track record.
The reporting also cited strong DeepSWE results for GLM-5.3-Flash against several named models. Those numbers should be treated as benchmark-specific claims, not as a definitive ranking of general model quality. The available evidence includes conflicting figures for the benchmark, and a coding score cannot measure every capability or production constraint. 4
For developers, the practical lesson is to separate three questions:
OpenRouter’s weekly data speaks most directly to the first question. It suggests that developers were willing to move substantial workloads toward a newly introduced Chinese model when the access conditions were attractive.
The same analysis reportedly placed Gemini 3.7 Flash at 3.95 trillion tokens, with growth of 120%, while GLM 5.2 and DeepSeek-V4-Pro-0423 left the leading group. These specific figures were not independently corroborated in the supplied sources, so they should be viewed as reported rather than settled platform-wide measurements.
Even with that qualification, the overall pattern is clear enough to matter: model rankings can change quickly when new entrants combine competitive capabilities with aggressive pricing, broad availability and an easy developer experience.
The week’s data points to a more contestable inference market. Premium U.S. proprietary models are no longer the only systems capable of attracting large international developer workloads through a shared routing layer. Chinese models can gain traffic by combining:
That does not prove that Chinese models have overtaken every competitor in capability, nor does it show that all of the traffic came from outside China. It does show that distribution and unit economics can change usage patterns rapidly. For teams choosing models, the relevant comparison is increasingly not “which model is best?” but “which model delivers the required result at acceptable cost, latency and risk?”
The August 24–30 snapshot therefore matters less as a permanent leaderboard than as evidence of how quickly inference demand can move. If the pattern persists after promotional access ends, it would provide stronger evidence that efficient Chinese models are winning durable developer adoption rather than simply benefiting from a launch-week spike.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
Global OpenRouter traffic reached 113 trillion tokens in August 24–30, up 21.11% week over week; Chinese models accounted for 55.16 trillion versus 17.07 trillion for U.S.
Global OpenRouter traffic reached 113 trillion tokens in August 24–30, up 21.11% week over week; Chinese models accounted for 55.16 trillion versus 17.07 trillion for U.S. The top three models were reportedly all Chinese: Ox Alpha—later identified as Z.ai’s GLM 5.3 Flash—led with 15.7 trillion tokens, followed by DeepSeek V4 Flash and Xiaomi MiMo V2.5.
The result points to price, free previews, long context and open weight distribution becoming powerful drivers of global developer traffic—not just model benchmark performance.