In the July 20–26 OpenRouter snapshot, Xiaomi’s MiMo V2.5 reportedly processed about 10.5 trillion weekly tokens and led a Chinese top five. MiMo combined a 1 million token context window and native visual/audio understanding with public API availability, token plan pricing and later MIT licensed open weights for co...
Research answer

Create a landscape editorial hero image for this Studio Global article: How did Xiaomi’s MiMo‑V2.5 become the world’s most-used LLM by token volume by July 2026—reaching 10.5 trillion weekly and 31.2 trillion mon. Article summary: MiMo‑V2.5’s rise appears to be a distribution-and-economics win as much as a model-quality win: Xiaomi paired agent-oriented, native multimodal capability and a 1M-token context window with aggressively accessible API ac. Topic tags: general, general web, user generated. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fa
OpenRouter’s July 2026 leaderboard offered a striking picture of practical LLM demand: Xiaomi’s MiMo-V2.5 reportedly led the platform with roughly 10.5 trillion tokens in the week of July 20–26, while Chinese-developed models occupied the top five positions. MiMo’s rise was likely less about one isolated breakthrough than a combination of capable agent-oriented features, accessible deployment options and economics that fit high-volume workloads. 8
15
The headline is about tokens routed through OpenRouter, a multi-model access platform—not a complete count of every LLM request made globally. The reported July figures put MiMo at about 10.5 trillion weekly tokens and 31.2 trillion monthly tokens; one report calculated growth from 1.46 trillion weekly tokens in May to 10.46 trillion in July as roughly 616%. 15
Those figures are meaningful because they show what developers were choosing on a large routing platform. But tokens are not the same as API calls, users, revenue or business criticality. A model used for long-context processing or multi-step agents can generate far more tokens per task than a model used for short chat prompts.
The frequently cited 63.5% Chinese-versus-35.5% US split is also a platform-specific snapshot, reported for OpenRouter over a 28-day period ending July 26. It should not be described as a census of all global AI inference. 15
Xiaomi positioned MiMo-V2.5 around agentic capability and multimodal understanding. The company says the model supports native visual and audio understanding and up to 1 million tokens of context. 16
That matters for workloads such as software agents, document-heavy research, media analysis and tool-using assistants. These systems may repeatedly read context, plan a next step, call tools, inspect results and revise an answer. They are naturally token intensive.
MiMo also entered public testing on April 23, 2026, and Xiaomi initially offered it free for a limited time. 3 Fast availability can matter as much as a benchmark result: developers can test a model in real applications, change routing rules and scale usage without a long procurement cycle.
When several models meet a team’s quality threshold, the decisive questions become operational: What does a completed task cost? Can the service handle the workload? Is it compatible with an existing stack? Can the organization control deployment?
Xiaomi explicitly marketed token-plan pricing and token efficiency for the MiMo family. 1 More importantly, it later released the V2.5 series under the permissive MIT license, allowing commercial inference deployment, secondary training and fine-tuning without additional authorization.
2
Open weights do not guarantee adoption, but they expand the choices available to customers. A team can use a hosted API, self-host, fine-tune or keep an alternative deployment path. That flexibility is particularly valuable for high-volume applications, where unit economics and vendor dependence can become material.
For the July 20–26 period, reports based on OpenRouter data placed MiMo-V2.5 first, followed by DeepSeek V4 Flash, Tencent HY3, Zhipu GLM-5.2 and DeepSeek V4 Pro. 8 Other July reporting similarly described Chinese-origin models as accounting for more than 60% of routed traffic on the platform.
19
26
The practical takeaway is not that Chinese models had conclusively surpassed every US model on every measure. Rather, the ranking shows that model adoption had become highly contestable at the inference layer. Models that are capable enough, inexpensive to run and easy to integrate can attract enormous production volume even if frontier-model leadership remains debated.
There was also evidence of cross-border demand: CNBC reported that US companies’ share of OpenRouter tokens spent on Chinese models had remained above 30% each week since February 8, reaching as high as 46%. 23 That makes the shift harder to explain as domestic usage alone.
MiMo’s brief lead illustrates a key distinction in AI markets:
The second contest rewards a wider set of attributes: price-performance, context capacity, multimodal inputs, integration options and commercial terms. MiMo brought several of those attributes together, while China’s model ecosystem supplied multiple alternatives competing for the same developer workloads.
That does not erase US strength in closed frontier models, nor does token volume measure safety, reliability, profitability or enterprise importance. It does show that the center of gravity for high-volume routed inference can move quickly when capable models are available on favorable economic and deployment terms.
MiMo-V2.5’s July ranking should be read as a distribution-and-economics milestone. Its reported 10.5 trillion weekly tokens reflected an offering built for agent and multimodal workloads, with a 1 million-token context window, accessible rollout and a later open commercial license. 2
8
16
The Chinese top-five result is therefore best understood as evidence that real-world LLM adoption increasingly rewards deployable, cost-conscious systems—not as proof that any one country has settled the broader race for AI leadership.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
In the July 20–26 OpenRouter snapshot, Xiaomi’s MiMo V2.5 reportedly processed about 10.5 trillion weekly tokens and led a Chinese top five.
In the July 20–26 OpenRouter snapshot, Xiaomi’s MiMo V2.5 reportedly processed about 10.5 trillion weekly tokens and led a Chinese top five. MiMo combined a 1 million token context window and native visual/audio understanding with public API availability, token plan pricing and later MIT licensed open weights for commercial deployment and fine tuning.
Token volume is shaped by pricing, long contexts and agent loops, so it should not be treated as a proxy for revenue, unique users, safety or frontier research leadership.