Alibaba's Qwen 3.7 Max, launched in May 2026, is a 2.8-trillion-parameter mixture-of-experts (MoE) flagship that appeared to match or exceed Anthropic's Fable 5 on benchmarks, for roughly one-sixth the cost . For high-volume use, Alibaba's Qwen3.6 Flash costs just $0.19 per million input tokens .
Zhipu AI's GLM-5.2, released June 13, 2026, uses a 744-billion-parameter MoE architecture (40 billion active) with a 1-million-token context window and was ranked the #1 open-weight model on the Artificial Analysis Intelligence Index . Priced at $1.40 per million input tokens and $4.40 per million output, it undercuts U.S. frontier models by roughly 5-6x .
Moonshot AI's Kimi K3, unveiled July 16, 2026, is the largest open-weight AI system ever announced: 2.8 trillion total parameters, a 1-million-token context window, and native visual understanding . Moonshot promised full open weights by July 27, 2026 . Its API pricing sits at $3 per million input tokens and $15 per million output, making it the most expensive Chinese model—yet still below comparable U.S. offerings .
The cost gap between Chinese and U.S. frontier models is extreme:
| Model | Input Price (per 1M tokens) | Output Price (per 1M tokens) |
|---|---|---|
| DeepSeek V4-Flash | $0.14 | $0.28 |
| Alibaba Qwen3.6 Flash | $0.19 | $1.13 |
| Zhipu GLM-5.2 | $1.40 | $4.40 |
| Alibaba Qwen3.7 Max | $1.25 | $3.75 |
| Moonshot Kimi K3 | $3.00 | $15.00 |
| Anthropic Claude Fable 5 (est.) | ~$15.00 | ~$50.00 |
| OpenAI GPT-5.5 (est.) | ~$5.00 | ~$15.00 |
Sources: Citing previous table and
The numbers tell the story: DeepSeek's V4-Flash costs an estimated 3 cents per benchmark test, versus $1.86 for GPT-5.5 and $3.33 for Claude Fable 5 . Venture capitalist Marc Andreessen noted that GLM-5.2 was "the first Chinese AI model to match and often beat the American big lab public AI models with no compromises" — at a fraction of the price .
The second prong of China's strategy is open-weight release. Models like GLM-5.2 (MIT license), Kimi K3 (promised Modified MIT license), and the Qwen series are freely downloadable, meaning anyone can run, tweak, or fine-tune them for no per-token cost . This "giving models away" strategy hollows out the mid-market where U.S. firms used to charge premium prices .
A key turning point came in June 2026, when the Trump administration asked Anthropic to cease operations of its two most advanced AI systems (Fable and Mythos) on national security grounds . Within two weeks, Z.ai unveiled a model that closely rivaled Anthropic's offerings — more affordable and with no U.S. restrictions . Chinese companies saw an opening and rushed to fill it .
The U.S. response has been fragmented. Anthropic complied with the government request. Google is reportedly struggling with the latest version of its Gemini model as Chinese models close the gap . No major coordinated U.S. policy countermeasure has been announced; analysts note that export controls on chips have not stopped Chinese firms from achieving competitive performance through architectural efficiency (e.g., MoE, hybrid attention) rather than brute compute .
Meanwhile, U.S. companies are increasingly testing and adopting Chinese models for cost reasons, further eroding the home-field advantage of American AI labs .