| Alibaba | Qwen 3.7 Max / Plus / Flash | May–June 2026 | 2.8T-param MoE flagship; Qwen3.7 Max at $1.25/$3.75 per M tokens; Qwen3.6 Flash at $0.19/$1.13 |
| Moonshot AI | Kimi K3 | July 16, 2026 | 2.8T total params, ~50B active, 1M context; promised fully open-source by July 27 |
| Z.ai | Fable 5 rival | Late June 2026 | Unveiled shortly after Anthropic ceased two advanced models at U.S. government request |
The cost gap is extreme. DeepSeek's V4-Flash charges $0.14 per million input tokens — more than 100 times cheaper than Anthropic's Claude Fable 5 to run . Alibaba's Qwen3.6 Flash costs $0.19/M input, while U.S. frontier models like GPT-5.5 and Claude Opus 4.8 cost roughly 7–12x more at comparable capability levels
. Moonshot's K3, a frontier-grade model, offers API pricing of $3/M input and $15/M output, with open weights promised for free self-hosting
.