| Variant | Input (off-peak) | Output (off-peak) | Cache Hit |
|---|---|---|---|
| V4-Pro | $0.435 | $0.87 | 90% discount |
| V4-Flash | $0.14 | $0.28 | 90% discount |
DeepSeek also ran a 75% discount promotion on V4-Pro through early May 2026 . For comparison, GPT-5.5 output pricing is ~$30/MTok, making V4-Flash roughly 107× cheaper on output
.
DeepSeek V4's benchmark results tell a nuanced story:
Important caveat from NIST CAISI evaluation (May 2026): DeepSeek's self-reported benchmarks claim parity with GPT-5.4 and Opus 4.6, but CAISI's independent evaluation found V4 "substantially less capable" on non-public benchmarks, suggesting some self-reported scores may be inflated . Artificial Analysis's Intelligence Index rates GPT-5.5 (high) at 53 vs. DeepSeek V4 Pro (Reasoning, High Effort) at 41* — GPT-5.5 is more intelligent and faster (69 vs. 56 tok/s), but DeepSeek is dramatically cheaper ($0.18 vs. $4.35/MTok under reasoning mode)
.
Alibaba's Qwen team previewed Qwen 3.8 Max at the World AI Conference in Shanghai on July 19, 2026 . It is the team's first flagship model above 1 trillion parameters.
No pay-as-you-go pricing has been published. Access is currently via subscription through Qwen Chat and select API partners. BenchLM lists Qwen 3.8 Max Preview pricing as "Not listed" . Some third-party trackers show auto-pricing at $1.50/$5.00 per million input/output tokens, but this is not confirmed by Alibaba
.
Alibaba states Qwen 3.8 is "second only to Fable 5" (Anthropic's top Claude model), which would place it ahead of GPT-5.5 and DeepSeek V4 in Alibaba's own ranking . Gains are claimed in coding, professional productivity, full-stack development, data analysis, and office workflows over Qwen 3.7 Max
.
However, no independent benchmarks have been published. BenchLM, a benchmark tracking site, tracks 321 benchmarks for Qwen 3.8 Max Preview with 0 verified results as of July 20, 2026 . As one analysis put it, "Qwen 3.8 is worth testing, but it is too early to treat it as a stable production replacement for Kimi K3, GPT-5.6 Sol, or Claude Fable 5"
.
For context, Qwen's predecessors set a strong foundation: Qwen 3.6-Max-Preview (April 20, 2026) was a 35B total / 3B active parameter text-only model with 262K context that claimed #1 on six coding and agent benchmarks . Qwen 3.7-Max (May 2026) expanded to a 1M token context window with reasoning agent capabilities
. But neither was generally considered frontier-level outside of agentic coding.
Alibaba explicitly positions Qwen 3.8 as "second only to Fable 5" . If this claim holds up under independent evaluation, it would place Qwen 3.8 ahead of GPT-5.5 and well ahead of DeepSeek V4. But with zero verified benchmarks, this remains a marketing claim, not an established fact. Anthropic's Fable 5 appears to be the current recognized frontier leader per Alibaba's own framing
.
DeepSeek's aggressive pricing has driven a race to the bottom on API costs. V4-Flash at $0.14/M input is radically cheaper than any Western frontier model. Alibaba has not yet set Qwen 3.8 pricing, but Qwen 3.6 was already priced aggressively (~$1.25/$7.50 input/output) .
The open-weight MIT license for DeepSeek V4 means self-hosting is also possible, further compressing margins for closed-source providers . As one analysis put it, the practical effect is that "developers can now access near-frontier capability (on coding tasks) at a fraction of the cost of GPT-5.5 or Claude."