On August 13, 2026, DeepSeek announced API price increases of 50%–1,100% for its V4 Pro and V4 Flash models, with new peak/off peak billing taking effect at 16:00 UTC on August 16, 2026 [17][19]. Both V4 models feature a 1M token context window, MIT license, and up to 384K max output tokens.
Research answer

Create a landscape editorial hero image for this Studio Global article: What new API pricing did DeepSeek announce for its V4 models, when does it take effect, and what are the key details about the models' conte. Article summary: On August 13, 2026, DeepSeek announced major API price increases for its V4-Pro and V4-Flash models, with new peak/off-peak billing tiers taking effect at **16:00 UTC on August 16, 2026** [1][4]. The increases range from. Topic tags: general, academic, documentation, general web, news. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks,
On August 13, 2026, DeepSeek announced major API price increases for its V4-Pro and V4-Flash models, introducing peak/off-peak billing tiers that take effect at 16:00 UTC on August 16, 2026 . The increases range from 50% to 1,100% above current rates, depending on the model, token type, and time band
. This marks a significant shift for a company that built its reputation on aggressively low pricing.
DeepSeek is the first major AI API provider to introduce structural time-of-day surge pricing, modeled after electricity peak/off-peak billing . Off-peak rates are set at 50% of peak rates, meaning peak-hour API calls cost double
. The peak windows (UTC) are reportedly 01:00–04:00 and 06:00–10:00
.
| Token Type | Current Price | New Off-Peak | New Peak |
|---|---|---|---|
| Input (cache miss) | $0.435 | ~$0.44–$0.49 (est.) | ~$0.87–$0.98 (est.) |
| Output | $0.87 | $1.98 | $3.96 |
| Token Type | Current Price | New Off-Peak | New Peak |
|---|---|---|---|
| Input (cache miss) | $0.14 | ~$0.21 (est.) | ~$0.42 (est.) |
| Output | $0.28 | $0.63 | $1.26 |
(Input cache-miss prices for the new tiers are estimated from the confirmed output multipliers. Output jumps are confirmed by multiple sources .)
Cached input tokens remain dramatically cheaper. For V4-Pro, cache-hit input stays at approximately $0.004 per 1M tokens, a 100x discount that incentivizes developers to optimize their caching strategies .
Both V4 models are text-only, Mixture-of-Experts (MoE) architectures released under the MIT license, making them fully open-weight and self-hostable .
| Specification | DeepSeek V4-Pro | DeepSeek V4-Flash |
|---|---|---|
| Total parameters | 1.6 trillion | 284 billion |
| Active parameters per token | ~49 billion | ~13 billion |
| Context window | 1,000,000 tokens | 1,000,000 tokens |
| Max output length | 384,000 tokens | 384,000 tokens |
| License | MIT | MIT |
Both models support three reasoning modes: Non-think (fast, intuitive), Think High (logical analysis with chain-of-thought), and Think Max (deep reasoning) .
The new prices reshape DeepSeek's competitive positioning. The comparison below uses standard short-context API rates per 1M tokens for all providers.
| Model | Input (per 1M) | Output (per 1M) |
|---|---|---|
| DeepSeek V4-Pro (peak) | ~$0.87–$0.98 | $3.96 |
| DeepSeek V4-Flash (peak) | ~$0.42 | $1.26 |
| OpenAI GPT-5.6 Luna | $0.20 | $1.20 |
| Anthropic Claude Fable 5 | $10.00 | $50.00 |
(Sources: DeepSeek , OpenAI
, Anthropic
)
OpenAI cut GPT-5.6 Luna prices by 80% on July 30, 2026, to $0.20 input and $1.20 output per 1M tokens . The competitive picture flipped dramatically:
Where DeepSeek was once the clear budget option, it now faces a pricing leader in OpenAI's most affordable tier .
Anthropic's flagship reasoning model is priced at $10 input and $50 output per 1M tokens . At those levels, DeepSeek remains the value play:
Even at their new higher rates, both DeepSeek models remain dramatically more affordable than Anthropic's top-tier offering .
xAI's Grok 4.6 pricing was not available from the provided sources, so a direct comparison cannot be made.
DeepSeek cited soaring demand straining inference capacity as the primary reason for the hike . The V4 models gained massive adoption after their preview launch in April 2026, creating GPU capacity bottlenecks that the company could no longer absorb at its previous ultra-low rates
.
Several market forces converged:
DeepSeek's price increases mark the end of its era as the overwhelmingly cheapest frontier API provider. The company is now closer to parity with OpenAI's most affordable model on some dimensions while still undercutting Anthropic's premium tier by an order of magnitude. For developers, the combination of peak/off-peak pricing, deep caching discounts, and MIT-licensed open weights means that workload architecture — rather than raw model choice — will increasingly determine the final API bill.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
On August 13, 2026, DeepSeek announced API price increases of 50%–1,100% for its V4 Pro and V4 Flash models, with new peak/off peak billing taking effect at 16:00 UTC on August 16, 2026 [17][19].
On August 13, 2026, DeepSeek announced API price increases of 50%–1,100% for its V4 Pro and V4 Flash models, with new peak/off peak billing taking effect at 16:00 UTC on August 16, 2026 [17][19]. Both V4 models feature a 1M token context window, MIT license, and up to 384K max output tokens.
DeepSeek cited soaring demand and GPU capacity bottlenecks as the primary reason for the hike, becoming the first major AI API provider to introduce structural time of day surge pricing [1][22].