DeepSeek replaced flat per token rates with peak/off peak pricing on August 16, 2026, increasing V4 Pro output from $0.87 to $3.96 per million tokens at peak — a 355% rise.
Research answer

Create a landscape editorial hero image for this Studio Global article: What pricing changes did DeepSeek implement for its V4 API models on August 16, 2026, how does the new peak and off-peak tiered structure co. Article summary: **Pi Agent** led in both success rate (20/30) and cost efficiency ($0.028/task).. Topic tags: general, news, general web, user generated. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fake numbers, clickbait thumbnails, icons, and tiny thumbnail layouts. Make it useful as an illustrative visual, not as factual evidence
On August 16, 2026, DeepSeek introduced one of the most significant pricing changes in the AI API market, replacing flat per-token rates with a peak/off-peak structure that pushed costs up by as much as 1,555% for certain token types . The same week, independent testing firm Composio revealed that DeepSeek's top-ranked V4 Flash model completed only 53.8% of real-world agent tasks across eight different orchestration harnesses, exposing a wide gap between leaderboard scores and production reliability
. Together, these developments force enterprises to reconsider the total cost and orchestration maturity required to deploy DeepSeek V4 at scale.
Starting at 16:00 UTC on August 16, 2026, DeepSeek switched from a single flat rate to a two-tier pricing model for both V4-Flash and V4-Pro . Peak hours are 01:00–04:00 UTC and 06:00–10:00 UTC (7 hours total per day); all other hours are off-peak
. Off-peak rates are exactly half of peak rates, but every token category still saw a net increase versus the previous flat pricing
.
| Model | Token type | Previous flat rate | Off-peak (new) | Peak (new) | Effective change |
|---|---|---|---|---|---|
| V4-Flash | Input (cache miss) | $0.14 | $0.21 | $0.42 | +50% off-peak, +200% peak |
| V4-Flash | Input (cache hit) | $0.0028 | $0.007 | $0.014 | +150% off-peak, +400% peak |
| V4-Flash | Output | $0.28 | $0.42 | $0.84 | +50% off-peak, +200% peak |
| V4-Pro | Input (cache miss) | $0.435 | $0.99 | $1.98 | +128% off-peak, +355% peak |
| V4-Pro | Input (cache hit) | $0.003625 | $0.03 | $0.06 | +728% off-peak, +1,555% peak |
| V4-Pro | Output | $0.87 | $1.98 | $3.96 | +128% off-peak, +355% peak |
Source pricing data from DeepSeek's published rate card and independent trackers .
The most extreme increase is V4-Pro cache-hit input, which went from $0.003625 to $0.06 at peak — a roughly 1,555% increase . DeepSeek said the change was made to allocate resources more efficiently and encourage users to schedule non-urgent workloads during off-peak windows
.
In August 2026, independent testing firm Composio evaluated DeepSeek V4 Flash across eight different agent orchestration harnesses on 30 deliberately difficult, multi-step workflows spanning live tools including Gmail, GitHub, Slack, and Google Sheets . Each task had a 900-second timeout, and all harnesses used the same hosted Composio MCP tools
.
The overall result: across 240 total runs, only 129 succeeded — a 53.8% overall pass rate. Only 6 of the 30 workflows were completed successfully by every harness .
| Harness | Pass rate (out of 30 tasks) | Cost per successful task |
|---|---|---|
| Pi Agent | 20/30 (66.7%) | $0.028 |
| Prime Agent | ~15/24 (62.5% valid runs) | $0.131 |
| OMP (Oh My Pi) | 17/30 (56.7%) | $0.103 |
| Claude Code | 16/30 (53.3%) | $0.195 |
| Codex | 16/30 (53.3%) | $0.081 |
| Deep Agents (LangChain) | 16/30 (53.3%) | $0.045 |
| Hermes Agent | 15/30 (50.0%) | $0.056 |
| OpenCode | 14/30 (46.7%) | $0.067 |
*Data from Composio's published results .
Pi Agent led in both success rate (20/30) and cost efficiency ($0.028/task) . A different harness won on each metric — success rate, cost, and speed — meaning orchestration choice matters as much as model capability
. The 20-point gap between the best and worst harness demonstrates extreme sensitivity to the agent framework, tool configuration, caching, retries, and provider stack
.
The combination of higher prices and uneven real-world results has reshaped analyst views on DeepSeek V4's enterprise fit.
VentureBeat noted that the same V4 Flash model produced substantially different results depending on the harness, tooling, and caching stack — suggesting enterprises must invest in orchestration maturity alongside model selection . Composio's own commentary pointed out that harness choice alone changed success by three tasks, cost by almost 3x, and speed by 2.2x
.
Even off-peak, every token type is more expensive than before. Analysts say this changes the ROI calculus for high-volume agentic workloads, particularly for customer support agents and code-generation pipelines that previously relied on DeepSeek's aggressively low pricing .
DeepSeek's official agent benchmarks (82.7 on Terminal-Bench 2.1, 70.3 on Toolathlon-Verified) look strong, but the 53.8% real-world pass rate is a reminder that leaderboard scores do not guarantee reliability in production environments with live APIs, rate limits, and stateful workflows .
API traffic routes through servers in China, creating compliance challenges for regulated enterprises. Analysts advise careful architectural planning for data governance, audit trails, and tool-permission controls . For regulated industries, self-hosting the open weights remains the only viable path to production
.
BayTech Consulting's enterprise framework recommends DeepSeek V4 for non-sensitive document summarization, open-repository code generation, and high-volume content drafting — but calls for rigorous workload-by-workload evaluation before committing to production .
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
DeepSeek replaced flat per token rates with peak/off peak pricing on August 16, 2026, increasing V4 Pro output from $0.87 to $3.96 per million tokens at peak — a 355% rise.