The increase varies by model, token type, and time of use. Reuters reported that the new rates range from 50% to 1,100% above previous prices.
The largest percentage changes apply to cached-input pricing, which starts from a very small base. For developers, the more practical impact depends on the mix of cache hits, uncached inputs, generated outputs, and the hours when workloads run. V4-Pro output, for example, is about 128% more expensive off-peak than the previous $0.87 promotional rate and about 355% more expensive at peak, based on the listed prices.
DeepSeek remains inexpensive relative to many closed US models even after the increase. Earlier reporting listed V4-Pro output at roughly $0.87 per million tokens under the promotional price, compared with substantially higher historical prices for leading US systems. The price change therefore weakens DeepSeek’s cost advantage without eliminating it.
For a developer generating 10 million V4-Pro output tokens per month:
That means the output-only bill increases by approximately $11.10 to $30.90 per month, depending on timing. The calculation excludes input tokens, cache-hit charges, taxes, and any other platform fees.
For a production application, the bigger issue may be workload composition rather than output alone. A system with long prompts, low cache reuse, or traffic concentrated in peak windows could experience a materially different increase. Scheduling batch jobs during off-peak hours and tracking cache-hit rates become more important under the new model.
OpenAI announced a contrasting change for ChatGPT. GPT-5.6 Luna became the default model for Free and Go users, while OpenAI also announced unlimited text chats and a new Think button for harder questions.
The headline requires a qualification: unlimited access applies to text chats, not every ChatGPT capability. Limits remain for file uploads, images, voice, image generation, and other tools, and the offer is subject to abuse safeguards.
This is a consumer-distribution strategy rather than an API-pricing strategy. OpenAI is reducing friction for everyday use and keeping a capable model inside the default ChatGPT experience. ChatGPT had recently passed 1 billion weekly users, according to TechCrunch’s report on the announcement.
Low token prices helped Chinese models gain developer usage. One Vercel AI Gateway measure found that Chinese-built open-weight models processed 29% of production tokens in June 2026, up from roughly one-ninth in April, while accounting for less than 4% of spending. The figure describes traffic on that gateway, not the entire global AI market.
That distinction is important. Token share measures how much work flows through a system; it does not by itself establish revenue, profit, customer loyalty, or durable market power. Developers can route workloads between providers, and price-sensitive traffic can move quickly when rates change.
DeepSeek’s new schedule suggests that winning usage is no longer the only objective. Time-based pricing gives the company a way to allocate capacity toward busy periods and charge more when demand is concentrated. Its official documentation describes the change as a resource-allocation measure and encourages users to schedule tasks according to actual usage.
The move also arrives as other Chinese AI companies pursue paid offerings or consider higher prices. Reporting on the broader Chinese market described Moonshot AI and ByteDance as introducing paid plans and Alibaba as preparing to follow, pointing to a wider shift from aggressive subsidy toward monetization.
That does not prove that DeepSeek or its peers are profitable. It does show that market share built through low prices must eventually be converted into revenue per token, recurring contracts, or another sustainable business model.
OpenAI’s free Luna rollout emphasizes distribution instead of the cheapest possible API bill. By making text chat broadly available and adding a visible Think control, OpenAI is trying to make ChatGPT a persistent consumer habit while reserving tighter limits for more expensive features.
The strategic contrast is clear:
These strategies can coexist. A developer may choose DeepSeek for high-volume workloads while an individual uses ChatGPT for general-purpose assistance. But each company is trying to make its ecosystem harder to replace: DeepSeek through cost-performance and developer routing, OpenAI through product familiarity and consumer reach.
DeepSeek is still a low-cost option, but its new V4 pricing makes “cheapest model” an incomplete description. Developers should compare the full bill across cache-hit inputs, cache-miss inputs, output volume, and peak-hour concentration.
The practical questions are now:
The broader China–US AI competition is therefore becoming less about who can offer the lowest headline rate. DeepSeek is testing whether Chinese labs can turn cost-performance leadership into sustainable unit economics. OpenAI is testing whether free, habitual consumer access can produce durable ecosystem control. In both cases, the central challenge is the same: convert enormous usage into a business that can keep funding frontier-model compute.