Open weight models reached 62% of Vercel AI Gateway token volume on the Saturday before August 25, then held a 54% share on Tuesday versus 46% for proprietary models. DeepSeek V4 Flash was the main catalyst: its updated weights materially improved agentic performance while preserving an efficiency focused price adva...
Research answer

Create a landscape editorial hero image for this Studio Global article: What drove the record surge in usage of low-cost Chinese open-weight AI models—particularly DeepSeek’s latest lightweight model—on Vercel’s. Article summary: The surge was driven by developers shifting high-volume workloads to inexpensive Chinese open-weight models, led by DeepSeek V4 Flash: a lightweight model whose updated weights improved agentic performance substantially.. Topic tags: general, general web, user generated, news. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts wi
A cost-and-capability tradeoff is reshaping traffic on Vercel’s AI Gateway. Open-weight models accounted for 54% of gateway token volume on Tuesday, August 25, 2026, compared with 46% for proprietary models. On the preceding Saturday, their share reached a reported record of 62%, leaving closed models with 38%.4
The immediate catalyst was DeepSeek V4 Flash, whose updated weights improved performance on coding, tool use, and agentic workflows without turning it into a premium-priced model.19
| Reported point in August | Open-weight models | Proprietary models |
|---|---|---|
| Preceding Saturday | 62% | 38% |
| Tuesday, August 25 | 54% | 46% |
That means open-weight models held an 8-point lead on Tuesday and a 24-point lead at the earlier record. The comparison is specifically about tokens processed through the gateway. It should not be read as a measure of revenue, profit, or the number of businesses using each model.
Vercel said DeepSeek V4 Flash began running updated weights by default on AI Gateway, with its Terminal-Bench score rising to 82.7 from 56.9 in the April preview—a 25.8-point improvement.19 The update required no change to the model ID or application code, reducing the friction for existing users to receive the newer weights.19
The model’s appeal was therefore broader than simply being inexpensive. Its improved agentic performance made it more suitable for high-volume coding and tool-use workloads, while its efficiency-tier positioning gave developers a way to control inference costs. Vercel describes the model as supporting reasoning, tool use, and implicit caching, with a context window of 1 million tokens.27
Pricing helped reinforce that positioning. Reuters reported that DeepSeek’s V4 Pro launched at roughly nine times V4 Flash’s input price and 14 times its output price, highlighting the gap between the company’s efficiency model and premium offering.17 Independent comparisons also show a substantial price advantage for V4 Flash, although benchmark results vary by test and configuration.25
The available reporting links the decline in proprietary-model volume to weaker business demand for Anthropic’s more expensive Fable 5.4 That explanation is plausible for workloads where developers care more about throughput and unit economics than maximum model capability, but the available gateway figures do not prove a single cause for every customer or request.
Fable 5 remains positioned for long-running, ambiguous, multi-step tasks, according to Vercel’s product description.3 In other words, the August traffic change looks less like a universal verdict that open models are better and more like a routing decision: use a lower-cost model for large volumes of suitable work, and reserve premium proprietary models for tasks that justify their higher price.
The most important caveat is the difference between usage and revenue. Earlier Vercel reporting showed open-weight models rising to 29% of gateway token volume while accounting for under 4% of spend.44 That gap exists because token counts do not capture the price paid per input or output token.
So the August crossover signals that open-weight models were processing more of the gateway’s workload, not necessarily generating more business for providers. Proprietary models can retain a disproportionate share of spending even while cheaper open models handle more tokens.
The data points toward a more deliberate model-routing strategy:
The broader lesson from Vercel’s August traffic is that open-weight models can win production volume when they combine adequate capability with sharply lower costs. DeepSeek V4 Flash supplied the clearest example, but the underlying shift is toward matching each workload with the least expensive model that performs reliably enough—not replacing every proprietary model at once.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
Open weight models reached 62% of Vercel AI Gateway token volume on the Saturday before August 25, then held a 54% share on Tuesday versus 46% for proprietary models.
Open weight models reached 62% of Vercel AI Gateway token volume on the Saturday before August 25, then held a 54% share on Tuesday versus 46% for proprietary models. DeepSeek V4 Flash was the main catalyst: its updated weights materially improved agentic performance while preserving an efficiency focused price advantage.
The shift does not mean open models captured most gateway spending; earlier Vercel data showed proprietary providers could still dominate spend despite handling fewer tokens.