This is the story of that launch — what the model can actually do, how much it costs, and why it represents a significant moment in the ongoing US-China AI competition.
Kimi K3 scored 57.1 on the independent Artificial Analysis Intelligence Index v4.1, placing it in 3rd place globally behind Claude Fable 5 (60) and GPT-5.6 Sol (59) . That gap is roughly 3 points — tight by historical standards, but clear: the broadest available intelligence test still favors the top US models
. Moonshot itself acknowledged this, stating publicly that K3 "still trails the most powerful proprietary models" from Anthropic and OpenAI on overall aggregate performance
.
Where K3 pulls ahead is on specific, real-world agentic benchmarks:
| Benchmark | Kimi K3 | Claude Fable 5 | GPT-5.6 Sol | Claude Opus 4.8 |
|---|---|---|---|---|
| Terminal-Bench 2.1 | 88.3 | 84.6 | 88.8 | 84.6 |
| DeepSWE | 73.0 | 70.0 | 67.5 | 59.0 |
| BrowseComp | 91.2 | 88.0 | 90.4 | 84.3 |
| Automation Bench | 75.6 | — | — | — |
| SpreadsheetBench 2 | 34.8 | 34.7 | 32.4 | 31.6 |
On these five agentic benchmarks plus Program Bench (77.8), K3 won 5 of 6 outright, beating both Fable 5 and GPT-5.6 Sol on those specific tests . It also substantially outperformed Claude Opus 4.8 on every published benchmark — a meaningful result because Opus 4.8 is priced at $5 input / $25 output per million tokens, making it a direct cost competitor
.
On the Arena blind human preference platform built by UC Berkeley researchers, K3 topped the coding leaderboard shortly after launch . Arena's crowd-comparison evaluations showed K3 taking first place in front-end web development with a score of 1,679, ahead of Fable 5 (1,631) and GPT-5.6 Sol (1,618)
.
On GPU kernel optimisation — the techniques that improve how AI models interact with hardware — Moonshot reported that K3 "substantially outperformed" Opus 4.8, GPT-5.6 Sol, and GPT-5.5 . This is commercially significant because better kernel optimization means lower latency and better throughput on the same hardware.
The model uses a sparse MoE architecture with 896 experts, of which only 16 are activated per token, which makes it roughly comparable to GPT-5.6 Sol's inference cost despite being 2.8 trillion parameters total . It also introduces Kimi Delta Attention (KDA) and Attention Residuals (AttnRes) — novel architectural components designed to improve long-context reasoning
.
One of the most market-disruptive aspects of Kimi K3 is its API pricing, which undercuts the top US models by a significant margin:
| Model | Input (per 1M tokens) | Output (per 1M tokens) |
|---|---|---|
| Kimi K3 | $3.00 | $15.00 |
| GPT-5.6 Sol | $5.00 | $30.00 |
| Claude Fable 5 | $10.00 | $50.00 |
| Claude Opus 4.8 | $5.00 | $25.00 |
Kimi K3 costs ~40% less than GPT-5.6 Sol and ~70% less than Claude Fable 5 at list rates . The three models sit on a clean 3.3× ladder: Fable 5 costs 3.3 times Kimi K3 on both input and output, with GPT-5.6 Sol almost exactly in between
.
Additionally, K3 has a cache-hit input rate of $0.30 per million tokens — 10x cheaper than its cache-miss rate — for context queries that have been served before . On a per-task basis, K3 is estimated to cost 50-65% less than the US flagships
.
Moonshot's pricing for K3 represents a notable shift: it is now priced like a frontier model rather than a budget option . Previous Kimi models like K2.5 and K2.6 were priced at $0.60 input / $2.50-4.00 output, making K3 roughly 3-4× more expensive than its own predecessors
. However, it remains cheaper than the top-tier US models while delivering comparable or better performance on specific agentic tasks.
On July 19, 2026 — roughly 48 hours after launch — Moonshot AI announced it was temporarily pausing new subscriptions for Kimi K3. The company posted on X: "Kimi K3 has received far more love than we expected, and our GPUs are feeling it. Over the past 48 hours, demand has pushed close to the limits of our current capacity" .
The company said it would prioritize compute for existing subscribers while adding capacity as fast as possible, planning to reopen subscription spots in batches . Existing users retained full access, and enterprise/API services continued
.
Multiple news outlets noted that this incident highlights China's ongoing compute constraints due to US chip export restrictions . The South China Morning Post wrote that the situation "underscores the challenges Chinese AI labs face in serving their products to global users"
. The model's full weights, which would allow third parties to run it on their own hardware, were not due until July 27 — leaving Moonshot as the sole provider of inference capacity during the launch window
.
Moonshot used the pause to split its membership into two plans: Kimi Membership (for web, app, and Work) and Kimi Code Membership (for coding workflows), a restructuring designed to better allocate compute resources .
The launch of Kimi K3 had ripple effects beyond AI benchmarks. The Straits Times directly reported that K3 "fueled a tech rout," and multiple outlets linked the release to a selloff in US semiconductor and AI stocks . While specific figures about the Philadelphia Semiconductor Index's direction or Nvidia's stock movement could not be independently confirmed within the available source set, the broader market context of the selloff is well-documented.
On the corporate side, Reuters reported on July 20, 2026, that the subscription pause "comes as the company seeks fresh funding, with sources saying it is moving toward an initial public offering in Hong Kong" . However, no source within the available evidence confirmed a specific shareholder resolution for a $30 billion valuation target. The Reuters report mentions an IPO push without specifying a target valuation
.
Full model weights for Kimi K3 are scheduled for public release on July 27, 2026 . This is a critical detail because it means developers and organizations can eventually download, run, inspect, fine-tune, and own the model rather than just renting it through an API
. The two-step rollout — API first, open weights eleven days later — is the shape of the entire announcement
.
Prior to the open-weight release, developers can access K3 through kimi.com, the Kimi API, or OpenRouter . The model speaks the OpenAI SDK, supports tool calling, structured output, and automatic context caching, and has a 1-million-token context window with no long-context surcharge
.
That said, running a 2.8-trillion-parameter model locally is not trivial. The file size is estimated at roughly 1.5 TB, and recommended hardware includes 64+ enterprise NVIDIA H100 GPUs . For most teams, the API will remain the practical access point.
Kimi K3 is not the model that dethrones GPT-5.6 Sol or Claude Fable 5 on overall intelligence — Moonshot itself says so. But it is the model that proves an open-weight system can compete with frontier proprietary models on specific, commercially valuable tasks while costing a fraction of the price. It is the largest open-weight model ever released, and it arrived with enough real-world demand to crash a company's GPU infrastructure in two days.
That combination — frontier-tier performance on agentic benchmarks, aggressive pricing, open weights, and a clear signal of demand — makes K3 a significant milestone in the AI landscape, regardless of where it sits on any single leaderboard.