| Specification | Grok 4.5 | Grok 4.6 |
|---|---|---|
| Parameters | 1.5 trillion | 2 trillion |
| Foundation model | V9 (completed training May 26, 2026) | Next-gen (architecture details not disclosed) |
| Status | Private beta at SpaceX & Tesla since June 28; API launched July 8 | Training finishing this week; launch expected August |
| Pricing (API) | $2/1M input, $6/1M output tokens | Not yet announced |
| Reported performance | Self-reported as near/exceeding Claude Opus; no independent benchmarks published | Claims to exceed Grok 4.5; aims to surpass Kimi K3 |
A critical caveat: No independent third-party benchmarks have been published for Grok 4.5 — all performance claims are xAI's own . One xAI engineer noted that Grok 4.5 used supplemental "Cursor data" (coding-focused data) not included in initial training, which was "not quite as good as having it in initial training"
.
The timing of Musk's announcement is directly tied to Moonshot AI's launch of Kimi K3 on July 16, 2026 . Kimi K3 is a 2.8-trillion-parameter, open-weight Mixture-of-Experts (MoE) model with a 1-million-token context window and native vision capabilities, making it the largest open-weight AI model ever released and the first in the "3T-class"
. It activates only 16 of its 896 experts per token, a design choice that reduces compute costs
.
Musk's posts explicitly frame Grok 4.6 as the response, stating the goal is to surpass Kimi K3 in performance while maintaining Grok's speed and efficiency advantages . Grok 4.6 at 2T parameters is smaller than Kimi K3's 2.8T, so xAI is betting on architecture efficiency and inference optimization rather than raw parameter count
.
This is where the story gets important: Grok 4.5 has no independent benchmark scores. All its competitive claims are self-reported by xAI. In contrast, Kimi K3 has been independently evaluated and ranks #3 globally on comprehensive leaderboards, behind only Anthropic's Claude Fable 5 and OpenAI's GPT-5.6 Sol . It has been described as performing "neck-and-neck with the most powerful proprietary systems from US rivals"
. K3 outperforms Claude Fable 5 on certain coding and web development benchmarks
.
On pricing, Kimi K3's API is reported to be roughly half that of comparable frontier models like Claude Fable 5 and GPT-5.6 Sol . Exact per-token pricing was not found from high-authority sources.
The key takeaway: No direct head-to-head independent benchmark comparisons exist between any Grok model and Kimi K3. All xAI's performance statements are unverified claims made by Musk on social media.
Musk's July 18 announcement fits into a much larger plan. Here's what xAI has publicly committed to:
Alongside the model race, xAI open-sourced Grok Build on July 15–16, 2026 under the Apache 2.0 license, publishing its Rust-based CLI coding agent on GitHub at xai-org/grok-build . Grok Build features 8 parallel subagents, a plan-first workflow, and Arena Mode. The open-sourcing move follows a privacy controversy: security researchers found that a prior version uploaded users' entire working directories to xAI's servers
. By opening the code, xAI aimed to build trust and accelerate community adoption.
Musk separately confirmed on July 8 that xAI is making daily improvements to Grok Build driven by user feedback . This positions xAI to directly challenge coding-agent incumbents like Cursor and GitHub Copilot.
Musk's July 18 Grok 4.6 announcement is a clear competitive counterpunch to Moonshot's Kimi K3 launch just two days earlier. xAI's monthly release cadence is faster than any other frontier lab's, and the Grok 5 roadmap signals much larger models ahead. But the lack of independent benchmarks for any xAI model means all competitive claims remain unverified until third-party testing emerges. In a race defined by transparency as much as scale, that gap matters.