On September 14, Musk said the 2.5 trillion parameter Grok 4.8 would finish pretraining that week and begin reinforcement learning, while Grok 5 was positioned as xAI’s AGI candidate. Musk put the delayed Grok 4.7 roughly on par with Anthropic’s Opus 5.0—not Opus 5.1—and said multimodal performance still needed work.
Published byEdited with GPT-5.6 TerraImages generated with GPT Image 2
Research answer

Create a landscape editorial hero image for this Studio Global article: What did Elon Musk claim on September 14 about xAI’s roadmap toward artificial general intelligence—specifically, that the 2.5-trillion-para. Article summary: Musk’s September 14 roadmap was an aspirational sequence, not a substantiated AGI forecast: Grok 4.8, a claimed 2.5-trillion-parameter model using a new C++ training stack, would finish pretraining that week and then ent. Topic tags: general, general web, user generated. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fa
Musk’s September 14 xAI roadmap described an ambitious sequence of unreleased Grok models: Grok 4.8 would complete pretraining during the week and move into reinforcement learning (RL), while Grok 5 was presented as the model that could reach artificial general intelligence (AGI). But the announcement supplied no public benchmark results, model card, operational definition of AGI, or delivery date. It should therefore be read as a founder forecast—not a demonstrated AGI result. 1
2
4
Musk described Grok 4.8 as a 2.5-trillion-parameter model trained on xAI’s new C++ software stack. He said pretraining would finish that week, followed by RL. 1
2
In the same discussion, he sketched a capability ladder for the models after Grok 4.7:
That wording matters. “Maybe better than anything” is deliberately less concrete than a measured performance claim, and the AGI framing did not specify what abilities, reliability threshold, autonomy, safety standard, or evaluation suite would qualify a model as AGI. Reporting characterized Grok 5 as xAI’s claimed AGI breakthrough, but the underlying roadmap remained short on testable details. 2
10
Musk was more restrained about the nearer-term Grok 4.7. He said it should be “roughly on par with Opus 5.0, not 5.1,” performing better in some areas and worse in others. He also acknowledged that xAI still needed to improve multimodal performance. 3
8
That is a notably narrower claim than declaring overall leadership. It also makes the later roadmap inherently provisional: 4.8, 4.9, and 5 were discussed before developers had a public 4.7 release to evaluate independently.
On September 2, Musk had said Grok 4.7 was about 10 days away, implying a release around September 12. By September 11, he said it needed “a few more days.” 7
20
His explanation was an RL-tuning problem: xAI may have penalized response length too heavily, causing the model to stop early on difficult tasks it could otherwise solve and to check its work insufficiently. Musk framed that as a possible diagnosis, rather than a final technical explanation. 20
21
This is more than a scheduling detail. It demonstrates the distinction between completing a large pretraining run and shipping a dependable product. RL behavior, evaluation, multimodal quality, reliability, safety work, deployment readiness, and pricing can all affect when—and whether—a model is released.
As of September 14, Grok 4.7 had not been released publicly. Reports noted that xAI’s latest documented flagship remained Grok 4.6, released on August 12, with listed API pricing of $2 per million input tokens and $6 per million output tokens. 3
11
At the same time, Claude Fable 5.1, Gemini 3.8 Flash, and GPT-6 Astra were reported as September releases or announced offerings. 34
45 That competitive context raises the stakes for xAI’s roadmap, but it does not validate comparative performance claims for unreleased Grok models.
The reported 2.5-trillion-parameter scale and new training stack may indicate substantial infrastructure investment, but neither parameter count nor compute capacity independently proves capability, reliability, or AGI. The Grok 4.7 delay is a practical reminder that post-training behavior can be decisive even after a model has been trained. 1
20
For developers, customers, and investors, the useful checkpoints are more concrete than roadmap language:
Until those materials appear, Musk’s September 14 statements establish xAI’s intended direction: 4.8 after pretraining and RL, a stronger 4.9, and a highly ambitious Grok 5. They do not establish that any of those models has reached the promised capability level. 2
4
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
On September 14, Musk said the 2.5 trillion parameter Grok 4.8 would finish pretraining that week and begin reinforcement learning, while Grok 5 was positioned as xAI’s AGI candidate.
On September 14, Musk said the 2.5 trillion parameter Grok 4.8 would finish pretraining that week and begin reinforcement learning, while Grok 5 was positioned as xAI’s AGI candidate. Musk put the delayed Grok 4.7 roughly on par with Anthropic’s Opus 5.0—not Opus 5.1—and said multimodal performance still needed work.
The missed roughly September 12 target for Grok 4.7 illustrates why large training runs and parameter counts alone do not establish a reliable release date or capability level.