September reports described suspected Gemini 4 Pro checkpoints on LMArena under a Gemini 3.8 Flash label, including a claimed 256,000 token output limit.
Published byEdited with GPT-6 LunaImages generated with GPT Image 2
Research answer

Create a landscape editorial hero image for this Studio Global article: What has been reported about the suspected Gemini 4 Pro checkpoints appearing on LMArena under Gemini 3.8 Flash labels—including the Septemb. Article summary: The LMArena sightings are reports of **suspected** Gemini 4 Pro checkpoints, not proof that Google put Gemini 4 Pro on the platform. Google DeepMind’s clearest reported confirmation is narrower: Gemini 4 has entered earl. Topic tags: general, general web, user generated. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fa
Reports of unusually capable models on LMArena have fueled speculation that Google may be testing Gemini 4 Pro under a gemini-3.8-flash label. The sightings, reported token limit and tester impressions are not confirmation of the model’s identity or official specifications. Separately, Google DeepMind leader Koray Kavukcuoglu has reportedly said Gemini 4 is in early post-training and that the team hopes to release an early version soon, without giving a date. 17
22
34
Reports placed one model labeled gemini-3.8-flash on LMArena on September 17, then described another sighting later in the month. Observers suspected the checkpoints might be related to Gemini 4 Pro based on their outputs. But a platform label does not establish what model is running behind it, and the reports do not confirm that Google put Gemini 4 Pro on LMArena. 33
34
38
“Argon” has circulated as an alleged internal codename for the suspected checkpoint. A reported screenshot also appeared to show a 256,000-token output limit, compared with 64,000 tokens previously. These are leak-based claims, not Gemini 4 Pro specifications confirmed by Google. 33
38
Testers described strong SVG generation, interactive games and detailed 3D-style visuals. One account said the model took about 2.4 minutes to generate a complex visual scene in a high-thinking setting; another described extended reasoning that took several minutes. These are anecdotal task reports, not standardized speed or design benchmarks, so they do not establish that the model is generally faster or better than other systems. 33
38
47
At The Information’s AI Agenda Live Summit, DeepMind leader Koray Kavukcuoglu reportedly said Gemini 4 had entered early post-training. Coverage of his remarks also says engineers were using the model internally in the Antigravity coding tool. 17
22
54
Kavukcuoglu reportedly said the team intended to release an early post-training version as soon as possible and hoped to do so “much earlier” than the end of 2026. He did not announce a launch date, so a specific month or a release timed ahead of a competitor remains speculation. 17
22
28
48
Post-training is not the same as a public launch or a completed safety evaluation. The reporting describes safety testing and guardrails as part of the work, but the sources cited here do not provide Gemini 4-specific safety-test results. 22
54
The LMArena label, the “Argon” codename, the 256,000-token output limit and the performance impressions should all be treated as unverified. The clearest reported update from DeepMind is about development stage and release intent: Gemini 4 is in early post-training, an early version is a goal, and no public release date has been set. 17
22
34
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
September reports described suspected Gemini 4 Pro checkpoints on LMArena under a Gemini 3.8 Flash label, including a claimed 256,000 token output limit.