Reports describe Google employees trying an internal Gemini 4 build called Carbon in the company’s Jetski coding environment. One tester reportedly said its coding “feels like Opus 5.5,” Anthropic’s model for long-running agentic coding—but that is an early impression, not a published head-to-head result.
29
25 There is no confirmed Carbon benchmark or public release plan in the available reporting.
48
What employees reportedly said about Carbon
Business Insider reported that employees had been testing Carbon, citing internal documents and screenshots.
29 The reported feedback was positive, including one employee’s tentative comparison with Claude Opus 5.5; that employee also said more testing was needed.
29
25
That distinction matters. A subjective coding impression can suggest that a build is promising, but it does not establish how it performs across tasks, how consistently it works in extended agent workflows, or whether it matches a competing model under controlled conditions. No Carbon-specific comparative benchmark is included in the available reporting.
29
48
The report also does not establish Google’s public plans for Carbon. Its eventual product name, relationship to Argon, and whether it will be released at all remain unconfirmed.
25
48
How Carbon relates to Gemini 4 Argon
The codenames describe internal model versions, not a confirmed sequence of public products. Business Insider reporting says Google tested Gemini 4 versions associated with the names Argon, Barium, and Carbon. A report based on an internal document says the Barium-B build was selected for the public model announced as Gemini 4 Argon.
25
That does not tell us exactly how Carbon fits into the lineup. It could be a later checkpoint or another version in the Gemini 4 family, but the reporting does not confirm its precise lineage or whether it would ship under the Argon name.
25
48
Google announced Gemini 4 Argon on September 30, 2026, initially rolling it out to trusted cyber defenders through the Fairwind Program.
36
4 Coverage on October 9 still described Argon as not publicly available, and identified paid API customers and Google AI Ultra subscribers as intended groups for a later expansion—not as users with confirmed general access at that point.
48
16
Google’s launch materials report a 77.9% score on DeepSWE v1.1 and an increase in Argon’s output-token limit from 64,000 to 1 million.
4 Those are Argon details; they are not Carbon results. Third-party coverage also listed introductory API pricing of $2 per million input tokens and $10 per million output tokens, but the available material here does not include an official pricing schedule.
5
Argon’s coding debate is separate from Carbon’s early feedback
Separate reporting on Argon described a gap between benchmark results and some employees’ experience using the model on practical coding tasks, including front-end work. Google disputed the characterization that Argon was weak at coding.
2
9 That disagreement concerns Argon and should not be treated as evidence about Carbon’s performance.
Another account says early internal Argon versions reminded an employee of Anthropic’s older Opus 5 on some coding tasks, while Carbon later drew the more favorable Opus 5.5 comparison.
30 Both are reported employee impressions, not controlled evaluations. Google’s official announcement presents Argon as a frontier model for complex, long-horizon workflows, but that claim does not independently verify Carbon’s reported performance.
36
The careful takeaway is that Carbon sounds promising to at least one reported tester, while its performance and product status remain open questions. Until there are results for Carbon itself—or a confirmed release—claims that it matches Claude Opus 5.5 should be understood as preliminary internal feedback, not a verified model comparison.
25
29
48