Google’s Gemini 4 Argon has drawn two different assessments: the company points to leading benchmark results, while some employees reportedly say the model is less dependable when they use it for real coding work. The gap matters, but the available reporting describes mixed internal opinions—not a consensus that Gemini 4 is behind its rivals.
11
Why benchmark scores may not settle the coding question
Google says Gemini 4 performed strongly on several benchmarks, including a security test on which it outscored OpenAI’s Astra. Employees familiar with the model’s internal evaluations reportedly describe a different experience with some practical coding tasks, including front-end development.
11
14
Those accounts don’t invalidate the benchmark results, and the benchmark results don’t resolve every question about performance on day-to-day engineering work. They point to a difference between scores on selected evaluations and how useful or consistent a model feels on specific tasks.
Google disputes the criticism—and staff opinions differ
Google said it would be inaccurate to characterize Gemini 4 as underperforming in coding, while highlighting its benchmark results. Separately, the employee accounts reported in the coverage are not unanimous: some staff reportedly believe Anthropic’s Fable and OpenAI’s Astra are improving faster, while others think Gemini 4 has caught up with leading models.
11
12
13
That disagreement makes it premature to treat either the internal criticism or the company’s benchmark case as a complete verdict on Gemini 4’s capabilities.
The Gemini 3.5 Pro plan adds release context
Google had planned to release Gemini 3.5 Pro in June 2026 but abandoned that version and moved to Gemini 4, according to reports. That history adds context to questions about the company’s development and release pace, but it does not, by itself, explain Gemini 4’s reported coding weaknesses or establish that the model is inferior.
6
13
Why investors paid attention
After the report, Alphabet shares gave back an earlier gain of more than 2%, ending up about 0.5% in the reported Wednesday session. The price movement shows that the news coincided with a change in market sentiment; it does not prove the employee concerns are correct or establish that those concerns caused the move.
1
14
The broader investor question is whether Google can compete with OpenAI and Anthropic in flagship AI models. The reporting frames Gemini 4 as part of that competition, but does not identify specific effects on individual Google products. Any wider impact remains a question, not a demonstrated outcome.
10