Bloomberg reported that Gemini 4 did well on industry benchmarks but struggled with some coding tasks in employee testing. Earlier Bloomberg reporting described Google employees’ worries about competition from Anthropic and OpenAI, and a delay to Gemini 3.5 Pro for coding improvements.
Published byEdited with GPT-6 LunaImages generated with GPT Image 2
Research answer

Create a landscape editorial hero image for this Studio Global article: What did Bloomberg report about the gap between Gemini 4’s strong benchmark scores and its performance on real-world coding tasks, why does. Article summary: Bloomberg reported that Gemini 4 scores well on standard AI benchmarks, but some employees testing it say it struggles with certain coding tasks in actual use. That gap matters because Google is preparing to launch a fla. Topic tags: general, news, general web, user generated. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts wi
Bloomberg reported that Gemini 4 performed well on commonly used AI benchmarks but did less well when employees tried it on real tasks, including some coding work. The report captured internal doubts ahead of the model’s release—not a definitive assessment of every capability. Google disputed the criticism, and a same-day report said the company had unveiled Gemini 4 Argon, though it did not give a timetable for general availability. 7
5
10
People with direct access to the effort told Bloomberg that Gemini 4 scored well on industry benchmarks but struggled with certain coding tasks in practice. The accounts were anonymous, so the reported gap is best understood as employee feedback rather than a published, independent evaluation of the model. 7
That distinction matters: benchmark results and reports from people using a model on particular tasks are different kinds of evidence. The Bloomberg account suggests the strong benchmark scores did not resolve every internal concern about coding performance. It does not establish that Gemini 4 performs poorly across all real-world uses. 7
The issue is consequential because Google is competing to build AI coding tools, an area where Bloomberg has reported concern inside the company about rivals’ progress. In April, Bloomberg reported that Google leaders were anxious about falling behind, particularly as Anthropic offered coding tools described by sources as more effective and popular with businesses. 4
A separate July report described Gemini 3.5 Pro as months behind schedule while Google worked to improve its capabilities, especially coding. Employees cited in that report worried that Anthropic and OpenAI were producing models that exceeded Gemini’s capabilities. 1 Reuters also reported that Gemini 3.5 Pro was intended to help Google catch up in AI coding tools and agentic AI tasks.
2
Those earlier reports provide context, but they concern Gemini 3.5 Pro and broader competitive worries—not proof that Gemini 4 has the same shortcomings or that Google has lost the race. Better coding performance could matter for Google’s efforts in coding tools and agentic tasks; the available sources do not establish how Gemini 4’s reported issues affect specific products or customers. 2
4
7
The available reporting supports saying that Gemini 3.5 Pro was delayed while Google worked on its capabilities, particularly coding. It does not establish that Google abandoned the model. 1
2
The distinction is important: a delay is not evidence that a model was canceled, and the reports about Gemini 3.5 Pro should not be treated as confirmation of Gemini 4’s performance. The Bloomberg report on Gemini 4 describes a separate set of employee concerns. 1
7
Google called it “inaccurate” to characterize Gemini 4 as underperforming in coding, according to Bloomberg’s report as relayed by ZeroHedge. That response conflicts with the anonymous employee accounts, and the sources provided do not include a public, independent test that settles the disagreement. 5
7
The timing also needs context: while Bloomberg described internal skepticism as Gemini 4 approached release, a same-day report said Google had unveiled Gemini 4 Argon and had not announced a timetable for general public availability. 7
10
After the report, Alphabet shares gave back some of their gains. They moved from more than 2% higher to about 0.5% higher near the close, according to a report citing Investing.com. That was a reduction in the day’s gains, not a decline below the previous close. 9
The reporting points to a real question for Google: whether Gemini 4’s strong benchmark performance carries over to the coding tasks employees tried. But the evidence is a mix of anonymous internal accounts and Google’s rebuttal, not a complete public evaluation. The earlier Gemini 3.5 Pro delay adds competitive context; it does not prove that the model was abandoned or that Gemini 4 is broadly behind its rivals. 1
5
7
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
Bloomberg reported that Gemini 4 did well on industry benchmarks but struggled with some coding tasks in employee testing.
Bloomberg reported that Gemini 4 did well on industry benchmarks but struggled with some coding tasks in employee testing. Earlier Bloomberg reporting described Google employees’ worries about competition from Anthropic and OpenAI, and a delay to Gemini 3.5 Pro for coding improvements.
After the report, Alphabet shares pared gains from more than 2% to about 0.5% near the close; they were still higher on the day.