What did Bloomberg report about the gap between Gemini 4’s strong benchmark scores and its performance on real-world coding tasks, why doesAn illustration of the difference between benchmark results and reported performance on coding tasks.
AI 提示詞
Create a landscape editorial hero image for this Studio Global article: What did Bloomberg report about the gap between Gemini 4’s strong benchmark scores and its performance on real-world coding tasks, why does. Article summary: Bloomberg reported that Gemini 4 scores well on standard AI benchmarks, but some employees testing it say it struggles with certain coding tasks in actual use. That gap matters because Google is preparing to launch a fla. Topic tags: general, news, general web, user generated. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts wi
openai.com
彭博報導指出,Gemini 4 在業界常用的 AI 基準測試中表現不錯,但部分直接接觸專案的員工表示,模型實際處理某些編碼任務時不如測試成績所呈現。這些說法來自要求匿名的內部人士,反映的是員工回饋,並非對模型所有能力的公開、獨立評估。 7
這項疑問對 Google 尤其重要,因為 AI 編碼工具是公司與其他業者競爭的領域之一。彭博今年 4 月曾報導,Google 內部主管擔心公司在 AI 編碼工具競賽中落後;消息人士認為,Anthropic 提供的相關工具對企業更有效,也更受歡迎。 4
與 Gemini 3.5 Pro 的延遲有何關係?
今年 7 月,彭博報導 Gemini 3.5 Pro 的推出較原定時程延後數月,Google 當時正設法改善模型能力,尤其是編碼表現。報導引述的員工也擔心 Anthropic 和 OpenAI 的模型能力超越 Gemini。 1 路透社則指出,Gemini 3.5 Pro 的目標之一,是協助 Google 在 AI 編碼工具及代理式 AI 任務方面追趕競爭對手。 2
這些先前報導提供了 Google 面臨的競爭背景,但談的是 Gemini 3.5 Pro 的延遲與更廣泛的競爭疑慮,不能據此認定 Gemini 4 也有相同問題,或 Google 已在競賽中落敗。現有報導亦不足以證明 Gemini 3.5 Pro 已遭放棄:報導支持的說法是它延後推出,Google 當時仍在改善能力。 124
更好的編碼能力可能影響 Google 發展編碼工具及代理式 AI 任務的努力,但現有來源沒有說明 Gemini 4 的相關疑慮會如何影響特定產品或客戶。 27
報導提出了 Google 必須面對的問題:Gemini 4 的基準測試成績,能否在員工嘗試的編碼工作中同樣體現?目前可見的證據主要是匿名內部人士的說法,以及 Google 的反駁,還不是完整的公開評估。Gemini 3.5 Pro 過去的延遲可作為競爭背景,但不代表該模型已遭放棄,也不足以證明 Gemini 4 整體落後對手。 157