核心新功能:Google 正式將「即時影像生成與編輯」功能直接整合進對話式 AI 服務「Gemini Live」中 [8]。 操作方式:在與 Gemini Live 語音對話時,可啟動相機分享模式,將鏡頭對準主體,並以自然語言下達生成或編輯影像的指令 [8]。

Create a landscape editorial hero image for this Studio Global article: What real-time image generation and editing capability has Google added to Gemini Live, how does it work on Android and iOS, what technology. Article summary: ## What Google Added to Gemini Live. Topic tags: general, documentation, general web, user generated, education. Reference image context from search candidates: Reference image 1: visual subject "Google has begun the deployment of Gemini's innovative real-time AI video functionalities, enabling the platform to interpret visual input from a user's device" source context "Google's Gemini update that can tell you live what it sees through your camera is now rolling out - PhoneArena" Reference image 2: visual subject "Smartphones must have user-replaceable batteries by 2027. But not your iPhone. Here's why" source context "Google's Gemini update that can tell you l
Google 在 2026 年開發者大會(I/O)上,為旗下 AI 助理 Gemini 注入了一項極具野心的功能:在對話中即時生成與編輯影像。這不只是單純的 AI 繪圖,而是將創作流程無縫融入你與 AI 的即時對話中。
Google 已正式將直接的即時影像生成和編輯能力,整合進其對話式 AI 模式「Gemini Live」中 。這代表著,當你開啟 Gemini Live 進行口語對話時,不再只能回覆文字,而是可以:
這項功能的精髓在於,它將過去繁瑣的「拍張照 -> 上傳 -> 輸入提示詞 -> 等待生成」的流程,壓縮成一氣呵成的對話體驗。
目前,此功能已經推送到 Android 和 iOS 的 Gemini App 上。
要特別留意的是,既有 Gemini App 中的影像生成功能,是基於文字或圖片上傳的轉換,而這次在 Gemini Live 中的新功能,則是將這個創作與編輯的過程,帶入了即時語音與影像的對話情境中 。
這一切的技術核心,仰賴 Google DeepMind 的最新影像模型:Gemini 2.5 Flash Image,它有個可愛的內部暱稱 「nano-banana」,是 Google 目前最先進的影像生成與編輯模型 。它的關鍵能力包括:
此模型已透過 API 和 Google AI Studio 提供給開發者,每百萬輸出 token 的定價為 30 美元,每張圖片計算為 1290 個輸出 token 。
Google 在 I/O 2026 上的一系列重磅發布,清晰地展現了其圍繞著 Gemini 打造一條「從理解到創造」的統一多模態管線的野心。除了 Gemini Live 的影像功能,還有更多重大更新:
這是大會上最受矚目的新模型。官方宣稱 Gemini Omni 是一個能「從任何輸入創造出任何東西」的模型,而第一步就是影片 。
Gemini 3.5 Flash 成為 Gemini App 和 Google 搜尋「AI 模式」的新預設模型 。其核心賣點是 「以 Flash 系列的速度,提供媲美旗艦模型的智慧」
。
Google 的策略核心,並非只開發一個強大的影像或影片模型,而是建立深度整合的即時多模態管線。
Google 的優勢在於其生態系統的垂直整合能力,從底層的 Gemini 模型、上層的應用程式(搜尋、Workspace),到終端設備(Android 手機),形成一個能直接面向數十億用戶的封閉迴路。關鍵挑戰則在於,當這些功能大規模上線後,其實際體驗能否如展示般流暢且可靠 。
從「nano-banana」到「Omni」,Google 正將 AI 助理的角色從「問答機」轉變為一個能理解你的意圖、看見你的世界,並即時為你創造內容的「全能夥伴」。
Studio Global AI
Use this topic as a starting point for a fresh source-backed answer, then compare citations before you share it.
核心新功能:Google 正式將「即時影像生成與編輯」功能直接整合進對話式 AI 服務「Gemini Live」中 [8]。
核心新功能:Google 正式將「即時影像生成與編輯」功能直接整合進對話式 AI 服務「Gemini Live」中 [8]。 操作方式:在與 Gemini Live 語音對話時,可啟動相機分享模式,將鏡頭對準主體,並以自然語言下達生成或編輯影像的指令 [8]。
Android 與 iOS 體驗:此功能已推送至 Android 和 iOS 版的 Gemini App,目前已知的操作流程為在 Live 對話中啟用相機並下達指令 [8]。