Google I/O 2026 於 5 月 19 日宣布兩款重要 AI 模型:Gemini Omni(可從文字、圖片、音訊、影片等多種輸入生成最高 10 秒的影片)與 Gemini 3.5 Flash(首款結合邊疆級智慧與自主行動能力的模型)。 Gemini Omni Flash 已於 6 月底進入公開預覽版,能在 YouTube Shorts 免費使用,並支援對話式多輪編輯與角色一致性。
研究答案

Create a landscape editorial hero image for this Studio Global article: Search & fact-check with cited sources for What new AI media models did Google announce on Tuesday, what are their names, capabilities, spee. Article summary: On Tuesday, May 19, 2026 (Google I/O), Google announced two major new AI media models: **Gemini Omni** (a multimodal video generation model) and **Gemini 3.5 Flash** (a frontier intelligence + agentic model). A third, **. Topic tags: general, general web, user generated, documentation. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks,
在 2026 年 5 月 19 日的 Google I/O 大會上,Google 宣布了兩款定位截然不同的重大 AI 模型:Gemini Omni 是能從任何輸入生成影片的多模態影片生成模型,而 Gemini 3.5 Flash 則是 Google 首款將邊疆級(frontier-level)智慧與自主行動能力(agentic capabilities)結合的模型。此外,第三款模型 Gemini Omni Flash 已於 6 月底進入公開預覽版,成為 Omni 家族中首個可投入生產的版本 。以下為完整解析。
它是什麼?能做什麼?
Gemini Omni 能接受文字、圖片、音訊和現有影片作為輸入,並生成高品質的短影片片段,同時包含同步音訊,且模型的判斷基礎來自 Gemini 對真實世界的知識 。首發版本 Gemini Omni Flash 目前能原生輸出最高約 10 秒的影片片段。圖片和音訊的輸出功能則規劃在未來推出
。此模型支援對話式的多輪編輯:你可以先產生一段片段,再用自然語言修改,系統會在不同版本之間維持角色外觀與物理邏輯的一致性
。
速度與成本
Google 將 Omni 描述為一款「高速」模型,適合快速生成影片 。截至 2026 年 6 月,開發者 API 的單次調用定價尚未公布。不過,Gemini Omni Flash 在 YouTube Shorts Remix 內(限 18 歲以上使用者)可以免費使用;若要在 Gemini App 中使用,則需訂閱 Google AI Plus、Pro 或 Ultra 方案
。
可用平台
gemini-omni-flash-preview 模型 首個產業合作夥伴
在 I/O 大會上,Google 宣布與 Believe / TuneCore 合作,為音樂人、製作人與詞曲創作者提供 Lyria 3 Pro 及 Gemini Omni 工具 。
它是什麼?能做什麼?
Gemini 3.5 Flash 是 Google 首款結合邊疆級智慧與自主行動能力的模型 。它在程式碼撰寫、自主任務執行的能力,以及多模態基準測試中,都超越了上一代的 Gemini 3.1 Pro
。此模型支援最高 100 萬個 token 的情境視窗(context window),每次回應最多可生成 66,000 個 token
。
速度
Google 指出,Gemini 3.5 Flash 在每秒輸出 token 數上,比競爭對手的邊疆級模型快 4 倍 。執行長 Sundar Pichai 形容:「這是一款能力非常強的模型,已達到邊疆級水準並可與最優秀的模型匹敵,但它仍然非常快速」
。
定價(截至 2026 年 6 月,官方 Gemini API 定價)
| Token 類型 | 每百萬 tokens 價格 |
|---|---|
| 輸入 | $1.50 美元 |
| 輸出(含思考 tokens) | $9.00 美元 |
| 快取輸入 | $0.15 美元 |
| 免費方案 | 無限免費使用(有速率限制) |
可用平台
產業合作夥伴
截至 2026 年 6 月,Gemini 3.5 Flash 尚未有公開命名的外部企業合作夥伴(與 Omni 的 Believe/TuneCore 合作不同)。它主要作為平台模型,開放給所有開發者透過 API 與雲端服務使用。
| 模型 | 主要功能 | 速度 | API 成本(輸入/輸出,每百萬 tokens) | 主要可用平台 | 主要合作夥伴 |
|---|---|---|---|---|---|
| Gemini Omni / Omni Flash | 多模態影片生成與編輯 | 高速 | 尚未公佈(YouTube Shorts 免費;Gemini App 需訂閱) | YouTube Shorts、Gemini App、Google Flow、Flow Music、Gemini API(預覽版) | Believe / TuneCore |
| Gemini 3.5 Flash | 邊疆級智慧 + 自主行動能力 | 比競爭對手快 4 倍 | $1.50 / $9.00 美元 | Gemini App、Google 搜尋、Android Studio、Vertex AI、Gemini API | 暫無公開合作夥伴 |
Studio Global AI
這個頁面包含附來源佐證的答案,你可以在 Studio Global 內繼續追問。
Google I/O 2026 於 5 月 19 日宣布兩款重要 AI 模型:Gemini Omni(可從文字、圖片、音訊、影片等多種輸入生成最高 10 秒的影片)與 Gemini 3.5 Flash(首款結合邊疆級智慧與自主行動能力的模型)。
Google I/O 2026 於 5 月 19 日宣布兩款重要 AI 模型:Gemini Omni(可從文字、圖片、音訊、影片等多種輸入生成最高 10 秒的影片)與 Gemini 3.5 Flash(首款結合邊疆級智慧與自主行動能力的模型)。 Gemini Omni Flash 已於 6 月底進入公開預覽版,能在 YouTube Shorts 免費使用,並支援對話式多輪編輯與角色一致性。
Gemini 3.5 Flash 輸出速度比競爭對手快 4 倍,API 定價為每百萬 tokens 輸入 1.50 美元、輸出 9.00 美元,比上一代 Gemini 3.1 Pro 便宜約 40%。