Google’s Gemini 3.8 Live is the scale and cost focused voice model, while Gemini 3.8 Live Extended Thinking is designed to reason and make asynchronous tool calls in the background while it keeps speaking. The practical shift is not just faster voice responses: Google is pitching an agent that can sustain a live con...
Published byEdited with GPT-5.6 TerraImages generated with GPT Image 2
Research answer

Create a landscape editorial hero image for this Studio Global article: What did Google announce about Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, including their real-time conversational, visual-under. Article summary: Google introduced two generally available, native audio-to-audio models for real-time voice agents: Gemini 3.8 Live for lower-cost, high-volume dialogue, and Gemini 3.8 Live Extended Thinking for more difficult, multi-st. Topic tags: general, general web, user generated, documentation. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks,
Google has introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two native audio-to-audio models for real-time voice agents and live dialogue. The standard model is aimed at scale, cost efficiency, fluid dialogue, and visual grounding; the Extended Thinking variant targets harder, multi-step tasks that benefit from background reasoning. 10
2
The defining feature belongs to Gemini 3.8 Live Extended Thinking. Google says the model can reason in the background and plan or call asynchronous tools while continuing to stream an audio response. It may use natural conversational fillers while a longer-running tool completes, rather than holding the entire conversation until a result arrives. 1
2
That makes “simultaneous reasoning” more than a claim about quick turn-taking. In Google’s design, speech generation, background planning, and asynchronous tool activity can proceed in parallel. A voice agent could therefore continue explaining a next step while it waits for data or works through a multi-step process. Whether that produces a better experience will depend on tool reliability, the agent’s dialogue design, and whether the speech remains useful rather than merely filling silence. 1
2
Google positions the standard Gemini 3.8 Live as the choice for lower-latency, high-volume conversational experiences without reasoning-induced delays. It combines live dialogue with visual grounding, allowing an agent to use visual context alongside the conversation. 10
45
Google representatives have also said the model supports 97 languages and can switch between languages in a conversation. Developers should confirm language quality, regional availability, and performance for their particular accents and use cases during evaluation. 16
Launch coverage reported Speech-to-Speech Quality Index scores of 76.0 for Gemini 3.8 Live and 82.6 for Gemini 3.8 Live Extended Thinking, compared with 81.5 for OpenAI’s GPT-Live-1 using medium effort. 25
Those figures make Extended Thinking the higher-scoring of the three in that reported comparison, but they should not be interpreted as an independently audited measure of overall production quality. A voice-agent deployment must also handle interruptions, multilingual and accented speech, tool accuracy, latency under load, safety, observability, and total cost per successfully resolved task. The reported numbers are one evaluation signal, not proof that one system is best for every contact-center or assistant workflow. 25
Google’s official pricing table groups both new models under the Live API. On the paid standard tier, it lists audio input at $3.00 per million tokens, or $0.005 per minute, and audio output at $12.00 per million tokens, or $0.018 per minute. Text input is listed at $0.75 per million tokens and text output, including thinking tokens, at $4.50 per million tokens. 37
That distinction matters: token prices cited elsewhere for Gemini 3.8 Flash are not automatically the prices for the Live models. For budgeting, use the current Live API table and model documentation, then test real sessions, because a voice agent’s cost depends on its mix of input audio, output audio, visual context, text, and any associated tools. 37
Google DeepMind lists both models as generally available. The published availability table includes the Gemini app, Google AI Studio, the Gemini API, and Google Enterprise Agent Platform for both models. It lists Google Search Live for Extended Thinking and Google Workspace for Gemini 3.8 Live. 54
For developers, this creates a choice between a standard live-dialogue model and a higher-reasoning model within the same native audio-to-audio family. For enterprises, the key question is not only which model is available, but which product surface, data controls, geographic region, and integration path apply to the intended deployment. 54
The central competitive proposition is a more integrated real-time stack: spoken interaction, visual grounding, background reasoning, and asynchronous tools in one model family. In principle, that can reduce the amount of custom stitching between speech recognition, a text model, tool orchestration, and speech synthesis—and can preserve conversational continuity while work is underway. 1
10
OpenAI’s GPT-Live-1 remains a relevant point of comparison, particularly for teams evaluating full-duplex conversational behavior. But the reported benchmark comparison is too narrow to settle a platform decision. A credible evaluation should use representative calls and measure interruption handling, tool-call correctness, multilingual performance, safety controls, recovery from failures, latency, and cost per completed outcome. 25
The supplied launch material does not establish a precise SynthID audio-watermarking policy for Gemini 3.8 Live or Extended Thinking. Google says SynthID has expanded to watermarking and identifying text generated by the Gemini app and web experience, but that statement does not confirm that every audio stream from these Live models is watermarked, nor does it specify detection or rollout terms. 59
The same caution applies to ecosystem claims: integration announcements and partner support can change quickly. Teams planning a production deployment should verify current model documentation, pricing, platform availability, and applicable service terms before committing to an architecture. 37
54
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
Google’s Gemini 3.8 Live is the scale and cost focused voice model, while Gemini 3.8 Live Extended Thinking is designed to reason and make asynchronous tool calls in the background while it keeps speaking.
Google’s Gemini 3.8 Live is the scale and cost focused voice model, while Gemini 3.8 Live Extended Thinking is designed to reason and make asynchronous tool calls in the background while it keeps speaking. The practical shift is not just faster voice responses: Google is pitching an agent that can sustain a live conversation while it plans, waits on tools, and processes results.