Gemini 3.5 Transcribe is Google’s new speech to text model and a faster successor to Chirp 3: coverage citing Artificial Analysis reports a 70% reduction in time to final transcription, with 4.0% streaming and 2.6% no... It is designed to turn imperfect speech into usable prose by removing filler words, incorporatin...
Published byEdited with GPT-5.6 LunaImages generated with GPT Image 1.5
Research answer

Create a landscape editorial hero image for this Studio Global article: What did Google announce about Gemini 3.5 Transcribe, including how it improves on the previous Chirp 3 speech-to-text engine in accuracy an. Article summary: Google introduced Gemini 3.5 Transcribe as its most precise speech-to-text model, positioning it as a faster, more intelligent successor to Chirp 3 that converts spoken input into polished, structured text rather than a . Topic tags: general, general web, documentation. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fak
Google has introduced Gemini 3.5 Transcribe as a speech-to-text model designed to do more than reproduce audio word for word. It is intended to convert informal, interrupted speech into polished and structured text, making it better suited to dictation, notes, messages, and voice-controlled editing than a purely literal transcript. Google describes it as its most precise speech-to-text model yet and positions it as a major step beyond Chirp 3. 1
12
The change is useful—but also important to understand. A system that removes “um,” resolves corrections, and infers what a speaker meant can produce better writing while preserving less of the speaker’s exact wording.
Gemini 3.5 Transcribe is built around what Google calls intelligent transcription. Instead of treating every hesitation or abandoned phrase as text that must remain in the final output, it can clean up disfluencies, incorporate a speaker’s correction, add punctuation, and apply formatting. It can also respond to voice editing instructions and turn unstructured speech into more readable prose. 1
8
12
That makes the model particularly relevant for everyday voice input. Someone can speak in rough notes or conversational fragments and receive text that is closer to a finished message, document, or form response.
The model is also designed to handle background noise, technical terminology, and specialized vocabulary. Google says users can supply custom vocabulary so names, products, and domain-specific terms are more likely to be rendered with the intended spelling. It automatically detects more than 85 languages. 1
7
For pre-recorded audio, Google lists timestamps and speaker diarization for up to three speakers, giving developers a way to connect portions of a transcript with individual participants. 1
Google presents Gemini 3.5 Transcribe as a substantial improvement over its previous Chirp 3 transcription model. Reporting that cites Artificial Analysis measurements lists a 2.6% word-error rate for non-streaming audio, a 4.0% rate for streaming audio, and a 70% reduction in the time required to reach a final transcript. 3
4
9
Those figures should not be treated as universal performance guarantees. Word-error rate depends on the language, accent, recording quality, background noise, and evaluation set. Google’s own materials highlight results from the FLEURS benchmark, including a 5.50% word-error rate in streaming mode, which is not directly comparable with every third-party measurement because the test conditions differ. 1
The practical takeaway is clearer than any single benchmark: Google is targeting both lower latency for live interactions and cleaner final output for recorded or dictated speech.
Google describes separate ways to use the model depending on the application.
The live option is intended for applications that need a continuous audio stream and rapid responses. Google’s Live API supports low-latency, real-time voice interactions and can provide transcriptions of both user input and model output. 12
20
This type of integration could support voice assistants, live captions, conversational interfaces, and applications that need text while a person is still speaking. The API reference also distinguishes unary, streaming, and real-time interaction patterns for Gemini applications. 24
For recorded audio, developers can upload or reference an audio file and submit it through Gemini’s API interaction flow. This path is better suited to processing an existing recording, where timestamps, formatting, custom vocabulary, and speaker attribution are more important than sub-second response times. 1
12
18
The broader Gemini API documentation recommends the Interactions API as the standard interface for new Gemini applications, while the Live API is intended for continuous, low-latency voice and vision experiences. 19
20
Gemini 3.5 Transcribe is not only a developer preview. Reporting says it already powers Rambler, the voice-dictation feature in Android’s Gboard, and the Gemini app on macOS. 2
13
26
Google also says speech-to-text support is coming to Chrome. The planned capability would let users dictate into web text fields, including messages, posts, forms, and prompts, rather than limiting voice input to a dedicated app. 1
8
13
The model arrives alongside Gemini 3.5 Live and Gemini 3.5 Live Experimental as part of Google’s broader Gemini Audio grouping. In that lineup, Transcribe is focused on turning speech into text, while the Live models are aimed at interactive audio experiences. 12
26
Gemini 3.5 Transcribe’s strongest feature is also its main limitation. Removing filler words and integrating self-corrections makes output easier to read, but it changes the relationship between the audio and the transcript.
A polished result may omit meaningful hesitation, preserve only the corrected version of a statement, or smooth away an individual’s speaking style. That is often desirable when dictating a message or drafting notes. It is less desirable when the exact wording matters.
For interviews, legal or medical records, research, accessibility documentation, or quotations, users should retain the original audio and verify important passages. Unless an implementation offers a clearly verbatim mode, Gemini 3.5 Transcribe should be understood as an intelligent rewriting transcription system—not simply a neutral audio-to-text recorder.
Google’s announcement moves speech recognition closer to voice-assisted writing. Gemini 3.5 Transcribe combines transcription with cleanup, formatting, language detection, vocabulary customization, and speaker information, while offering separate workflows for live and recorded audio. 1
12
For developers, that could reduce the amount of post-processing required after speech is recognized. For users, it could make dictation feel less like speaking carefully to a machine and more like giving a rough draft to an editor.
The remaining question is not whether the model can produce cleaner text. It is whether each application needs clean text—or an exact record of what was said.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
Gemini 3.5 Transcribe is Google’s new speech to text model and a faster successor to Chirp 3: coverage citing Artificial Analysis reports a 70% reduction in time to final transcription, with 4.0% streaming and 2.6% no...
Gemini 3.5 Transcribe is Google’s new speech to text model and a faster successor to Chirp 3: coverage citing Artificial Analysis reports a 70% reduction in time to final transcription, with 4.0% streaming and 2.6% no... It is designed to turn imperfect speech into usable prose by removing filler words, incorporating self corrections, applying formatting, recognizing custom vocabulary, and detecting more than 85 languages.
Developers can build with separate real time and pre recorded audio workflows, while the technology is already appearing in Gboard’s Rambler dictation feature and the Gemini app for macOS, with Chrome support planned.
Gemini 3.5 Transcribe is Google’s new speech to text model and a faster successor to Chirp 3: coverage citing Artificial Analysis reports a 70% reduction in time to final transcription, with 4.0% streaming and 2.6% no... It is designed to turn imperfect speech into usable prose by removing filler words, incorporatin...
Published byEdited with GPT-5.6 LunaImages generated with GPT Image 1.5
Research answer

Create a landscape editorial hero image for this Studio Global article: What did Google announce about Gemini 3.5 Transcribe, including how it improves on the previous Chirp 3 speech-to-text engine in accuracy an. Article summary: Google introduced Gemini 3.5 Transcribe as its most precise speech-to-text model, positioning it as a faster, more intelligent successor to Chirp 3 that converts spoken input into polished, structured text rather than a . Topic tags: general, general web, documentation. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fak
Google has introduced Gemini 3.5 Transcribe as a speech-to-text model designed to do more than reproduce audio word for word. It is intended to convert informal, interrupted speech into polished and structured text, making it better suited to dictation, notes, messages, and voice-controlled editing than a purely literal transcript. Google describes it as its most precise speech-to-text model yet and positions it as a major step beyond Chirp 3. 1
12
The change is useful—but also important to understand. A system that removes “um,” resolves corrections, and infers what a speaker meant can produce better writing while preserving less of the speaker’s exact wording.
Gemini 3.5 Transcribe is built around what Google calls intelligent transcription. Instead of treating every hesitation or abandoned phrase as text that must remain in the final output, it can clean up disfluencies, incorporate a speaker’s correction, add punctuation, and apply formatting. It can also respond to voice editing instructions and turn unstructured speech into more readable prose. 1
8
12
That makes the model particularly relevant for everyday voice input. Someone can speak in rough notes or conversational fragments and receive text that is closer to a finished message, document, or form response.
The model is also designed to handle background noise, technical terminology, and specialized vocabulary. Google says users can supply custom vocabulary so names, products, and domain-specific terms are more likely to be rendered with the intended spelling. It automatically detects more than 85 languages. 1
7
For pre-recorded audio, Google lists timestamps and speaker diarization for up to three speakers, giving developers a way to connect portions of a transcript with individual participants. 1
Google presents Gemini 3.5 Transcribe as a substantial improvement over its previous Chirp 3 transcription model. Reporting that cites Artificial Analysis measurements lists a 2.6% word-error rate for non-streaming audio, a 4.0% rate for streaming audio, and a 70% reduction in the time required to reach a final transcript. 3
4
9
Those figures should not be treated as universal performance guarantees. Word-error rate depends on the language, accent, recording quality, background noise, and evaluation set. Google’s own materials highlight results from the FLEURS benchmark, including a 5.50% word-error rate in streaming mode, which is not directly comparable with every third-party measurement because the test conditions differ. 1
The practical takeaway is clearer than any single benchmark: Google is targeting both lower latency for live interactions and cleaner final output for recorded or dictated speech.
Google describes separate ways to use the model depending on the application.
The live option is intended for applications that need a continuous audio stream and rapid responses. Google’s Live API supports low-latency, real-time voice interactions and can provide transcriptions of both user input and model output. 12
20
This type of integration could support voice assistants, live captions, conversational interfaces, and applications that need text while a person is still speaking. The API reference also distinguishes unary, streaming, and real-time interaction patterns for Gemini applications. 24
For recorded audio, developers can upload or reference an audio file and submit it through Gemini’s API interaction flow. This path is better suited to processing an existing recording, where timestamps, formatting, custom vocabulary, and speaker attribution are more important than sub-second response times. 1
12
18
The broader Gemini API documentation recommends the Interactions API as the standard interface for new Gemini applications, while the Live API is intended for continuous, low-latency voice and vision experiences. 19
20
Gemini 3.5 Transcribe is not only a developer preview. Reporting says it already powers Rambler, the voice-dictation feature in Android’s Gboard, and the Gemini app on macOS. 2
13
26
Google also says speech-to-text support is coming to Chrome. The planned capability would let users dictate into web text fields, including messages, posts, forms, and prompts, rather than limiting voice input to a dedicated app. 1
8
13
The model arrives alongside Gemini 3.5 Live and Gemini 3.5 Live Experimental as part of Google’s broader Gemini Audio grouping. In that lineup, Transcribe is focused on turning speech into text, while the Live models are aimed at interactive audio experiences. 12
26
Gemini 3.5 Transcribe’s strongest feature is also its main limitation. Removing filler words and integrating self-corrections makes output easier to read, but it changes the relationship between the audio and the transcript.
A polished result may omit meaningful hesitation, preserve only the corrected version of a statement, or smooth away an individual’s speaking style. That is often desirable when dictating a message or drafting notes. It is less desirable when the exact wording matters.
For interviews, legal or medical records, research, accessibility documentation, or quotations, users should retain the original audio and verify important passages. Unless an implementation offers a clearly verbatim mode, Gemini 3.5 Transcribe should be understood as an intelligent rewriting transcription system—not simply a neutral audio-to-text recorder.
Google’s announcement moves speech recognition closer to voice-assisted writing. Gemini 3.5 Transcribe combines transcription with cleanup, formatting, language detection, vocabulary customization, and speaker information, while offering separate workflows for live and recorded audio. 1
12
For developers, that could reduce the amount of post-processing required after speech is recognized. For users, it could make dictation feel less like speaking carefully to a machine and more like giving a rough draft to an editor.
The remaining question is not whether the model can produce cleaner text. It is whether each application needs clean text—or an exact record of what was said.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
Gemini 3.5 Transcribe is Google’s new speech to text model and a faster successor to Chirp 3: coverage citing Artificial Analysis reports a 70% reduction in time to final transcription, with 4.0% streaming and 2.6% no...
Gemini 3.5 Transcribe is Google’s new speech to text model and a faster successor to Chirp 3: coverage citing Artificial Analysis reports a 70% reduction in time to final transcription, with 4.0% streaming and 2.6% no... It is designed to turn imperfect speech into usable prose by removing filler words, incorporating self corrections, applying formatting, recognizing custom vocabulary, and detecting more than 85 languages.
Developers can build with separate real time and pre recorded audio workflows, while the technology is already appearing in Gboard’s Rambler dictation feature and the Gemini app for macOS, with Chrome support planned.