Gemini 3.5 Transcribe Google का नया speech to text मॉडल है। Artificial Analysis के आंकड़ों के अनुसार, यह Chirp 3 की तुलना में अंतिम ट्रांसक्रिप्ट तैयार करने का समय 70% घटाता है; word error rate स्ट्रीमिंग में 4.0% और... यह मॉडल रुक रुककर बोलने, filler words हटाने, self corrections को समझने, formatting लागू करने, cus...
प्रकाशितकर्ताGPT-5.6 Luna से संपादितGPT Image 1.5 से चित्र बनाए गए
शोध उत्तर

Create a landscape editorial hero image for this Studio Global article: What did Google announce about Gemini 3.5 Transcribe, including how it improves on the previous Chirp 3 speech-to-text engine in accuracy an. Article summary: Google introduced Gemini 3.5 Transcribe as its most precise speech-to-text model, positioning it as a faster, more intelligent successor to Chirp 3 that converts spoken input into polished, structured text rather than a . Topic tags: general, general web, documentation. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fak
Google ने Gemini 3.5 Transcribe पेश किया है—ऐसा speech-to-text मॉडल जो ऑडियो को केवल शब्द-दर-शब्द लिखने के बजाय उसे अधिक साफ, व्यवस्थित और उपयोगी टेक्स्ट में बदलने की कोशिश करता है। कंपनी इसे अब तक का अपना सबसे सटीक speech-to-text मॉडल बताती है और इसे पिछले Chirp 3 मॉडल से बड़ा सुधार मानती है। 1
12
इसका फायदा dictation, नोट्स, संदेश लिखने और voice-controlled editing जैसे कामों में दिख सकता है। लेकिन एक अहम बात समझना जरूरी है: जो सिस्टम “उम्”, “आ…” जैसी झिझक हटाता है, अधूरे वाक्य ठीक करता है और वक्ता का आशय समझने की कोशिश करता है, वह बेहतर लेखन तो दे सकता है—पर हर शब्द का बिल्कुल वैसा रिकॉर्ड नहीं रखेगा।
Gemini 3.5 Transcribe को Google के शब्दों में “intelligent transcription” के लिए बनाया गया है। यानी यह हर hesitation या अधूरे वाक्य को अंतिम टेक्स्ट में बनाए रखने के बजाय filler words हटा सकता है, वक्ता की correction को शामिल कर सकता है, punctuation और formatting जोड़ सकता है और बिखरी हुई बातचीत को अधिक पठनीय गद्य में बदल सकता है। मॉडल voice editing instructions पर भी प्रतिक्रिया दे सकता है। 1
8
12
इसका मतलब है कि उपयोगकर्ता बिना पूरी तरह सधे हुए बोले भी rough notes या बातचीत जैसे fragments दे सकता है और बदले में संदेश, दस्तावेज़ या form response के करीब का टेक्स्ट पा सकता है।
मॉडल background noise, technical terminology और specialized vocabulary को संभालने के लिए भी तैयार किया गया है। Google के अनुसार, उपयोगकर्ता custom vocabulary दे सकते हैं, जिससे नामों, उत्पादों और किसी खास क्षेत्र के शब्दों की spelling अधिक सही रहने की संभावना बढ़ती है। यह 85 से अधिक भाषाओं का स्वतः पता लगा सकता है। 1
7
Pre-recorded ऑडियो के लिए Google timestamps और अधिकतम तीन speakers की speaker diarization उपलब्ध कराता है। इससे डेवलपर ट्रांसक्रिप्ट के हिस्सों को अलग-अलग वक्ताओं से जोड़ सकते हैं। 1
Google Gemini 3.5 Transcribe को अपने पिछले Chirp 3 transcription मॉडल से काफी बेहतर बता रहा है। Artificial Analysis के मापों का हवाला देने वाली रिपोर्टों के अनुसार, non-streaming ऑडियो में word-error rate 2.6%, streaming ऑडियो में 4.0% और final transcript तैयार होने में लगने वाले समय में 70% कमी दर्ज की गई है। 3
4
9
इन आंकड़ों को हर भाषा, accent या recording condition के लिए सार्वभौमिक गारंटी नहीं माना जाना चाहिए। Word-error rate भाषा, उच्चारण, ऑडियो की गुणवत्ता, background noise और evaluation set के आधार पर बदल सकता है। Google की अपनी सामग्री FLEURS benchmark पर streaming mode में 5.50% word-error rate बताती है। अलग-अलग test conditions होने के कारण इसे हर third-party measurement से सीधे तुलना नहीं किया जा सकता। 1
फिर भी व्यावहारिक तस्वीर स्पष्ट है: Google live बातचीत के लिए कम latency और recorded या dictated speech के लिए अधिक साफ final output—दोनों पर ध्यान दे रहा है।
Google application के प्रकार के आधार पर मॉडल इस्तेमाल करने के अलग तरीके बताता है।
Live विकल्प उन ऐप्स के लिए है जिन्हें लगातार आती audio stream पर तेज प्रतिक्रिया देनी होती है। Gemini Live API low-latency, real-time voice interactions को support करता है और user input तथा model output—दोनों की transcription उपलब्ध करा सकता है। 12
20
इससे voice assistants, live captions, conversational interfaces और ऐसे ऐप बनाए जा सकते हैं जिन्हें व्यक्ति के बोलते समय ही टेक्स्ट चाहिए। Gemini API reference unary, streaming और real-time interaction patterns के बीच भी अंतर बताता है। 24
Recorded audio के लिए डेवलपर audio file upload या reference करके उसे Gemini API interaction flow के जरिए process कर सकते हैं। यह तरीका पहले से मौजूद recording के लिए अधिक उपयुक्त है, जहां timestamps, formatting, custom vocabulary और speaker attribution, sub-second response time से ज्यादा महत्वपूर्ण हो सकते हैं। 1
12
18
Google की व्यापक Gemini API documentation नए Gemini applications के लिए Interactions API को standard interface बताती है। वहीं लगातार चलने वाले, low-latency voice और vision experiences के लिए Live API का इस्तेमाल किया जाता है। 19
20
Gemini 3.5 Transcribe अभी केवल developer preview तक सीमित नहीं है। रिपोर्टों के अनुसार, यह Android के Gboard में Rambler voice-dictation फीचर और macOS के Gemini ऐप को power कर रहा है। 2
13
26
Google ने यह भी कहा है कि speech-to-text सुविधा Chrome में आने वाली है। इसके बाद उपयोगकर्ता किसी dedicated ऐप तक सीमित हुए बिना web text fields में बोलकर लिख सकेंगे—जैसे messages, posts, forms और prompts। 1
8
13
यह मॉडल Gemini 3.5 Live और Gemini 3.5 Live Experimental के साथ Google के व्यापक “Gemini Audio” समूह का हिस्सा है। इस lineup में Transcribe का मुख्य काम speech को text में बदलना है, जबकि Live मॉडल interactive audio experiences पर केंद्रित हैं। 12
26
Gemini 3.5 Transcribe की सबसे मजबूत खूबी ही उसकी सबसे बड़ी सीमा भी है। Filler words हटाने और self-corrections को एकीकृत करने से output पढ़ने में आसान हो जाता है, लेकिन audio और transcript के बीच का रिश्ता बदल जाता है।
साफ-सुथरा परिणाम किसी वक्ता की महत्वपूर्ण झिझक को हटा सकता है, किसी statement का केवल corrected version रख सकता है या बोलने की व्यक्तिगत शैली को कम कर सकता है। संदेश लिखने, notes बनाने या form भरने के लिए यह अक्सर उपयोगी है। लेकिन जहां exact wording मायने रखती है, वहां यही व्यवहार समस्या बन सकता है।
Interview, कानूनी या medical records, research transcripts, accessibility documentation और direct quotations के लिए original audio सुरक्षित रखना और महत्वपूर्ण अंशों की पुष्टि करना बेहतर होगा। जब तक किसी implementation में स्पष्ट verbatim mode उपलब्ध न हो, Gemini 3.5 Transcribe को neutral audio-to-text recorder नहीं, बल्कि intelligent rewriting transcription system समझना चाहिए।
Google speech recognition को voice-assisted writing के और करीब ले जा रहा है। Gemini 3.5 Transcribe transcription के साथ cleanup, formatting, language detection, vocabulary customization और speaker information को जोड़ता है, जबकि live और recorded audio के लिए अलग workflows उपलब्ध कराता है। 1
12
डेवलपर के लिए इसका अर्थ है कि speech recognition के बाद होने वाले post-processing का काम कम हो सकता है। उपयोगकर्ताओं के लिए dictation ऐसा महसूस हो सकता है जैसे वे मशीन को सावधानी से बोलने के बजाय किसी editor को rough draft दे रहे हों।
लेकिन अंतिम सवाल यह नहीं है कि मॉडल कितना साफ टेक्स्ट बना सकता है। असली सवाल यह है कि किसी application को साफ टेक्स्ट चाहिए—या बोले गए हर शब्द का सटीक रिकॉर्ड।
Studio Global AI
इस पृष्ठ में एक स्रोत-समर्थित उत्तर शामिल है जिसे आप Studio Global के अंदर जारी रख सकते हैं।
Gemini 3.5 Transcribe Google का नया speech to text मॉडल है। Artificial Analysis के आंकड़ों के अनुसार, यह Chirp 3 की तुलना में अंतिम ट्रांसक्रिप्ट तैयार करने का समय 70% घटाता है; word error rate स्ट्रीमिंग में 4.0% और...
Gemini 3.5 Transcribe Google का नया speech to text मॉडल है। Artificial Analysis के आंकड़ों के अनुसार, यह Chirp 3 की तुलना में अंतिम ट्रांसक्रिप्ट तैयार करने का समय 70% घटाता है; word error rate स्ट्रीमिंग में 4.0% और... यह मॉडल रुक रुककर बोलने, filler words हटाने, self corrections को समझने, formatting लागू करने, custom vocabulary इस्तेमाल करने और 85 से अधिक भाषाओं का स्वतः पता लगाने के लिए बनाया गया है।
डेवलपर real time और pre recorded ऑडियो के लिए अलग workflows अपना सकते हैं। यह तकनीक Android Gboard के Rambler dictation फीचर और macOS के Gemini ऐप में इस्तेमाल हो रही है; Chrome सपोर्ट जल्द आने वाला है।
Gemini 3.5 Transcribe Google का नया speech to text मॉडल है। Artificial Analysis के आंकड़ों के अनुसार, यह Chirp 3 की तुलना में अंतिम ट्रांसक्रिप्ट तैयार करने का समय 70% घटाता है; word error rate स्ट्रीमिंग में 4.0% और... यह मॉडल रुक रुककर बोलने, filler words हटाने, self corrections को समझने, formatting लागू करने, cus...
प्रकाशितकर्ताGPT-5.6 Luna से संपादितGPT Image 1.5 से चित्र बनाए गए
शोध उत्तर

Create a landscape editorial hero image for this Studio Global article: What did Google announce about Gemini 3.5 Transcribe, including how it improves on the previous Chirp 3 speech-to-text engine in accuracy an. Article summary: Google introduced Gemini 3.5 Transcribe as its most precise speech-to-text model, positioning it as a faster, more intelligent successor to Chirp 3 that converts spoken input into polished, structured text rather than a . Topic tags: general, general web, documentation. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fak
Google ने Gemini 3.5 Transcribe पेश किया है—ऐसा speech-to-text मॉडल जो ऑडियो को केवल शब्द-दर-शब्द लिखने के बजाय उसे अधिक साफ, व्यवस्थित और उपयोगी टेक्स्ट में बदलने की कोशिश करता है। कंपनी इसे अब तक का अपना सबसे सटीक speech-to-text मॉडल बताती है और इसे पिछले Chirp 3 मॉडल से बड़ा सुधार मानती है। 1
12
इसका फायदा dictation, नोट्स, संदेश लिखने और voice-controlled editing जैसे कामों में दिख सकता है। लेकिन एक अहम बात समझना जरूरी है: जो सिस्टम “उम्”, “आ…” जैसी झिझक हटाता है, अधूरे वाक्य ठीक करता है और वक्ता का आशय समझने की कोशिश करता है, वह बेहतर लेखन तो दे सकता है—पर हर शब्द का बिल्कुल वैसा रिकॉर्ड नहीं रखेगा।
Gemini 3.5 Transcribe को Google के शब्दों में “intelligent transcription” के लिए बनाया गया है। यानी यह हर hesitation या अधूरे वाक्य को अंतिम टेक्स्ट में बनाए रखने के बजाय filler words हटा सकता है, वक्ता की correction को शामिल कर सकता है, punctuation और formatting जोड़ सकता है और बिखरी हुई बातचीत को अधिक पठनीय गद्य में बदल सकता है। मॉडल voice editing instructions पर भी प्रतिक्रिया दे सकता है। 1
8
12
इसका मतलब है कि उपयोगकर्ता बिना पूरी तरह सधे हुए बोले भी rough notes या बातचीत जैसे fragments दे सकता है और बदले में संदेश, दस्तावेज़ या form response के करीब का टेक्स्ट पा सकता है।
मॉडल background noise, technical terminology और specialized vocabulary को संभालने के लिए भी तैयार किया गया है। Google के अनुसार, उपयोगकर्ता custom vocabulary दे सकते हैं, जिससे नामों, उत्पादों और किसी खास क्षेत्र के शब्दों की spelling अधिक सही रहने की संभावना बढ़ती है। यह 85 से अधिक भाषाओं का स्वतः पता लगा सकता है। 1
7
Pre-recorded ऑडियो के लिए Google timestamps और अधिकतम तीन speakers की speaker diarization उपलब्ध कराता है। इससे डेवलपर ट्रांसक्रिप्ट के हिस्सों को अलग-अलग वक्ताओं से जोड़ सकते हैं। 1
Google Gemini 3.5 Transcribe को अपने पिछले Chirp 3 transcription मॉडल से काफी बेहतर बता रहा है। Artificial Analysis के मापों का हवाला देने वाली रिपोर्टों के अनुसार, non-streaming ऑडियो में word-error rate 2.6%, streaming ऑडियो में 4.0% और final transcript तैयार होने में लगने वाले समय में 70% कमी दर्ज की गई है। 3
4
9
इन आंकड़ों को हर भाषा, accent या recording condition के लिए सार्वभौमिक गारंटी नहीं माना जाना चाहिए। Word-error rate भाषा, उच्चारण, ऑडियो की गुणवत्ता, background noise और evaluation set के आधार पर बदल सकता है। Google की अपनी सामग्री FLEURS benchmark पर streaming mode में 5.50% word-error rate बताती है। अलग-अलग test conditions होने के कारण इसे हर third-party measurement से सीधे तुलना नहीं किया जा सकता। 1
फिर भी व्यावहारिक तस्वीर स्पष्ट है: Google live बातचीत के लिए कम latency और recorded या dictated speech के लिए अधिक साफ final output—दोनों पर ध्यान दे रहा है।
Google application के प्रकार के आधार पर मॉडल इस्तेमाल करने के अलग तरीके बताता है।
Live विकल्प उन ऐप्स के लिए है जिन्हें लगातार आती audio stream पर तेज प्रतिक्रिया देनी होती है। Gemini Live API low-latency, real-time voice interactions को support करता है और user input तथा model output—दोनों की transcription उपलब्ध करा सकता है। 12
20
इससे voice assistants, live captions, conversational interfaces और ऐसे ऐप बनाए जा सकते हैं जिन्हें व्यक्ति के बोलते समय ही टेक्स्ट चाहिए। Gemini API reference unary, streaming और real-time interaction patterns के बीच भी अंतर बताता है। 24
Recorded audio के लिए डेवलपर audio file upload या reference करके उसे Gemini API interaction flow के जरिए process कर सकते हैं। यह तरीका पहले से मौजूद recording के लिए अधिक उपयुक्त है, जहां timestamps, formatting, custom vocabulary और speaker attribution, sub-second response time से ज्यादा महत्वपूर्ण हो सकते हैं। 1
12
18
Google की व्यापक Gemini API documentation नए Gemini applications के लिए Interactions API को standard interface बताती है। वहीं लगातार चलने वाले, low-latency voice और vision experiences के लिए Live API का इस्तेमाल किया जाता है। 19
20
Gemini 3.5 Transcribe अभी केवल developer preview तक सीमित नहीं है। रिपोर्टों के अनुसार, यह Android के Gboard में Rambler voice-dictation फीचर और macOS के Gemini ऐप को power कर रहा है। 2
13
26
Google ने यह भी कहा है कि speech-to-text सुविधा Chrome में आने वाली है। इसके बाद उपयोगकर्ता किसी dedicated ऐप तक सीमित हुए बिना web text fields में बोलकर लिख सकेंगे—जैसे messages, posts, forms और prompts। 1
8
13
यह मॉडल Gemini 3.5 Live और Gemini 3.5 Live Experimental के साथ Google के व्यापक “Gemini Audio” समूह का हिस्सा है। इस lineup में Transcribe का मुख्य काम speech को text में बदलना है, जबकि Live मॉडल interactive audio experiences पर केंद्रित हैं। 12
26
Gemini 3.5 Transcribe की सबसे मजबूत खूबी ही उसकी सबसे बड़ी सीमा भी है। Filler words हटाने और self-corrections को एकीकृत करने से output पढ़ने में आसान हो जाता है, लेकिन audio और transcript के बीच का रिश्ता बदल जाता है।
साफ-सुथरा परिणाम किसी वक्ता की महत्वपूर्ण झिझक को हटा सकता है, किसी statement का केवल corrected version रख सकता है या बोलने की व्यक्तिगत शैली को कम कर सकता है। संदेश लिखने, notes बनाने या form भरने के लिए यह अक्सर उपयोगी है। लेकिन जहां exact wording मायने रखती है, वहां यही व्यवहार समस्या बन सकता है।
Interview, कानूनी या medical records, research transcripts, accessibility documentation और direct quotations के लिए original audio सुरक्षित रखना और महत्वपूर्ण अंशों की पुष्टि करना बेहतर होगा। जब तक किसी implementation में स्पष्ट verbatim mode उपलब्ध न हो, Gemini 3.5 Transcribe को neutral audio-to-text recorder नहीं, बल्कि intelligent rewriting transcription system समझना चाहिए।
Google speech recognition को voice-assisted writing के और करीब ले जा रहा है। Gemini 3.5 Transcribe transcription के साथ cleanup, formatting, language detection, vocabulary customization और speaker information को जोड़ता है, जबकि live और recorded audio के लिए अलग workflows उपलब्ध कराता है। 1
12
डेवलपर के लिए इसका अर्थ है कि speech recognition के बाद होने वाले post-processing का काम कम हो सकता है। उपयोगकर्ताओं के लिए dictation ऐसा महसूस हो सकता है जैसे वे मशीन को सावधानी से बोलने के बजाय किसी editor को rough draft दे रहे हों।
लेकिन अंतिम सवाल यह नहीं है कि मॉडल कितना साफ टेक्स्ट बना सकता है। असली सवाल यह है कि किसी application को साफ टेक्स्ट चाहिए—या बोले गए हर शब्द का सटीक रिकॉर्ड।
Studio Global AI
इस पृष्ठ में एक स्रोत-समर्थित उत्तर शामिल है जिसे आप Studio Global के अंदर जारी रख सकते हैं।
Gemini 3.5 Transcribe Google का नया speech to text मॉडल है। Artificial Analysis के आंकड़ों के अनुसार, यह Chirp 3 की तुलना में अंतिम ट्रांसक्रिप्ट तैयार करने का समय 70% घटाता है; word error rate स्ट्रीमिंग में 4.0% और...
Gemini 3.5 Transcribe Google का नया speech to text मॉडल है। Artificial Analysis के आंकड़ों के अनुसार, यह Chirp 3 की तुलना में अंतिम ट्रांसक्रिप्ट तैयार करने का समय 70% घटाता है; word error rate स्ट्रीमिंग में 4.0% और... यह मॉडल रुक रुककर बोलने, filler words हटाने, self corrections को समझने, formatting लागू करने, custom vocabulary इस्तेमाल करने और 85 से अधिक भाषाओं का स्वतः पता लगाने के लिए बनाया गया है।
डेवलपर real time और pre recorded ऑडियो के लिए अलग workflows अपना सकते हैं। यह तकनीक Android Gboard के Rambler dictation फीचर और macOS के Gemini ऐप में इस्तेमाल हो रही है; Chrome सपोर्ट जल्द आने वाला है।