Released September 23, 2026, Spark ASR 2.0 is claimed to reduce transcription errors and produce more fluent text than Spark ASR 1.0 for roughly 10% higher inference cost. Its reported improvements span mixed Chinese–English speech, dialects, jargon, noise, low volume or fast speech, and children’s voices; a definit...
Published byEdited with GPT-6 SolImages generated with GPT Image 2
Research answer

Create a landscape editorial hero image for this Studio Global article: What is iFLYTEK Spark-ASR-2.0, released on September 23, 2026, and how does it compare with Spark-ASR-1.0 in recognition accuracy, text flue. Article summary: iFLYTEK released Spark-ASR-2.0 on September 23, 2026, as an upgrade to its speech-to-text model. It aims to produce not just a more accurate transcript than Spark-ASR-1.0, but more fluent, readable text; the reported gai. Topic tags: general, general web, user generated. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fa
iFLYTEK released Spark-ASR-2.0 on September 23, 2026, with a goal beyond transcribing speech accurately: producing text that reads more like finished prose. Launch reports describe improvements over Spark-ASR-1.0, but the performance figures and comparisons come from the company rather than an independently reproducible evaluation. 5
6
Recognition accuracy: iFLYTEK reports lower word error rates, particularly for Chinese–English mixed speech, dialects and noisy audio. It also reports improvements for specialist terminology, low-volume or fast speech, and children’s voices. The available launch reporting does not establish a reliable overall percentage-point accuracy gain. 5
6
Text fluency: The new model is intended to use context to trim redundant wording and make transcripts more coherent and readable. That is a different goal from a strictly verbatim transcript, and users who need an exact record should check how the output handles spoken wording. 5
7
Inference cost: iFLYTEK says overall inference cost is about 10% higher than for Spark-ASR-1.0. That is a claim about the cost of running the model, not a published 10% increase in API pricing. 5
6
The launch reporting identifies three approaches: cooperation between non-autoregressive recognition and LLM-enhanced autoregressive processing; joint augmentation of Chinese–English text and audio; and dynamic injection of context. iFLYTEK associates the combined approach with better handling of code-switching, dialects, technical terms, difficult acoustics and transcript formatting. The reporting does not isolate how much each technique contributes to any one improvement. 5
6
19
In particular, the cited difficult-acoustic cases are high noise and low-volume speech, along with fast speech and children’s voices. “Low-volume” should not be mistaken for a claim that the model can transcribe silent audio. 5
6
There is no well-established Spark-ASR-2.0 dialect count in the available evidence. One report says it covers 38 dialect and minority-language variants, while another describes recognition of dialects from 202 cities. Separately, iFLYTEK documentation advertises 202 dialects for a real-time transcription service without establishing that the figure applies to Spark-ASR-2.0. These descriptions should not be combined into a verified specification for the model. 3
6
17
Likewise, assertions that Spark-ASR-2.0 surpasses the industry’s best systems on dialects, noise or low-volume speech are reported company comparisons, not independently established rankings. 5
6
A gradual rollout in iFLYTEK Input Method was scheduled to begin September 24, 2026. Reports also describe API access and a trial entry point through the iFLYTEK Open Platform, with later integrations planned for AI glasses, smart office notebooks and iFLYTEK Tingjian transcription products. Those product plans do not confirm that every integration is already live. 6
9
18
For subsequent development, iFLYTEK has identified three experience targets: rejecting speech from irrelevant speakers, editing recognized text through voice commands, and presenting results with attention to emotion. The cited reporting does not provide a verified release date for those capabilities. 18
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
Released September 23, 2026, Spark ASR 2.0 is claimed to reduce transcription errors and produce more fluent text than Spark ASR 1.0 for roughly 10% higher inference cost.
Released September 23, 2026, Spark ASR 2.0 is claimed to reduce transcription errors and produce more fluent text than Spark ASR 1.0 for roughly 10% higher inference cost. Its reported improvements span mixed Chinese–English speech, dialects, jargon, noise, low volume or fast speech, and children’s voices; a definitive dialect count for this model is not established.