Launched September 28, 2026, Eleven v4 is the quality focused model for expressive narration, while v4 Turbo targets live agents with about 100 ms median inference latency. Both models support more than 90 languages; v4 adds expanded performance controls and brings Professional Voice Clones back, while Turbo is the...
Published byEdited with GPT-6 LunaImages generated with GPT Image 2
Research answer

Create a landscape editorial hero image for this Studio Global article: What are ElevenLabs’ new Eleven v4 and Eleven v4 Turbo text-to-speech models, how do their availability, architecture, voice cloning, langua. Article summary: ElevenLabs launched Eleven v4 and Eleven v4 Turbo on September 28, 2026: v4 prioritizes expressive, produced speech, while Turbo trades some emphasis on maximum quality for the speed needed by live voice agents. Together. Topic tags: general, general web, user generated. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fa
ElevenLabs’ September 28, 2026 launch splits its latest speech generation into two options: Eleven v4 for expressive, produced audio and Eleven v4 Turbo for real-time applications. Both support more than 90 languages and voice cloning, but their intended uses differ—and launch claims about quality and speed are not a substitute for testing with your own content. 5
11
17
| Eleven v4 | Eleven v4 Turbo | |
|---|---|---|
| Designed for | Content creation, audiobooks and character voiceovers where quality is the priority |
Real-time use, including conversational agents and interactive voice experiences |
| Performance emphasis | Expressive delivery and control over a script |
Lower-latency speech generation |
| Reported latency | No comparable figure cited in the available sources | About 100 ms median inference latency; about 150 ms to first speech, according to ElevenLabs |
| Language support | More than 90 languages |
More than 90 languages |
| Availability | ElevenAPI, ElevenAgents and ElevenCreative |
ElevenAPI, ElevenAgents and ElevenCreative |
Choose v4 when delivery, emotion and script control matter most. Choose Turbo when response speed is central to the application. The available documentation describes Turbo as retaining high quality while reducing latency, but does not provide a like-for-like quality comparison across every use case. 17
ElevenLabs describes v4 as using a new architecture intended to give more control over tone, pacing, emotion and character. The public launch materials describe those goals but do not provide enough technical detail to independently assess the architecture itself. 5
10
The model also expands the use of inline audio tags, which let creators guide how a line is delivered. ElevenLabs introduced inline tags with v3; TechCrunch reported that v4 supports stacking tags and following their sequence. 5 That offers more direction inside a script, but the result still needs to be checked against the intended performance.
Both new models support more than 90 languages. ElevenLabs says an Instant Voice Clone can be created from as little as 10 seconds of audio. Professional Voice Clones, which the company says were not supported in v3, return in v4. 1
5
11
For longer scripts, ElevenLabs says its context stitching feature helps keep pacing and delivery steady across a script of any length. That may be useful for audiobooks and other long-form work, but it is a product claim—not independent evidence that every voice remains consistent in every project. 11
ElevenLabs reports roughly 100 ms median inference latency for v4 Turbo and roughly 150 ms median time to first speech. These are different measurements: one describes inference latency, while the other refers to the wait before speech begins. Neither figure, on its own, establishes how quickly a complete voice-agent interaction will feel to a caller. 11
A live interaction also depends on the rest of the system, including processing the caller’s audio, generating the agent’s response and delivering it over the network. So treat the published latency figures as useful model-level specifications, then measure response times in the application you plan to deploy.
A launch report says Artificial Analysis’ Provider Voice Arena leaderboard placed Eleven v4 at No. 1. That is a result for v4 on a particular evaluation; it does not establish that v4 leads on every language, voice, long-form task or live-agent setup. 6 The cited ranking should not be applied to Turbo: a separate sourced public benchmark for v4 Turbo was not identified in the available benchmark information.
18
Cartesia and Inworld are among the other providers discussed in coverage of real-time voice tools. Their existence makes direct testing important: compare the voices and latency using the same scripts, network conditions and agent workflow rather than relying on vendor latency claims alone. 19
The model split reflects a broader product strategy: serve creators who want controlled, expressive narration while also targeting companies building real-time voice agents. That is a product-positioning signal, not proof that either model has won its market.
The company has reported more than $500 million in annual recurring revenue in early 2026, and its February 2026 funding round valued it at $11 billion. 32
33 TechCrunch later reported that the company was pacing at $600 million in annual recurring revenue and that it was reportedly valued at $22 billion; the reported valuation should not be confused with a confirmed new funding round. The same report discussed IPO timing as an ambition, not an announced listing.
28 TechCrunch also described Decagon, a former ElevenLabs customer, as a competitor—an example of competition extending into customer-facing voice products.
28
For teams evaluating the models, the practical takeaway is straightforward: start with the use case, then test the specific voice and workflow. Compare v4 and Turbo on the same script or agent task, measure end-to-end latency, and verify that cloning and long-form consistency meet your needs. The launch establishes a clear quality-versus-speed choice; it does not settle which model will perform best for your project.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
Launched September 28, 2026, Eleven v4 is the quality focused model for expressive narration, while v4 Turbo targets live agents with about 100 ms median inference latency.
Launched September 28, 2026, Eleven v4 is the quality focused model for expressive narration, while v4 Turbo targets live agents with about 100 ms median inference latency. Both models support more than 90 languages; v4 adds expanded performance controls and brings Professional Voice Clones back, while Turbo is the real time option.
A reported top leaderboard result applies to v4, not automatically to Turbo or every language and use case.