ElevenLabs has launched Eleven v4 and Eleven v4 Turbo, a new generation of text-to-speech models designed to improve expressive speech generation, voice consistency and real-time AI voice interactions.
Announced on September 28, 2026, the two models target different workloads. Eleven v4 focuses on high-quality speech generation for content creation, narration, dialogue and character performance, while Eleven v4 Turbo is designed for lower-latency applications such as conversational AI agents. ElevenLabs says the new models are built on a new architecture intended to better capture tone, pacing, emotion and character in generated speech.
The release moves ElevenLabs beyond the improvements introduced with Eleven v3, particularly in areas such as voice cloning, multilingual speech and control over how generated dialogue is performed.
Quick Summary
- ElevenLabs launched Eleven v4 and Eleven v4 Turbo for AI voice generation.
- Eleven v4 focuses on expressive speech, emotion, pacing, dialogue and speaker consistency.
- Eleven v4 Turbo targets real-time applications with approximately 100 ms median inference latency.
- The new models support 90+ languages, including languages such as Hindi, Marathi, Cantonese, Mongolian and Odia.
- Eleven v4 adds greater control over pronunciation, dialogue and vocal performance.
- Voice cloning has been improved for greater speaker fidelity and consistency.
- ElevenLabs is positioning the models for AI agents, customer support, narration, content creation and conversational applications.
- The launch includes promotional API pricing, but the promotional rates should be distinguished from permanent pricing.
What Eleven v4 Changes?
Eleven v4 is positioned as the higher-quality model in the new generation. According to ElevenLabs’ documentation, it improves output quality, voice accuracy, consistency, emotion, delivery and audio-tag control compared with Eleven v3.
Rather than treating text primarily as a sequence of words to pronounce, the model is designed to interpret instructions about how a line should be delivered. Users can use inline audio tags to influence elements such as emotion, reactions and delivery.
For example, a script can contain directions such as [whispering], [laughing] or [shouting], allowing the generated performance to change at specific points in the text. Eleven v4 supports these tags through both its speech-generation and dialogue workflows.
This makes the model particularly relevant to applications where the delivery of a line is as important as the words themselves, including character performances, audiobooks, narration and dialogue.
Read More - ElevenLabs: The Complete Guide to AI Voice, Audio & Features
Eleven v4 Turbo Targets Real-Time AI Voice Agents
Eleven v4 Turbo is the lower-latency variant of the release.
ElevenLabs says Turbo has a median inference latency of approximately 100 milliseconds, positioning it for real-time applications where generated speech needs to begin quickly. The company specifically highlights conversational agents and interactive voice experiences as target use cases.
That distinction is important because highly expressive speech generation and real-time response speed can require different optimization priorities.
For an audiobook or character performance, users may prioritize expressive quality and consistency. For a customer-service or scheduling agent, latency becomes much more important because delays can make conversations feel unnatural.
Eleven v4 Turbo is intended to address the latter category while retaining the expressive capabilities of the new model generation.
Voice Cloning Gets a Major Upgrade
Voice cloning is another major part of the Eleven v4 release.
ElevenLabs says v4 captures characteristics such as a speaker’s timbre, cadence and delivery more faithfully than earlier models. Both Instant Voice Cloning and Professional Voice Cloning are supported.
The company also says Instant Voice Clones can be created from short recordings, while Professional Voice Clones remain available for workflows that require greater consistency and fidelity.
The improvement matters beyond simple voice imitation. Maintaining a recognizable speaker identity across long-form narration, dialogue and regenerated sections is important for audiobooks, video production, games and other applications where the same synthetic speaker may need to produce large amounts of content.
ElevenLabs’ documentation also notes that the quality of the source recording becomes particularly important with v4 because the model more closely reproduces characteristics of the reference voice.
Eleven v4 Supports More Than 90 Languages
Multilingual speech is another major focus of the release.
ElevenLabs says Eleven v4 supports more than 90 languages, with coverage including languages such as Cantonese, Mongolian and Odia. Its documented language list also includes Hindi, Marathi, Bengali, Gujarati, Punjabi and several other languages relevant to users in South Asia.
The model also introduces changes to cross-language voice generation. When the generated language differs from the language of the reference voice, Eleven v4 is designed to produce natural speech in the target language rather than automatically carrying over the original accent.
This could be particularly useful for multilingual content production and voice agents serving customers across different language markets.
More Control Over Pronunciation and Dialogue
Eleven v4 also gives developers additional control over pronunciation through IPA support.
That is useful for names, technical terminology, brand names and other words that conventional pronunciation systems may render incorrectly.
The model also supports multi-speaker dialogue, with ElevenLabs describing the new generation as providing more natural conversational dynamics between speakers.
Combined with audio tags, pronunciation controls and improved speaker consistency, the changes make v4 more suited to scripted dialogue rather than simple one-line text-to-speech generation.
Eleven v4 Availability and Launch Pricing
Eleven v4 and Eleven v4 Turbo are being made available through ElevenLabs’ ecosystem, including its creative tools, agent platform and API.
For the first two weeks following launch, ElevenLabs is offering promotional API pricing of $22 per 1 million characters for Eleven v4 and $11 per 1 million characters for Eleven v4 Turbo, according to the launch announcement. The company is also offering expanded access to Eleven v4 for eligible Creator+ users in ElevenCreative during the promotion.
The promotional rates should not be treated as permanent pricing. Developers considering migration from an existing ElevenLabs model will need to check the company’s current API pricing and model documentation before calculating ongoing production costs.
What About the Artificial Analysis #1 Claim?
ElevenLabs says Eleven v4 was ranked #1 by Artificial Analysis around the time of its launch.
That claim should be treated as a launch-time benchmark statement rather than a permanent ranking. The Artificial Analysis leaderboard currently accessible for this report does not yet show Eleven v4, while ElevenLabs’ v3 models remain listed on the leaderboard. Artificial Analysis explains that its TTS rankings are based on blind listener comparisons and Elo-style scoring.
As a result, the most accurate description is that ElevenLabs says v4 achieved the top Artificial Analysis ranking at launch, rather than stating that it is currently the #1 TTS model based on the live leaderboard.
Why Eleven v4 Matters for AI Voice?
The release reflects a broader shift in AI voice technology from simply producing intelligible speech toward controlling the performance of that speech.
ElevenLabs is targeting both sides of that equation with the new models: Eleven v4 emphasizes expressive and consistent generation, while Eleven v4 Turbo focuses on reducing latency for interactive systems.
That combination makes the models relevant to several growing AI applications, including voice agents, customer support, sales systems, scheduling assistants, games, audiobooks, dubbing and synthetic media.
The biggest change, however, is the level of control ElevenLabs is attempting to provide over generated performances. With audio tags, pronunciation controls, multilingual generation and improved speaker consistency, Eleven v4 is positioned as a broader speech-generation system rather than simply another incremental text-to-speech model.
ElevenLabs says the models will continue to evolve after launch, meaning performance and recommended workflows may change as the company continues training and refining the technology.
FAQs
1. What is Eleven v4?
Eleven v4 is ElevenLabs’ new-generation text-to-speech model focused on expressive speech, voice accuracy, consistency, dialogue and multilingual generation.
2. What is Eleven v4 Turbo?
Eleven v4 Turbo is the lower-latency variant designed primarily for real-time applications such as conversational AI agents and interactive voice experiences.
3. How many languages does Eleven v4 support?
ElevenLabs documents support for more than 90 languages, including Hindi, Marathi, Cantonese, Mongolian and Odia.
4. Does Eleven v4 support voice cloning?
Yes. Eleven v4 supports both Instant Voice Cloning and Professional Voice Cloning, with ElevenLabs reporting improved speaker accuracy and consistency.
5. How much does the Eleven v4 API cost?
During the initial two-week promotion, ElevenLabs lists Eleven v4 at $22 per 1 million characters and Eleven v4 Turbo at $11 per 1 million characters. Promotional pricing may change after the offer ends.
Also Read –
ElevenLabs Studio 4.0 Adds AI Video Editing
ElevenLabs Reception Launches as AI Receptionist
ElevenCreative by ElevenLabs: AI Platform for Voice, Music and Video
Source
ElevenLabs – Eleven v4 official page
ElevenLabs – Eleven v4 documentation
Artificial Analysis – Text-to-Speech Leaderboard
Artificial Analysis – ElevenLabs model analysis


