General Blog

ElevenLabs V4: What Changed, How It Works, and Who Should Use It

Learn what ElevenLabs V4 changes, including expressive speech, voice cloning, Audio Tags, 90+ languages, dialogue support, and V4 Turbo.

ElevenLabs V4: What Changed, How It Works, and Who Should Use It

ElevenLabs V4 is the company’s newest generation of text-to-speech technology, designed to produce more expressive speech, improve voice-cloning accuracy, handle more languages, and give creators finer control over how generated speech is performed. ElevenLabs released Eleven V4 alongside the lower-latency Eleven V4 Turbo on September 28, 2026.

For creators, developers, audiobook producers, game studios, and teams building conversational AI, the most important change is not simply better audio quality. V4 is designed to interpret the emotional and performance context of a script more effectively. Instructions such as whispers, excitement, hesitation, laughter, urgency, or other delivery cues can be placed directly into text through Audio Tags, while the model can also generate dialogue involving multiple speakers.

ElevenLabs currently positions the standard Eleven V4 model as its quality-focused option for content production, while Eleven V4 Turbo targets interactive applications where response speed matters. That distinction makes choosing between them relatively straightforward: prioritize V4 when expressive quality is the main goal and evaluate V4 Turbo when latency is critical.

Quick verdict

ElevenLabs V4 is a substantial text-to-speech upgrade focused on expressive delivery, stronger voice cloning, 90+ language support, Audio Tags, and natural multi-speaker dialogue. The standard V4 model is better suited to quality-first content production, while V4 Turbo is designed for real-time and interactive voice applications.

Key takeaways

  • Eleven V4 supports more than 90 languages and a 10,000-character text limit for standard speech generation
  • Audio Tags let creators direct emotion, pacing, reactions, and delivery using natural-language cues
  • Eleven V4 improves Instant and Professional Voice Cloning fidelity according to ElevenLabs
  • Multi-speaker dialogue is supported through ElevenLabs dialogue endpoints
  • Eleven V4 Turbo targets real-time applications with approximately 100 ms median inference latency

What is ElevenLabs V4?

ElevenLabs V4, identified as eleven_v4 in the API, is the latest high-quality text-to-speech model in ElevenLabs’ speech-generation lineup. ElevenLabs describes it as its most expressive model and the first generation after Eleven V3.

The model converts written text into spoken audio while considering more than pronunciation alone. Its design emphasizes tone, emotion, cadence, pacing, speaker identity, language, and contextual delivery. This matters because two speakers can read exactly the same words yet communicate completely different meanings depending on how the sentence is delivered.

ElevenLabs lists support for more than 90 languages for V4, compared with 70+ languages for Eleven V3. The standard V4 model supports a 10,000-character input limit and natural multi-speaker dialogue. It is available through ElevenLabs’ creator tools and API infrastructure.

There are two closely related versions:

  • Eleven V4: the quality-focused model intended for use cases such as audiobooks, character voices, narration, creative voiceovers, and other situations where expressive performance matters more than minimum latency.
  • Eleven V4 Turbo: the real-time-oriented variant designed for conversational agents and interactive experiences. ElevenLabs reports median inference latency of approximately 100 milliseconds for Turbo.

This means V4 should not be considered a universal replacement for every ElevenLabs speech model. For example, an application primarily concerned with very low cost, extremely long text inputs, or another specialized requirement may still find another model family more appropriate.

What changed from Eleven V3?

The transition from Eleven V3 to V4 focuses heavily on expressiveness, voice identity, language coverage, and controllability. ElevenLabs’ own documentation recommends that V3 users test V4 with their existing voices and content, while acknowledging that specific edge cases may still perform differently.

More expressive speech

One of V4’s central goals is making synthesized speech behave more like a performance instead of a neutral reading. The model can respond to contextual signals that suggest whether a line should sound dramatic, conversational, tender, frightened, amused, excited, or urgent.

This can be particularly useful for character dialogue, games, audiobooks, storytelling, advertising creative, podcasts, and entertainment workflows where merely pronouncing the words correctly is not enough.

Improved voice cloning

ElevenLabs says V4 improves the accuracy of both Instant Voice Clones and Professional Voice Clones. The goal is to retain characteristics such as timbre, cadence, delivery style, and speaker identity more faithfully.

Instant Voice Cloning remains useful when someone wants to establish a voice from a comparatively short reference sample, while Professional Voice Cloning is intended for higher-fidelity workflows. The quality of any cloned voice can still depend on the suitability and quality of the reference material, so a newer model cannot automatically fix every weak recording or poorly matched voice.

Expanded language coverage

Eleven V4 supports more than 90 languages according to current ElevenLabs documentation. Another notable behavior concerns accents across languages. ElevenLabs says that when generated speech uses the same language as the reference voice, the original accent is preserved. When generation switches into another language, V4 is designed to produce natural speech in the target language instead of automatically transferring the reference speaker’s original-language accent.

That could make V4 particularly relevant to multilingual media production, localized characters, dubbing-related workflows, and international content operations. Teams should still evaluate individual languages, voices, and accents before depending on the output in production.

Audio Tags are one of V4's most useful controls

Audio Tags provide a relatively simple way of directing performance without building a complicated control interface. A creator can place a natural-language instruction inside square brackets and position it near the dialogue it should affect.

For example, tags can indicate broad emotional or delivery changes such as whispering, shouting, curiosity, crying, laughter, sighing, or other reactions. ElevenLabs also allows creators to experiment with their own descriptive tags rather than restricting them to one fixed vocabulary.

Tags can be combined to give the model more context. A script might specify that a character should sound both nervous and quiet, or playful and excited. Punctuation can then provide another layer of pacing control.

This workflow is important because it makes voice direction accessible to writers and creators who may not want to manipulate detailed synthesis parameters. Instead of expressing performance numerically, users can write instructions closer to the language they would use when directing a human performer.

Expert tip

Treat Audio Tags as performance direction rather than guaranteed commands. Start with a small number of clear cues, generate several versions when necessary, and only add more specific direction when the model is not producing the intended delivery.

V4 does not use SSML break tags

Teams migrating an existing speech workflow should pay particular attention to how V4 handles prompting. ElevenLabs documentation states that Eleven V4 and Eleven V3 do not support SSML break tags.

Instead, creators are encouraged to use Audio Tags, punctuation, ellipses, sentence structure, and similar textual techniques to influence pauses and pacing. That makes the scripting workflow more natural for many creators, but it may require changes for applications built around SSML-heavy templates.

V4 also exposes Stability and Similarity as its main voice settings. ElevenLabs documentation says the Style and Speed sliders available in some other model workflows are not available in Eleven V4.

Multi-speaker dialogue is a major V4 use case

Eleven V4 supports natural multi-speaker dialogue through ElevenLabs’ Text to Dialogue capabilities. Rather than generating each line separately and assembling a conversation afterward, developers can create dialogue turns associated with different voice IDs.

Audio Tags can also be included inside individual turns to shape each speaker's delivery. This makes the capability relevant to fictional conversations, game dialogue, podcast-style productions, educational simulations, scripted interviews, interactive stories, and other multi-character experiences.

There is an important practical limitation: generated speech remains nondeterministic. ElevenLabs specifically notes in its dialogue documentation that multiple generations may sometimes be required to achieve the desired result. Applications where every generation must be identical or perfectly predictable should account for this behavior.

Where Eleven V4 fits in an AI audio workflow

V4 can be used from ElevenLabs’ creator-facing tools or integrated into applications through the ElevenAPI. Developers generating speech can specify the eleven_v4 model through supported speech endpoints. Multi-speaker projects can use the dedicated dialogue endpoints.

The appropriate workflow largely depends on what is being produced.

Audiobooks and long-form narration

Audiobook teams may benefit from improved emotional delivery, voice consistency, and language coverage. Character-focused projects can also make use of dialogue and Audio Tags. Long projects should still be divided thoughtfully into sections so that editors can manage quality and regenerate problematic passages without unnecessarily recreating an entire production.

Games and character dialogue

Games are an especially natural fit for expressive synthesis. Different emotional versions of a line can be generated without requiring a separate recording session for every variation. Multi-speaker support can also help with prototypes, dynamic scenes, and narrative content.

Production teams should nevertheless maintain human quality assurance for important character performances rather than assuming that one generation will always match the intended acting direction.

Video narration and creative content

Video creators can use V4 when narration needs more personality than traditional neutral text-to-speech. Emotional direction can help with storytelling, advertisements, explainers, social video, fictional scenes, and branded creative.

Conversational AI

Standard V4 prioritizes maximum quality, while V4 Turbo is specifically designed to bring the V4 model family into real-time applications. ElevenLabs reports approximately 100 ms median inference latency for V4 Turbo, making it the more relevant variant for voice agents and other interactive systems where waiting for a high-quality offline generation would hurt the user experience.

Eleven V4 versus Eleven V4 Turbo

AreaEleven V4Eleven V4 Turbo
Primary goalMaximum expressive qualityReal-time expressive speech
Typical useAudiobooks, content creation, character voiceoversVoice agents and interactive applications
Languages90+90+
Audio TagsSupportedSupported
Voice cloningSupportedSupported
Latency focusQuality-firstApproximately 100 ms median inference latency reported by ElevenLabs

The distinction is useful because speech applications often have competing requirements. A publisher generating a finished audiobook can usually tolerate additional generation time if the output is better. A conversational agent cannot: a long pause after every user statement can make the interaction feel unnatural.

Teams building both kinds of experiences should therefore test each model against the actual workflow rather than selecting a model solely because it has the newest name.

Important limitations and considerations

Although V4 expands ElevenLabs’ speech-generation capabilities, it does not remove the normal operational issues associated with generative audio.

Output can vary between generations. ElevenLabs describes its speech models as nondeterministic. A seed can improve consistency in some workflows, but subtle differences may remain.

Audio Tags are not guarantees. They influence performance, but results can vary by voice, script, language, reference material, and requested delivery.

Voice selection still matters. ElevenLabs notes that performance styles already represented in a voice's source material may be easier to reproduce than dramatically different deliveries.

SSML-dependent workflows need adjustment. V4 does not support SSML break tags, so applications depending on them need to use supported prompting and text-structure methods instead.

Generated voice does not remove rights considerations. Users remain responsible for having the necessary rights to source material, voices, scripts, and commercial content. ElevenLabs says generated audio belongs to the user, while commercial usage requires an appropriate paid plan and rights to the underlying input.

High-quality output should still be reviewed. Audiobooks, advertising, games, customer-facing agents, educational material, and other production use cases benefit from human listening and quality assurance before publication.

Who should use this

  • Creators producing expressive narration, character dialogue, podcasts, or video voiceovers
  • Developers integrating high-quality multilingual speech into applications
  • Audiobook and media teams that need emotion and multi-speaker dialogue
  • Voice-agent teams evaluating Eleven V4 Turbo for real-time conversations

Who should avoid this

  • Teams that require SSML break-tag compatibility without changing their workflow
  • Applications where nondeterministic voice generation is unacceptable
  • Users unwilling or unable to verify voice rights and commercial-use requirements

Should existing Eleven V3 users switch?

ElevenLabs itself recommends testing V4 as an upgrade from V3, and the published specifications provide several reasons to do so: broader language support, improved voice cloning, stronger emotional delivery, enhanced Audio Tag behavior, and expanded model capabilities.

That does not mean every V3 workflow should be changed without evaluation. Existing voices, scripts, integrations, latency targets, and production processes may respond differently. A safer migration approach is to create a representative test set containing the types of narration, dialogue, accents, languages, emotional cues, names, and difficult pronunciations used in production.

Generate that same set with the existing model and V4, then compare the results against the actual requirements of the project. For real-time applications, test V4 Turbo separately because latency is part of the product experience, not merely a technical benchmark.

Final thoughts

ElevenLabs V4 moves the company's text-to-speech technology further toward controllable voice performance rather than basic speech synthesis. More than 90 supported languages, improved voice cloning, natural-language Audio Tags, multi-speaker dialogue, and a dedicated Turbo variant make the model relevant to a wide range of creative and developer workflows.

The strongest reason to evaluate V4 is its combination of expressive control and voice identity. The biggest practical caution is that generative speech still requires experimentation and quality control: tags do not guarantee identical results, outputs are nondeterministic, and the suitability of a voice depends on the source material and target delivery.

For existing ElevenLabs users, a representative side-by-side test with V3 is more useful than assuming every project will improve automatically. New users should choose between standard V4 and V4 Turbo according to whether production quality or real-time responsiveness is the higher priority.

Frequently asked questions

Explore related SearchSagar pages

What is ElevenLabs V4?
ElevenLabs V4 is the latest quality-focused text-to-speech model from ElevenLabs. It is designed for expressive speech, improved voice cloning, multilingual generation, Audio Tags, and multi-speaker dialogue.
How many languages does ElevenLabs V4 support?
ElevenLabs currently documents support for more than 90 languages with Eleven V4.
What is the difference between Eleven V4 and Eleven V4 Turbo?
Eleven V4 prioritizes maximum speech quality for content production, while Eleven V4 Turbo is designed for real-time and interactive applications. ElevenLabs reports approximately 100 ms median inference latency for V4 Turbo.
Does ElevenLabs V4 support voice cloning?
Yes. ElevenLabs V4 supports both Instant Voice Cloning and Professional Voice Cloning, and ElevenLabs says the new model improves how faithfully cloned voices reproduce source characteristics such as timbre, cadence, and delivery.
Does ElevenLabs V4 support Audio Tags?
Yes. Audio Tags are natural-language instructions placed in square brackets that can guide emotions, delivery, reactions, pacing, and other performance characteristics.
Does ElevenLabs V4 support SSML?
ElevenLabs states that Eleven V4 does not support SSML break tags. It recommends using Audio Tags, punctuation, ellipses, and text structure to influence delivery and pacing instead.
Can ElevenLabs V4 generate conversations with multiple speakers?
Yes. Eleven V4 supports multi-speaker generation through ElevenLabs' Text to Dialogue capabilities, where individual dialogue turns can use different voice IDs and performance instructions.

Written by

Shikha Goyal

Author

We use optional analytics to understand site usage. No analytics loads until you consent.