NEWCreate 3D models from text or images — rig, animate, export.Try 3D
PopcraftPopcraft
Voice model by ElevenLabs

Eleven v4 + Eleven v4 TurboIt doesn't read the script.
It performs it.

ElevenLabs calls Eleven v4 "our most emotive text-to-speech model yet". Write the delivery into the script with inline tags, in 90+ languages, up to 10,000 characters a take. And v4 Turbo answers fast enough to hold a live conversation.

[whispers] Don't move. It's right behind you.Eleven v4TTSGenerate
90+ languages10,000 characters per takeInline audio tags
A voice actress at a condenser microphone in a dim recording booth, lit by a warm practical lamp
Character voicing
A dubbing stage: an actress watches a projected scene, her face lit by the screen
Dubbing
A control room at night, waveforms glowing on the monitors in front of an engineer
Direction
A game character: an armoured knight by firelight, mid-line
Game NPCs
A support agent on a headset call in an office at dawn
Voice agents
A silver-haired narrator reads aloud from a hardback book in an armchair under a reading lamp
Audiobooks

Hear it. The script is beside every clip.

Eight takes, generated in Popcraft. The words in orange brackets are the whole direction: there is no editing, no retakes stitched together, and nothing in the audio that is not on the page.

01

One sentence, five deliveries

[warm] The lighthouse keeper left a lamp burning for ships that never came.

[low, threatening] The lighthouse keeper left a lamp burning for ships that never came.

[rushed] The lighthouse keeper left a lamp burning for ships that never came.

[a tired detective who has heard it all before] The lighthouse keeper left a lamp burning for ships that never came.

[hushed and reverent, like a nature documentary narrator] The lighthouse keeper left a lamp burning for ships that never came.

02

A game NPC, turning on a line

[casual] You're looking for work, eh? Hmm, I might have something for you. [door creaking] [suddenly suspicious, lowering his voice] Wait. Who sent you? [pause] [delighted] Marta's cousin! Well why didn't you say so. [coins clinking] Half now, half when it's done.

03

One beat, four languages

[softly] Don't be afraid. We'll cross together when the morning comes.

[softly] 怖がらないで。朝が来たら、一緒に渡ろう。Japanese

[softly] Não tenha medo. Vamos atravessar juntos quando a manhã chegar.Brazilian Portuguese

[softly] 别害怕。等天亮了,我们一起过去。Mandarin

Same voice, same tag, the four languages ElevenLabs says gained the most in v4.

04

A support agent who sounds like one

[warm] Thanks so much for holding. I can see the duplicate charge right here. [reassuring] You don't need to do anything else; the refund goes out today. [clear, measured] And your eSIM for Reykjavik is active. Anything else I can check?

The register v4 Turbo is built to hold in a live call.

05

A whisper that stays a whisper

[whispers] Don't move. It's right behind you. [pause] [whispers] On three, we run. One. Two. [shouts] Run!

06

An audiobook cold open

[quiet, measured narration] When I was young, I thought the mountain stood because it was strong. [pause] I know better now. [long pause] It stood because it had learned where to bend.

v4 takes 10,000 characters in one request, about ten minutes of audio.

07

Monday morning

[sleepy drowsy voice] Mmm, five more minutes. [yawning] Is it really Monday already?

08

Direction in plain language

[like a sports commentator, speeding up] She takes it on the half-volley, past one, past two, she's through on goal. [amazed] Oh, that is extraordinary!

A tag doesn't have to come from a list. Describe the delivery and the model performs it.

Two models. One for the performance, one for the conversation.

Until now you picked between expressive voices and fast ones. Eleven v4 ends that trade: the standard model carries long-form work, and Turbo carries live calls.

For creators

Eleven v4

Built for content creation and long-form audio: audiobooks, character work, dubbing-length narration. The delivery you write is the delivery you get, scene sounds included.

  • 10,000characters per request, about ten minutes of audio in one take
  • 90+languages, up from 70+ in v3
  • Tagsemotion, pacing, reactions, accents and sound effects, written inline
  • IPAexact pronunciation, written between forward slashes
"The culmination of our latest research in expressive speech generation."ElevenLabs, launch post
For voice agents

Eleven v4 Turbo

The same expressiveness, tuned for real time: about 100 ms median inference latency and about 150 ms to first speech, by ElevenLabs' measurements.

  • ~100 msmedian inference latency
  • ~150 msmedian time to first speech
  • Livesupport, sales and scheduling calls that adapt their tone mid-conversation
"Faster than the average pause between two people talking."ElevenLabs, launch post

Direct it like talent in the booth

Everything below is in the script itself. No sliders, no post.

Audio tags

Write the emotion where it happens

Tags sit inline, in square brackets, and take effect from where they stand to the next tag. Six kinds: emotion, delivery, pacing, reactions, accents and sound effects. Stack cues with commas, and the model performs them in sequence.

[excited] Big news, everyone. [pause] [whispers] The summer menu drops Friday. [laughs]
A close-up of a condenser microphone in a haze of backlight
Languages

90+ languages, one voice

ElevenLabs calls expressive speech across languages one of the hardest problems in audio AI. v4 covers more than ninety, with the biggest gains in Japanese, Brazilian Portuguese, Mandarin and Cantonese. The same voice carries the same direction across all of them, so a dub keeps its casting.

A radio-drama cast of four actors gathered around one microphone
Long takes

10,000 characters in one request

Double what v3 accepted: about ten minutes of audio from a single generation. A chapter, a full support script or a long scene holds its pacing as one continuous read instead of stitched paragraphs.

A tall stack of manuscript pages beside a studio microphone under warm lamp light
Pronunciation

Say it exactly, with IPA

Names, drug names, place names, product names: write the pronunciation in IPA between forward slashes and the model says it your way, every time. For everything else, v4 set the highest pronunciation score Artificial Analysis has measured.

The ferry to Reykjavik /ˈreɪkjavɪk/ leaves at nine.
A script page marked up with pencil notation on a music stand
Voice clones

Professional Voice Clones are back

v3 never supported Professional Voice Clones. v4 does, so a produced voice you own can perform with the full tag vocabulary. Clones made on earlier models retrain for v4 with one click, with the voice owner's verified consent behind every clone.

Two identical reel-to-reel tape machines side by side in a studio rack

The story behind it

Every ElevenLabs generation added one thing. v4 is the first that doesn't make you choose between them.

  • Aug 2023

    Multilingual v2

    29 languages, and ElevenLabs leaves beta.

  • Jul 2024

    Turbo v2.5

    Built for speed.

  • Dec 2024

    Flash

    Latency drops to about 75 ms. Fast, but not expressive.

  • Jun 2025

    Eleven v3 alpha

    Audio tags, 70+ languages and dialogue arrive. Expressive, but v3's own launch post told real-time builders to stay on Flash.

  • Feb 2026

    Eleven v3 goes GA

    The expressive line matures. The speed trade-off stays.

  • Jun 2026

    The Warsaw preview

    At its Warsaw summit, ElevenLabs demos the next model whispering, shifting accents and changing emotion live on stage.

  • 28 Sep 2026

    Eleven v4 + v4 Turbo

    A new architecture ships both directions at once: the most expressive model ElevenLabs has built, and a Turbo cut fast enough for live conversation.

The proof

Independent first: Artificial Analysis runs blind listener arenas across every major speech model.

Artificial Analysis, October 2026Eleven v4Context
Provider Voice arena#1Ahead of Cartesia Sonic 3.6 and Gemini 3.8 Flash TTS. Eleven v3 sits at #17 on the same board.
Pronunciation Robustness91.7%The highest score Artificial Analysis has measured, on any model.
Controlled Voice arena#2Behind Qwen-Audio-3.1-TTS-Plus. v3 scored well below both.
Generation speed73.4 chars/sAgainst 42.5 for v3, measured by Artificial Analysis.

In ElevenLabs' own blind preference test, listeners chose v4 about 75% of the time against competing models. That figure is the vendor's; the arena rankings above are not. Reactions on launch day: the announcement thread and the Artificial Analysis results.

What changed from v3

Eleven v3Eleven v4
Architecturev3 familyEntirely new
Languages70+90+
Request length5,000 characters10,000 characters
Professional Voice ClonesNot supportedSupported
DirectionAudio tagsTags, plain-language direction, IPA
Real timeNot recommendedv4 Turbo, ~100 ms median inference

Scripts you wrote for v3 run on v4 unchanged.

FAQ

Do my existing cloned voices work on v4?

After a retrain, yes. Clones made on earlier models retrain for v4 with one click in ElevenLabs' voice library. Be ready for differences: a retrained voice can sound substantially different from its v3 version, flaws in the source recording are reproduced faithfully, and voices built with Voice Design may perform less expressively than recorded clones.

Do audio tags always land?

Not always. ElevenLabs' own docs call tag-following "not perfect yet": occasionally a delivery cue renders as a sound effect instead. You'll get the most reliable results with one tag per clause, placed before the words it affects, and comma-joined cues inside a single bracket when you want to layer them.

What happens when a voice speaks a language that isn't its own?

It speaks with a native accent of that language, not the voice's home accent. ElevenLabs says this is by design, and that an accent-keeping toggle is being researched.

Which controls does v4 have?

Stability and Similarity. There is no style or speed slider and no SSML; pauses are written as [pause] and [long pause] in the script, and everything else is tags.

v4 or v4 Turbo?

v4 for anything you render and keep: narration, characters, dubbing-length reads up to 10,000 characters. v4 Turbo for anything that talks back: voice agents and interactive apps, where its ~100 ms median inference latency keeps the conversation moving.

Write the performance. v4 delivers it.

Pick a voice, put the direction in brackets, and hear the take in seconds.

All audio on this page was generated in Popcraft with ElevenLabs voices; the full script for each clip is shown beside it. Images are AI-generated. Quoted lines are ElevenLabs' own, from the Eleven v4 launch post. Rankings are from Artificial Analysis, read October 2026.

Keep exploring

Related models