Eleven v4 + Eleven v4 TurboIt doesn't read the script. It performs it.
ElevenLabs calls Eleven v4 "our most emotive text-to-speech model yet". Write the delivery into the script with inline tags, in 90+ languages, up to 10,000 characters a take. And v4 Turbo answers fast enough to hold a live conversation.
Eight takes, generated in Popcraft. The words in orange brackets are the whole direction: there is no editing, no retakes stitched together, and nothing in the audio that is not on the page.
01
One sentence, five deliveries
[warm] The lighthouse keeper left a lamp burning for ships that never came.
[low, threatening] The lighthouse keeper left a lamp burning for ships that never came.
[rushed] The lighthouse keeper left a lamp burning for ships that never came.
[a tired detective who has heard it all before] The lighthouse keeper left a lamp burning for ships that never came.
[hushed and reverent, like a nature documentary narrator] The lighthouse keeper left a lamp burning for ships that never came.
02
A game NPC, turning on a line
[casual] You're looking for work, eh? Hmm, I might have something for you. [door creaking][suddenly suspicious, lowering his voice] Wait. Who sent you? [pause][delighted] Marta's cousin! Well why didn't you say so. [coins clinking] Half now, half when it's done.
03
One beat, four languages
[softly] Don't be afraid. We'll cross together when the morning comes.
[softly] 怖がらないで。朝が来たら、一緒に渡ろう。Japanese
[softly] Não tenha medo. Vamos atravessar juntos quando a manhã chegar.Brazilian Portuguese
[softly] 别害怕。等天亮了,我们一起过去。Mandarin
Same voice, same tag, the four languages ElevenLabs says gained the most in v4.
04
A support agent who sounds like one
[warm] Thanks so much for holding. I can see the duplicate charge right here. [reassuring] You don't need to do anything else; the refund goes out today. [clear, measured] And your eSIM for Reykjavik is active. Anything else I can check?
The register v4 Turbo is built to hold in a live call.
05
A whisper that stays a whisper
[whispers] Don't move. It's right behind you. [pause][whispers] On three, we run. One. Two. [shouts] Run!
06
An audiobook cold open
[quiet, measured narration] When I was young, I thought the mountain stood because it was strong. [pause] I know better now. [long pause] It stood because it had learned where to bend.
v4 takes 10,000 characters in one request, about ten minutes of audio.
07
Monday morning
[sleepy drowsy voice] Mmm, five more minutes. [yawning] Is it really Monday already?
08
Direction in plain language
[like a sports commentator, speeding up] She takes it on the half-volley, past one, past two, she's through on goal. [amazed] Oh, that is extraordinary!
A tag doesn't have to come from a list. Describe the delivery and the model performs it.
Two models. One for the performance, one for the conversation.
Until now you picked between expressive voices and fast ones. Eleven v4 ends that trade: the standard model carries long-form work, and Turbo carries live calls.
For creators
Eleven v4
Built for content creation and long-form audio: audiobooks, character work, dubbing-length narration. The delivery you write is the delivery you get, scene sounds included.
10,000characters per request, about ten minutes of audio in one take
90+languages, up from 70+ in v3
Tagsemotion, pacing, reactions, accents and sound effects, written inline
IPAexact pronunciation, written between forward slashes
"The culmination of our latest research in expressive speech generation."ElevenLabs, launch post
For voice agents
Eleven v4 Turbo
The same expressiveness, tuned for real time: about 100 ms median inference latency and about 150 ms to first speech, by ElevenLabs' measurements.
~100 msmedian inference latency
~150 msmedian time to first speech
Livesupport, sales and scheduling calls that adapt their tone mid-conversation
"Faster than the average pause between two people talking."ElevenLabs, launch post
Direct it like talent in the booth
Everything below is in the script itself. No sliders, no post.
Audio tags
Write the emotion where it happens
Tags sit inline, in square brackets, and take effect from where they stand to the next tag. Six kinds: emotion, delivery, pacing, reactions, accents and sound effects. Stack cues with commas, and the model performs them in sequence.
[excited] Big news, everyone. [pause][whispers] The summer menu drops Friday. [laughs]
Languages
90+ languages, one voice
ElevenLabs calls expressive speech across languages one of the hardest problems in audio AI. v4 covers more than ninety, with the biggest gains in Japanese, Brazilian Portuguese, Mandarin and Cantonese. The same voice carries the same direction across all of them, so a dub keeps its casting.
Long takes
10,000 characters in one request
Double what v3 accepted: about ten minutes of audio from a single generation. A chapter, a full support script or a long scene holds its pacing as one continuous read instead of stitched paragraphs.
Pronunciation
Say it exactly, with IPA
Names, drug names, place names, product names: write the pronunciation in IPA between forward slashes and the model says it your way, every time. For everything else, v4 set the highest pronunciation score Artificial Analysis has measured.
The ferry to Reykjavik /ˈreɪkjavɪk/ leaves at nine.
Voice clones
Professional Voice Clones are back
v3 never supported Professional Voice Clones. v4 does, so a produced voice you own can perform with the full tag vocabulary. Clones made on earlier models retrain for v4 with one click, with the voice owner's verified consent behind every clone.
The story behind it
Every ElevenLabs generation added one thing. v4 is the first that doesn't make you choose between them.
Aug 2023
Multilingual v2
29 languages, and ElevenLabs leaves beta.
Jul 2024
Turbo v2.5
Built for speed.
Dec 2024
Flash
Latency drops to about 75 ms. Fast, but not expressive.
Jun 2025
Eleven v3 alpha
Audio tags, 70+ languages and dialogue arrive. Expressive, but v3's own launch post told real-time builders to stay on Flash.
Feb 2026
Eleven v3 goes GA
The expressive line matures. The speed trade-off stays.
Jun 2026
The Warsaw preview
At its Warsaw summit, ElevenLabs demos the next model whispering, shifting accents and changing emotion live on stage.
28 Sep 2026
Eleven v4 + v4 Turbo
A new architecture ships both directions at once: the most expressive model ElevenLabs has built, and a Turbo cut fast enough for live conversation.
The proof
Independent first: Artificial Analysis runs blind listener arenas across every major speech model.
Artificial Analysis, October 2026
Eleven v4
Context
Provider Voice arena
#1
Ahead of Cartesia Sonic 3.6 and Gemini 3.8 Flash TTS. Eleven v3 sits at #17 on the same board.
Pronunciation Robustness
91.7%
The highest score Artificial Analysis has measured, on any model.
Controlled Voice arena
#2
Behind Qwen-Audio-3.1-TTS-Plus. v3 scored well below both.
Generation speed
73.4 chars/s
Against 42.5 for v3, measured by Artificial Analysis.
In ElevenLabs' own blind preference test, listeners chose v4 about 75% of the time against competing models. That figure is the vendor's; the arena rankings above are not. Reactions on launch day: the announcement thread and the Artificial Analysis results.
What changed from v3
Eleven v3
Eleven v4
Architecture
v3 family
Entirely new
Languages
70+
90+
Request length
5,000 characters
10,000 characters
Professional Voice Clones
Not supported
Supported
Direction
Audio tags
Tags, plain-language direction, IPA
Real time
Not recommended
v4 Turbo, ~100 ms median inference
Scripts you wrote for v3 run on v4 unchanged.
FAQ
Do my existing cloned voices work on v4?
After a retrain, yes. Clones made on earlier models retrain for v4 with one click in ElevenLabs' voice library. Be ready for differences: a retrained voice can sound substantially different from its v3 version, flaws in the source recording are reproduced faithfully, and voices built with Voice Design may perform less expressively than recorded clones.
Do audio tags always land?
Not always. ElevenLabs' own docs call tag-following "not perfect yet": occasionally a delivery cue renders as a sound effect instead. You'll get the most reliable results with one tag per clause, placed before the words it affects, and comma-joined cues inside a single bracket when you want to layer them.
What happens when a voice speaks a language that isn't its own?
It speaks with a native accent of that language, not the voice's home accent. ElevenLabs says this is by design, and that an accent-keeping toggle is being researched.
Which controls does v4 have?
Stability and Similarity. There is no style or speed slider and no SSML; pauses are written as [pause] and [long pause] in the script, and everything else is tags.
v4 or v4 Turbo?
v4 for anything you render and keep: narration, characters, dubbing-length reads up to 10,000 characters. v4 Turbo for anything that talks back: voice agents and interactive apps, where its ~100 ms median inference latency keeps the conversation moving.
Write the performance. v4 delivers it.
Pick a voice, put the direction in brackets, and hear the take in seconds.
All audio on this page was generated in Popcraft with ElevenLabs voices; the full script for each clip is shown beside it. Images are AI-generated. Quoted lines are ElevenLabs' own, from the Eleven v4 launch post. Rankings are from Artificial Analysis, read October 2026.