ElevenLabs Voice & Dubbing: The Essentials
Last updated: August 14, 2026
This guide covers the ElevenLabs voice and dubbing tools on Scenario. Dubbing v2 and the earlier Dubbing translate spoken audio or video into new languages while preserving each speaker's voice, tone, and timing. Voice Clone builds a reusable custom voice from a sample. Voice Isolator strips a clean voice out of a noisy recording, and Voice Changer re-voices a recording into another voice. Each model has its own section below with examples, parameters, and tips.
Dubbing v2
ElevenLabs Dubbing v2 is next-generation automatic dubbing. Give it a spoken audio or video clip, choose a target language, and it returns the same performance in that language while keeping the original speaker's voice, tone, emotion, and timing. Instead of a flat, re-read translation, it conditions directly on the source audio, so intonation and pacing carry across.
It covers 31 target languages on Scenario, auto-detects the source language, and can lock brand names, product names, and jargon in place with keyterms so they survive the translation verbatim. That makes it a fast path to localizing marketing, e-learning, game trailers, film, customer stories, and social content without re-recording a single line.
How ElevenLabs Dubbing v2 Works
Dubbing is a single-step transformation, not a prompt-driven generation. You provide one input file and one target language, and the model does the rest: it transcribes the speech, translates it, clones each speaker's voice, and regenerates the dialogue in the new language, timed to the original delivery.
Voice, emotion, and timing are preserved
Because the model listens to the source rather than a plain transcript, the dubbed track keeps the speaker's identity and performance. A calm explainer stays calm, an energetic ad stays energetic, and a dramatic trailer keeps its weight. In our test batch, every dubbed clip came back at the exact same length as its source (for example 26.8 seconds in, 26.8 seconds out), so the result drops straight back onto the original footage.
Source language auto-detection
You do not have to tell the model what language the clip is in. Source Language defaults to auto, so you can dub a Spanish testimonial into English or a French documentary into English without setting anything beyond the target. Set it explicitly only when a clip mixes languages or a rare accent trips detection.
Protecting names with Keyterms
Translation engines love to translate everything, including things that should never change: product names, brand names, and technical terms. Keyterms is a list of strings the model holds constant across the translation. In the German example below, the terms Scenario, Nexus, and Flux all survive intact inside fully natural German. See the onboarding example in the next section.
Audio or video input
The File input accepts either an audio file or a video file. When you pass a video, the model extracts the audio automatically, so a talking-head clip can go straight in. The examples in this article use audio, rendered here as waveform videos so they play inline.
Video Dubbing Examples
Dubbing v2 localizes video, not just audio. Feed it a clip with a voice-over and it returns the same footage with the narration translated into the target language, timed to the original delivery and keeping the speaker's tone. Every clip below is a single narrated video dubbed on ElevenLabs Dubbing v2.
Product ad, dubbed to Spanish
Flying-car commercial, English to Spanish · Open on Scenario
Anime trailer, dubbed to Japanese
Samurai trailer narration, English to Japanese · Open on Scenario
Sports broadcast, dubbed to French
Tennis commentary, English to French · Open on Scenario
Epic documentary, dubbed to Italian
Documentary narration, English to Italian · Open on Scenario
Creature spot, dubbed to German
Pet-dinosaur spot, English to German · Open on Scenario
Language Coverage
The same four source clips were dubbed across a wider spread of languages to sanity-check quality end to end. Every clip below was produced on ElevenLabs Dubbing v2 and opens in Scenario.
Spanish, product marketing: Open on Scenario
Japanese, game trailer: Open on Scenario
French, e-learning lesson: Open on Scenario
English, customer testimonial (from Spanish): Open on Scenario
German, software onboarding (keyterms): Open on Scenario
Portuguese, travel promo: Open on Scenario
Korean, tech keynote (keyterms): Open on Scenario
English, documentary narration (from French): Open on Scenario
Hindi, nonprofit announcement: Open on Scenario
Chinese, e-commerce livestream: Open on Scenario
Parameters
Dubbing v2 has a small, focused parameter set. There is no prompt: you choose an input file, a target language, and optionally a source language and keyterms.
file
Required. The audio or video clip to dub. Video is accepted directly and its audio is extracted automatically. Best results come from clean speech; heavy background noise or overlapping speakers make transcription and voice cloning harder.
targetLang
Required. The language to dub into, as a language code. Scenario exposes 31 targets, including English, Spanish, French, German, Italian, Japanese, Korean, Chinese, Portuguese, Hindi, Arabic, Russian, Dutch, Polish, Turkish, and more. Every example in this article was produced by changing only this value.
sourceLang
Optional, defaults to auto. The language of the input clip. Leave it on auto for almost everything, as shown in the Spanish-to-English and French-to-English examples above. Set it explicitly when a clip mixes languages or an unusual accent causes a misdetection.
keyterms
Optional. A list of terms to preserve verbatim across the translation, such as names, brands, and jargon (up to 1000 terms of 50 characters each). This is what keeps Scenario, Nexus, and Flux intact in the German example. Pass it as a list, even for a single term. Note that keyterms protects wording, not spelling systems: a Latin-script term stays as written, while a language like Korean or Japanese may render a name phonetically in its own script.
Use Cases
Marketing localization: turn one ad or product video into a dozen regional versions that keep the same voice and energy, no re-record required.
Games: localize trailers, cinematics, and character barks while preserving the dramatic performance of the original read.
E-learning and training: ship a course or onboarding walkthrough in every language your learners speak, with the instructor's voice intact.
Film and documentary: dub narration and interviews into new markets while keeping tone and pacing aligned to the picture.
Customer stories and social: take a testimonial recorded in one language and share it globally, or repurpose short-form video for new audiences.
Tips for Better Results
Feed it clean speech. Clear, well-recorded dialogue with minimal background noise produces the most faithful voice cloning and translation.
Leave Source Language on auto. Detection was reliable across our tests, including reversing direction from Spanish and French into English. Only override it when a clip is mixed-language.
List your keyterms up front. Add every product name, brand, and piece of jargon you want untouched before you run, rather than fixing them after.
Keep speakers from overlapping. One voice at a time transcribes and clones more accurately than crosstalk.
Pass video directly when you have it. There is no need to strip the audio first; the model extracts it, and the timing lines back up with your footage.
Match the target to the region, not just the language. Pick the target language that fits your audience, and preview a short clip before dubbing a long one.
Trust the timing. Outputs came back at the same duration as the source, so you can plan edits around the original length.
Known Limitations
Language set is a subset. ElevenLabs supports more languages than Scenario currently exposes. Scenario offers 31 targets today; if you need one outside that list, it is not available here yet.
Names in non-Latin scripts. Keyterms holds a term constant, but in languages with a different writing system a protected name may be rendered phonetically (for example a Latin-script product name appearing in Korean script). Check brand-critical clips.
Depends on input quality. Noisy recordings, music beds under the voice, or heavy overlapping speech reduce accuracy. Clean the source audio first when you can.
Invented or very rare words. Made-up names (fantasy places, coined product words) can shift under translation; add them to keyterms, and review the result.
Not a same-language voice swap. Dubbing translates into another language. To change the voice without translating, use a voice-changer or speech-to-speech model instead.
Voice Isolator and Voice Changer
Two ElevenLabs audio tools on Scenario, built to work as a pair. Voice Isolator strips background noise, music, and reverb so only a clean voice remains. Voice Changer takes that voice (or any clean recording) and replaces the speaker while keeping the exact words, timing, and emotion.
The short version
Noisy take? Run Voice Isolator first.
Different character voice? Feed the result into Voice Changer.
Neither model uses a text prompt. Both are speech-to-speech: audio in, audio out.
Which Model Should I Use?
Model | ID | Input | Best for |
|---|---|---|---|
ElevenLabs Voice Isolator Editing |
| Audio or video | Remove music, noise, and room tone; keep speech only |
ElevenLabs Voice Changer Editing |
| Audio | Swap the speaker to a preset or cloned voice without re-recording |
Use Voice Isolator when the mix is the problem. Use Voice Changer when the performance is right but the voice is wrong. For heavy noise, isolate before you change. For clean studio takes, skip straight to Voice Changer.
Parameters
Voice Isolator
Audio (required). The recording to clean. Built for speech: there must be a voice to extract. Accepts encoded files (MP3, WAV, and similar) and can take a video asset directly; Scenario pulls the audio track for you.
Input File Format. Leave on Encoded (MP3, WAV, etc.) for normal uploads. Switch to PCM 16-bit 16 kHz mono only if your file is already in that raw format for slightly faster processing.
Voice Changer
Audio (required). The performance to transform. Words, pauses, and emotion stay tied to the original read.
Voice or Public Voice (one required). Pick a cloned ElevenLabs voice from your project, or choose one of 21 presets (Adam, Bella, Charlie, Liam, Sarah, and others). If both are set, the cloned voice wins.
Remove Background Noise. Optional cleanup before conversion. Useful on noisy sources; skip it when the input is already studio-clean or you ran Voice Isolator first.
Stability, Similarity Boost, Style Exaggeration. Optional dials from 0 to 1. Leave empty for the voice defaults. Lower stability adds expressiveness. Higher similarity locks to the target voice. Higher style exaggeration pushes emotion but can reduce steadiness. Change one dial at a time when tuning.
Speaker Boost. Optional. Makes the output sound more like the chosen voice at the cost of slightly slower processing.
Output Format. Default
mp3_44100_128is a solid balance. Usemp3_44100_192or Opus 192 kbps for final delivery; lower bitrates for quick iteration.Seed. Optional. Lock a number to reproduce the same output on repeat runs with identical settings.
Input File Format. Same Encoded vs PCM choice as Voice Isolator.
How Voice Isolator Works
Upload audio or video, run the model, and get back an isolated speech track. Music, crowd noise, HVAC hum, and room reverb are stripped while timing and delivery stay intact.
Removes background music from speech recordings
Removes ambient noise (traffic, crowds, fans, AC)
Reduces reverb from echoey rooms
Preserves natural timing, tone, and emotion in the voice
Not a voice changer. Voice Isolator keeps the original speaker. To swap identity, run Voice Changer on the isolated file.
How Voice Changer Works
Upload a recording, pick a target voice, and generate. The model converts identity, not script: mispronunciations, breaths, and pacing from the source carry through unless you fix them upstream.
21 preset voices for instant character variety
Cloned voices for project-specific characters (create clones with ElevenLabs Voice Clone on Scenario)
Optional in-model noise removal when you skip the isolator step
MP3 or Opus export at multiple bitrates
Using the Two Models Together
Raw recording (noisy or mixed)
→ ElevenLabs Voice Isolator
→ ElevenLabs Voice Changer (preset or cloned voice)
→ Delivery-ready audio for game, film, or podcastCapture the take. Record the line in any environment, even a noisy one.
Isolate. Run Voice Isolator. Music and room tone drop out; speech remains.
Recast. Feed the clean file into Voice Changer. Pick a preset or clone.
Place in the edit. Export MP3 or Opus and drop into your timeline, game engine, or lipsync pipeline.
Examples
Game dialogue from director reads. Record lines in a small office. Isolate to remove hum, then Voice Changer with three cloned voices for three NPCs.
Podcast guest from a noisy cafe. Phone recording with cafe ambience. Voice Isolator alone can be enough; no recast required.
Animatic voice swap. Placeholder reads with the right emotion. Voice Changer recasts each line into final character voices for a test screening.
Music vocal extraction. Isolate a forward lead vocal for sampling or lipsync input. Works best on dry, upfront vocals.
Archival restoration. Old interview with tape hiss and HVAC noise. Isolate for a clean republish; optionally recast if the project needs a modern narrator tone.
Use Cases
Game development: prototype dialogue from director reads, recast into shipping character voices.
Animation and film previz: scratch tracks that sound final enough for animatic review.
Podcast and interview production: field captures cleaned to broadcast quality.
Music and remix: vocal extraction from existing songs.
Voice AI pipelines: clean speech for lipsync, transcription, or downstream models.
Marketing: A/B the same script across multiple brand voices without re-recording.
Archival restoration: clean legacy audio and optionally re-voice.
Tips for Better Results
Isolate before you change on noisy sources. Voice Changer's built-in noise removal helps, but Voice Isolator is stronger on music and heavy ambience.
Test a 30 second clip first. Validate isolation or voice match before processing a long file.
Read with the emotion you want kept. Voice Changer swaps the speaker, not the performance. Urgent lines should be read urgently.
Use clones for recurring characters. Presets are fast for exploration; clones stay consistent across a season or game.
Lock the seed when comparing voices. Same seed, same settings: the only variable is the target voice.
Add room tone back in the mix if needed. Isolated speech can feel dry for film. Blend a little ambience in your DAW after isolation.
Start from video when the audio only lives in a clip. Feed the video into Voice Isolator instead of extracting manually first.
Known Limitations
Voice Isolator needs speech to isolate. Pure music or crowd-only recordings have no voice target.
Heavily processed vocals are harder to split. Deep reverb tails, vocoder effects, and dense harmonies reduce fidelity.
Voice Changer is not a re-recorder. Pops, clipping, and mispronunciations in the source persist in the output.
No partial-file control. Both models process the full upload uniformly.
Cloning happens upstream. Create voices with ElevenLabs Voice Clone before selecting them in Voice Changer.
Plan access may apply. Both models carry access restrictions on some workspaces.
Open the models directly: ElevenLabs Voice Isolator · ElevenLabs Voice Changer
Voice Clone
You can build a custom ElevenLabs voice from your own audio on Scenario, then reuse that voice in speech and dubbing-style flows. Instant cloning is built for speed. Professional cloning is built for higher fidelity and may ask for an extra identity check. The PVC Verify model exists for that check: one short recording of the real speaker reading a captcha phrase.
How the clone flow works on Scenario
Open Voice Clone, add your samples (or rely on samples already on the model), then set Language to match the speech in those files. Choose Instant when you need a usable voice quickly for drafts, scratch narration, or rapid iteration. Choose Professional when you want the deeper pass and are ready for possible extra verification.
Add a Voice Name and optional Description so future you (and collaborators) know which character or brand this belongs to. If your recordings include air conditioning hum or cafe noise, try Remove Background Noise, but trim or re-record first if clips are shorter than five seconds.
When the provider asks for captcha verification after a Professional clone, switch to PVC Verify. Record one clean take of the owner reading the phrase exactly as displayed, then upload that file to Captcha Recording and run the job.
ElevenLabs Voice Clone PVC Verify (model_elevenlabs-voice-clone-pvc-verify)
This model is intentionally minimal. Scenario exposes a single audio upload that lines up with ElevenLabs captcha verification for Professional Voice Clones.
Captcha Recording (
captchaRecording): required. One audio file of the voice owner reading the captcha phrase shown in your verification flow.
The verify model does not replace Voice Clone. It does not upload training samples or pick Instant versus Professional. It only submits the verification recording when that step is required.
Examples
These walkthroughs show typical creative paths through the fields above. They are scenario-style guides, not a gallery of pinned outputs.
Game studio: hero bark-in and combat barks
A narrative designer exports ten clean lines from the voice actor session as WAV files. They run Voice Clone with Clone Type Instant, Language set to English, Voice Name set to the hero codename, and Description set to “confident, mid energy.��� They generate speech in a TTS step using the new voice for rapid cinematic prototyping. If marketing later requests a higher fidelity narrator match, they re-run with Professional and complete PVC Verify when prompted.
E-learning: one instructor voice across a course pack
An instructional designer collects two minutes of room tone-free narration from the SME. They enable Remove Background Noise only after confirming each clip is longer than five seconds. They pick Professional, set Language to match the course language, and add Gender and Age hints that match the SME on camera. After verification, lesson generation uses one consistent voice across dozens of scripts.
Social trailer: fast scratch track, then polish
A small team clones the founder’s voice with Instant for a same-day vertical trailer VO. They iterate script timing in Scenario. Before the paid boost, they run Professional on a longer studio session file, then upload the captcha read through PVC Verify so the public-facing export uses the verified clone.
Brand assistant: named voice for internal demos
A marketing lead uploads approved sample reads, sets Voice Name to the assistant product name, and writes Description as “warm, concise, no slang.��� Language matches the market they ship first. They keep Clone Type on Instant until legal signs off, then move to Professional for the launch kit.
Verification only: captcha redo after a failed attempt
A user’s first captcha take included overlapping roommates. They re-record in a closet with soft furnishings, read the phrase once without repeating lines, export MP3, and run PVC Verify again with only Captcha Recording filled.
Use Cases
Games and cinematics: Lock a character voice across branching dialogue and pick Instant for rapid blocking or Professional for hero scenes.
Marketing and product: Keep founder or brand spokesperson tone consistent across landing pages, demos, and update videos.
Film and podcast production: Build a narrator profile from clean dailies, then satisfy verification when Professional cloning requires it.
Education: Turn a teacher’s natural delivery into a reusable voice for slide narration and micro-lessons.
E-commerce and support: Produce on-brand explainers and FAQ audio that sounds like the same human spokesperson.
Social video: Draft VO fast with Instant, then swap to a verified Professional clone for sponsored posts.
Tips for Better Results
This article was prepared without a Step 3 example batch on Scenario. Tips 1 to 3 come directly from Scenario field rules. Tips 4 to 7 summarize ElevenLabs public help content about cloning and captcha verification. Always follow on-screen prompts in your own workspace if they differ.
Match Language to the speech you uploaded. The control sets how the clone interprets pronunciation, not a translation task.
If Remove Background Noise is on, keep every clip at least five seconds. Shorter files can block the job or weaken cleaning.
Use Gender and Age only for Professional when they help listeners. Leave them blank if you do not need that hint.
Instant is for speed-first drafts. ElevenLabs positions Instant Voice Cloning as the quick path from small amounts of clear audio.
Professional expects higher quality source audio. ElevenLabs documentation stresses clean studio-style material for Professional Voice Cloning because the model learns directly from your samples.
Read the captcha phrase once, cleanly. ElevenLabs help articles warn that repeating lines inside one recording can fail verification.
Match verification tone to training audio. ElevenLabs suggests sounding similar to the samples you already uploaded so the check can succeed.
Dubbing (v1)
ElevenLabs Dubbing enables video and audio localization by translating spoken audio in any video or audio file into a different language while preserving the original speaker's voice, emotion, and timing. The service automates transcription, translation, and voice synthesis.
Key capabilities:
Input support: raw video files, finished exports, or standalone audio files.
Vocal fidelity: maintains nuance, pace, and emotional energy from original performances.
Output: dubbed versions with translated audio replacing original speech.
Parameters
Parameter | Type | Required | Default | Description |
File | File (audio or video) | Yes | The audio or video asset to dub. | |
Target Language | String | Yes | en | Output language code (ISO 639-1). See supported languages below. |
Source Language | String | No | Auto-detect | Original file language. Leave empty for automatic detection. |
Number of Speakers | Integer | No | 0 | Speaker count. 0 = auto-detect. Max 10. |
Drop Background Audio | Boolean | No | false | Removes music and ambience from the dubbed output when enabled. |
Disable Voice Cloning | Boolean | No | false | Uses a generic voice instead of cloning the original speaker. |
Highest Resolution | Boolean | No | false | Returns the dubbed video in the highest resolution available. |
Supported Languages
The model supports 30 languages:
Arabic, Bulgarian, Chinese, Croatian, Czech, Danish, Dutch, English, Finnish, French, German, Hindi, Hungarian, Indonesian, Italian, Japanese, Korean, Malay, Norwegian, Filipino, Polish, Portuguese, Romanian, Russian, Slovak, Spanish, Swedish, Thai, Turkish, Ukrainian, Vietnamese.
Features
Voice Cloning
By default, the model clones the original speaker's voice for each dubbed track. For best results, speakers should have at least 30 seconds of clean, clear audio. Set Disable Voice Cloning to true to use a generic synthesized voice instead.
Multi-Speaker Support
The model automatically detects up to 10 distinct speakers. For scenes with overlapping dialogue, set the Number of Speakers manually to improve accuracy.
Background Audio Control
Enable Drop Background Audio to strip music and sound design from the output. This is useful when dubbing into languages where background audio may interfere with voice clarity, but note the control is binary: complex sound design that needs partial preservation requires separate post-production mixing.
Best Practices
Prioritize clean source audio for optimal voice cloning quality.
Set speaker count manually when dialogue overlaps between multiple speakers.
Leave Highest Resolution disabled during testing to reduce processing time.
Ensure each speaker has at least 30 seconds of clear audio for effective voice learning.
Known Limitations
Optimized for up to 10 distinct speakers.
Dubbed audio is time-stretched to fit original visuals. Longer target languages may sound slightly accelerated.
The background audio toggle is binary. Complex sound design requires separate post-production mixing.