ElevenLabs Music, Speech & Sound Effects: The Essentials
Last updated: August 14, 2026
This guide covers the ElevenLabs audio-generation tools on Scenario. Text-to-Speech turns written text into natural narration across many voices and languages. Music v2 and the earlier Music generate full songs from a prompt. Sound Effects generates SFX from a plain text description. Each model has its own section below with examples, parameters, and tips.
Text-to-Speech
Scenario’s text-to-speech tools let you create professional-quality voiceovers from text. Among the available models, three ElevenLabs models stand out: ElevenLabs v3, Multilingual v2 and Turbo 2.5. This guide explains each model’s unique strengths, how to choose between them, and tips for getting the best results.
Multilingual v2 supports 29 languages with rich emotional expression and natural prosody. Turbo 2.5 supports 32 languages (adds Vietnamese, Hungarian, Norwegian) with 3x faster generation for non-English languages and 25% faster English generation.
1. Overview
ElevenLabs v3 (alpha)
Eleven v3 is a research‑preview model that produces natural, life‑like speech with a high emotional range. Unlike v2 and Turbo, it is not designed for real‑time applications and is better suited to long‑form narration and expressive dialogue. The model supports more than 70 languages. Because it is in alpha, quality can vary; we recommend experimenting with longer prompts (≥250 characters) and multiple generations.
ElevenLabsMultilingual v2
Multilingual v2 is ElevenLabs’ most lifelike and emotionally rich production model. It delivers consistent voice quality and natural prosody across 29 languages, making it ideal for audiobooks, film dubbing, podcasts and other projects where emotional fidelity matters. V2 prioritizes quality over speed, resulting in higher latency and cost compared with Turbo. You can generate up to 10 000 characters per request.
Supported languages: English, Spanish, French, German, Italian, Portuguese, Russian, Japanese, Korean, Chinese (Mandarin), Arabic, Hindi, Dutch, Polish, Czech, Slovak, Ukrainian, Croatian, Romanian, Bulgarian, Greek, Finnish, Danish, Swedish, Norwegian, Hungarian, Turkish, Hebrew, Malay, Tamil.
ElevenLabsTurbo 2.5
Turbo 2.5 balances quality with low latency. It supports 32 languages, adding Vietnamese, Hungarian and Norwegian to the v2 language set. Generation speed is roughly three times faster than v2 for non‑English languages and 25 % faster for English, and the model is about 50 % cheaper per character. This makes Turbo ideal for real‑time conversational agents, interactive games and high‑volume projects. Turbo supports up to 40 000 characters per call and allows manual language enforcement via two‑letter ISO 639‑1 codes.
2. Model selection
2.1 Choosing ElevenLabs v3
Storytelling and character dialogue – Use v3 when you need highly expressive performances for multi‑speaker conversations, audiobooks or dramatic scenes. The model’s emotional range and contextual understanding provide realism that v2 and Turbo cannot match.
Non‑real‑time projects – v3 is not optimized for real‑time; its higher latency and 3 000‑character limit favour offline workflows where you can generate several takes and choose the best.
Multilingual content – With support for 70+ languages, v3 is a good choice when you need expressive narration in less‑common languages.
2.2 Choosing Multilingual v2
High‑fidelity narration – Select v2 when natural prosody and emotional nuance are paramount, such as for audiobooks, podcasts, voiceovers and educational content.
Stable quality – V2 maintains consistent voice personality across long passages and multiple languages.
Language coverage – Use v2 when your content is in one of its 29 supported languages and you require emotional richness.
Content length – v2’s 10 000‑character limit per call supports longer audio segments than v3.
2.3 Choosing Turbo 2.5
Real‑time interaction – Turbo’s low latency and cost make it suitable for chatbots, games and other interactive applications.
Cost efficiency – Its per‑character price is roughly half that of v2, making it economical for large volumes of speech.
Language flexibility – Turbo supports 32 languages and allows manual language selection via ISO codes.
Longest content – With a 40 000‑character limit, Turbo can generate extended scripts in a single call.
3. Key differences
Model | Latency & use | Emotional range & quality | Languages | Character limit |
ElevenLabs v3 (alpha) | Not real‑time; suited to offline projects | Highest emotional range and contextual understanding | 70+ languages | 3 000 characters |
Multilingual v2 | Higher latency and cost; prioritizes quality | Lifelike speech with rich emotional expression | 29 languages | 10 000 characters |
Turbo 2.5 | Low latency and 50 % cheaper per character | Balanced quality; less emotional nuance than v2 | 32 languages | 40 000 characters |
4. Interface controls & workflow
4.1 Text input
In the Scenario interface, type or paste your script in the text field. ElevenLabs models automatically detect the language and can handle multilingual content within a single generation. Use proper punctuation and capitalization to guide rhythm and emphasis; ellipses (…) add pauses and capitalization signals emphasis.
4.2 Voice selection
All three models share the same voice library. Choose a voice that matches the desired delivery; neutral voices tend to be more stable across languages. For v3, voice selection is especially critical because the model responds strongly to voice characteristics.
ElevenLabs' voice library including Aria, Roger, Sarah, Laura, Charlie, George, Callum, River, Liam, Charlotte, Alice, Matilda, Will, Jessica, Eric, Chris, Brian, Daniel, Lily, and Bill. Each voice works with both models.
Aria: female, expressive, social, engaging.
Roger: male confident, social, persuasive.
Sarah: female, expressive, social, energetic.
Laura: female, upbeat, social, lively.
Charlie: male, natural, conversational, relaxed.
George: male, warm, narration, trustworthy.
Callum: male, intense, character, dramatic.
River: non-binary, confident, social, modern.
Liam: male, articulate, narration, clear.
Charlotte: female, seductive, character, playful.
Alice: female, confident, news, formal.
Matilda: female, friendly, narration, calm.
Will: male, natural, narration, steady.
Jessica: female, expressive, conversational, youthful.
Eric: male, friendly, conversational, approachable.
Chris: male, casual, conversational, easy going.
Brian: male, deep, narration, serious.
Daniel: male, authoritative, news, commanding.
Lily: female, warm, narration, gentle.
Bill: male, trustworthy, narration, classic.
4.3 Generation parameters (v2 & Turbo)
Multilingual v2 and Turbo 2.5 provide the following controls:
Stability (0-1, default 0.5)
Controls consistency and predictability of speech. Higher values produce more stable, consistent output. Lower values allow more variation and expressiveness.Similarity Boost (0-1, default 0.5)
Enhances similarity to the selected voice characteristics. Higher values make the output more closely match the chosen voice profileStyle Exaggeration (0-1, default 0)
Controls emotional intensity and expressiveness. Higher values increase dramatic emphasis and emotional range. More effective with Multilingual v2Speed (0.7-1.2, default 1)
Adjusts speech rate. Values below 1.0 slow down speech, above 1.0 speed it up. Extreme values may affect quality.Timestamps Toggle
When enabled, returns timestamps for each word in the generated speech, useful for synchronization applications.
4.4 Advanced Features
Advanced features include Previous Text/Next Text fields for chaining long scripts and a Language Code parameter to enforce a specific language.
Previous Text / Next Text
Allows chaining multiple text segments for longer content generation while maintaining voice consistencyLanguage Code
Manually specify language using ISO 639-1 codes to enforce specific language pronunciation when automatic detection isn't sufficient.
5. Best practices (all ElevenLabs models)
Text Formatting
Use proper punctuation for natural pacing. Periods create pauses, commas add brief breaks, and exclamation points increase energy. Break long paragraphs into shorter sentences for better flow.
Voice Consistency
Use the same voice and similar parameter settings across related content. Enable Timestamps when you need precise synchronization with other media.
Parameter Tuning
Start with default settings and adjust based on results. Higher Stability for consistent content, higher Style Exaggeration for dramatic effect, adjusted Speed for pacing preferences.
Model Selection
Use Multilingual v2 for final production content where quality matters. Use Turbo 2.5 for prototyping, real-time applications, or when speed is the priority.
6. Best practices (ElevenLabs v3)
ElevenLabs v3 introduces unique settings and tags for fine‑grained emotional control:
Longer prompts – Prompts shorter than ~250 characters may yield inconsistent output; longer prompts improve stability.
Stability modes – v3 offers Creative, Natural and Robust modes. Creative provides expressive output but may hallucinate; Natural balances expressiveness and accuracy; Robust is highly stable but less responsive. Use Creative or Natural when employing audio tags.
Audio tags – Use tags such as
[laughs],[whispers],[sarcastic],[curious],[excited]or sound effects like[gunshot],[applause],[clapping]to control emotion and add effects. Some tags may work better with certain voices; test combinations to find what works.Punctuation and capitalization – Ellipses create pauses; capitalization adds emphasis; proper punctuation improves natural rhythm.
7. Practical prompt examples
Audiobook narration (v2) – “Chapter One: The discovery. Sarah walked through the ancient library, her footsteps echoing in the silence….”. Use high Stability and Similarity Boost values; moderate Style Exaggeration (0.1–0.3) for subtle emotion.
Emotional dialogue (v2/v3) – “I can’t believe you’re leaving! After everything we’ve been through together, how can you just walk away like this means nothing?”. Add tags like
[crying]or[angry]in v3; increase Style Exaggeration in v2.Educational content (v2) – “Today we’ll explore the fascinating world of quantum physics. Don’t worry if it seems complex at first – we’ll break it down step by step.”. Choose a calm voice and set Stability high.
Conversational AI (Turbo) – “Hi there! How can I help you today? I’m here to answer your questions and assist with whatever you need.”. Lower Stability and increase Speed for snappier responses; set Language Code to ensure the desired language.
Multilingual prompts – “Welcome to our international conference. Bienvenue à notre conférence internationale. Bienvenidos a nuestra conferencia internacional.”. Use v2 or Turbo with automatic language detection or specify a Language Code for each segment.
6. Optimization Settings by Use Case
6.1 Audiobook/Podcast (Multilingual v2)
For audiobook and podcast production, choose expressive voices like Aria or Sarah that can convey narrative emotion effectively. Set Stability between 0.6-0.8 to ensure consistent narration throughout longer content, while using a Similarity Boost of 0.7 to maintain strong voice consistency across chapters or episodes. Apply moderate Style Exaggeration of 0.3-0.5 for natural expression that engages listeners without overwhelming the content. Keep Speed between 0.9-1.0 to create a comfortable listening pace that allows for proper comprehension and enjoyment.
6.2 Conversational AI (Turbo 2.5)
Conversational AI applications work best with natural-sounding voices like Charlie or Laura that feel approachable and friendly. Use Stability settings of 0.4-0.6 to allow slight variation that makes conversations feel more human and less robotic. Set Similarity Boost to 0.5 for balanced consistency that maintains character while allowing natural speech variation. Keep Style Exaggeration minimal at 0-0.2 to maintain a neutral, professional tone appropriate for most conversational contexts. Speed should be set to 1.0-1.1 to create a responsive feel that matches natural conversation pacing.
6.3 Dramatic Content (Multilingual v2)
Dramatic content requires expressive voices like George or Jessica that can handle emotional range and character depth. Lower Stability to 0.3-0.5 to allow for emotional variation that brings characters to life and supports dramatic storytelling. Use Similarity Boost of 0.6 to maintain character consistency while allowing for emotional expression. Increase Style Exaggeration to 0.6-0.8 for dramatic emphasis that enhances the emotional impact of the content. Reduce Speed to 0.8-0.9 for dramatic pacing that gives weight to important moments and allows emotional beats to resonate.
6.4 Educational Content (Both Models)
Educational content benefits from clear, articulate voices like Brian or Alice that prioritize comprehension and clarity. Set Stability to 0.7 for consistent delivery that helps students focus on the content rather than vocal variations. Use Similarity Boost of 0.6 for reliability that ensures consistent voice characteristics across lessons or modules. Apply moderate Style Exaggeration of 0.2-0.4 to create engaging but clear speech that maintains student interest without distracting from the educational material. Set Speed to 0.9 to optimize for comprehension, giving students time to process complex information while maintaining engagement.
9. Creative Applications
Audio Content Creation
Create audiobook narrations, podcast intros, video voiceovers, and educational content. Use Previous Text/Next Text features for longer content while maintaining voice consistency.
Interactive Applications
Build conversational AI systems, voice assistants, interactive games, and real-time communication tools. Turbo 2.5's speed makes it ideal for responsive applications.
Multilingual Projects
Develop content for global audiences using automatic language detection or manual language specification. Both models maintain voice characteristics across different languages.
10. Troubleshooting
Quality Issues
If speech sounds robotic, lower Stability and increase Style Exaggeration. If pronunciation is incorrect, try different punctuation or manual Language Code specification for Turbo 2.5.
Speed vs Quality Balance
For real-time applications needing better quality, try Turbo 2.5 with higher Stability settings. For high-quality content needing faster generation, use Multilingual v2 with optimized parameters.
Voice Consistency
Use identical parameter settings and the same voice selection across related content. The Similarity Boost setting helps maintain consistent voice characteristics.
Conclusion
Scenario’s ElevenLabs portfolio features Eleven v3, Multilingual v2 and Turbo 2.5, offering a spectrum from high expressiveness to real‑time efficiency. By understanding each model’s strengths, selecting suitable voices, tuning generation parameters and crafting well‑structured prompts, you can produce professional‑quality audio tailored to your use case.
Music v2
ElevenLabs Music v2 generates original, studio-quality music from your words. Two models share the same engine: one takes a single text prompt and writes a whole track, the other lets you build a song section by section with precise control over structure, lyrics, and style. Both are trained on licensed data and cleared for commercial use.
30 s lo-fi instrumental bed from Music v2.
Which Model Should I Use?
Model | ID | How you drive it | Best for |
|---|---|---|---|
ElevenLabs Music v2 Prompt |
| One text prompt + duration | Fast, complete tracks: beds, loops, song sketches, background music |
ElevenLabs Music Advanced v2 Composer |
| Ordered list of sections | Full songs with defined structure, lyrics, and section-level control |
Rule of thumb: reach for Music v2 when you want a great track fast from one description. Reach for Advanced v2 when you need to control how the song is built, intro into verse into chorus into bridge into outro, with different styles or lyrics in each part.
How to Use the Models
Music v2: one prompt, one track
Write one prompt that names the genre, mood, instruments, and tempo, then set how long you want it. The model composes and produces the full track, including vocals unless you turn them off. Listen the music.
prompt: Warm, mellow lo-fi hip-hop with dusty vinyl crackle, soft Rhodes piano,
laid-back boom-bap drums and upright bass. Relaxed study mood, around 75 BPM.
durationSeconds: 30
forceInstrumental: true
outputFormat: mp3_44100_192For songs with singing, leave forceInstrumental off and the model writes and performs its own topline. Naming a tempo in BPM and listing specific instruments gives tighter, more predictable results than vague adjectives alone.
prompt: Upbeat indie-pop with jangly electric guitars, punchy live drums,
warm bass, bright female vocals about summer nights
durationSeconds: 30
outputFormat: mp3_44100_192Advanced v2: section-by-section songs
On ElevenLabs Music Advanced v2:
Advanced v2 builds a song from an ordered list of sections. The text field uses three kinds of content:
[Square brackets]mark structure:[Intro],[Verse],[Chorus],[Bridge],[Outro].Plain text is treated as sung lyrics. Anything without brackets will be sung.
{Curly braces}are performance directions that guide the music without being sung.
For an instrumental section, do not write arrangement notes as plain text. Put them in {curly braces} and add "vocals" to that section's negativeStyles. Plain text will be sung aloud.
sections:
- text: "[Intro] {soft solo piano, sparse, rubato}"
durationSeconds: 8
positiveStyles: ["acoustic pop", "warm", "fingerpicked guitar"]
negativeStyles: ["drums", "vocals"]
contextAdherence: "high"
- text: "[Verse] Morning light spills through the open door, I'm chasing dreams I've never had before"
durationSeconds: 20
positiveStyles: ["acoustic pop", "gentle female vocals", "singer-songwriter"]
contextAdherence: "high"
- text: "[Chorus] So I'll run, run, run toward the rising sun, this is just the start of something we've begun"
durationSeconds: 20
positiveStyles: ["uplifting", "catchy", "full band", "vocal harmonies"]
contextAdherence: "high"
- text: "[Outro] {soft fade with gentle humming}"
durationSeconds: 12
positiveStyles: ["gentle", "acoustic", "warm"]
contextAdherence: "medium"60 s acoustic-pop song with intro, verse, chorus, and outro. asset_oQVoeqnV9gQctTWcmfkXAg1N
Set contextAdherence to high so each section listens to its neighbors and the song flows as one piece. Keep sections inside a single musical identity: this model is for structuring one song, not for blending unrelated genres.
Using both models together
Sketch with Music v2 to find a genre and mood quickly, then rebuild in Advanced v2 with explicit sections so you control where the drop, chorus, or lyrics land. Both accept a seed, so lock a result you like and iterate on the parts around it.
Parameters
ElevenLabs Music v2
prompt
Required. Up to 2048 characters. Describe mood, genre, instruments, tempo, and any vocal direction.
durationSeconds
Optional. Default 30. Range 3 to 180 seconds. Sets the track length.
forceInstrumental
Optional. Default false. Set true to generate music with no vocals.
outputFormat
Optional. Default mp3_44100_128. Twelve MP3 or Opus options up to 192 kbps. Higher bitrates mean better quality and larger files.
seed
Optional. Random by default. Reuse the same seed and settings to reproduce a track exactly.
ElevenLabs Music Advanced v2
Advanced v2 replaces the single prompt with a sections array (1 to 20 ordered sections). It keeps outputFormat and seed from the base model, but has no single global prompt.
sections[].text
Section content. Use [labels] for structure, plain text for sung lyrics, and {directions} for performance notes that should not be sung.
sections[].durationSeconds
Optional per section. Default 10. Range 3 to 120 seconds.
sections[].positiveStyles / sections[].negativeStyles
Up to 10 style tags each. Positive tags steer the section toward a sound; negative tags steer it away (for example "vocals" on an instrumental part).
sections[].contextAdherence
Optional. Default medium. Values: low, medium, high. Controls how closely a section follows its neighbors.
Use Cases
Games: looping level themes, boss-fight builds, retro chiptune, and tension beds tuned to a scene.
Marketing: ad and social music beds, upbeat corporate backing, and short stingers in any genre.
Film and trailers: cinematic orchestral cues, emotional piano scenes, and hybrid trailer builds.
Education: simple sing-along songs and calm background music for lessons and explainers.
Content and e-commerce: lounge, bossa nova, and ambient beds for storefronts, podcasts, and videos.
Tips for Better Results
Name a tempo and real instruments. "Around 120 BPM, jangly guitars, punchy live drums" beats "upbeat song" every time.
Turn off vocals when you want a bed. Use
forceInstrumental: trueon Music v2, or add"vocals"tonegativeStyleson each Advanced section.Keep one identity per Advanced song. Vary structure and dynamics across sections, not the genre. One coherent style sounds far better than a genre mash.
Use
contextAdherence: highby default. Drop to medium only for slow builds. Low adherence with contrasting styles produces incoherent results.Mind the three text types in Advanced.
[labels]for structure,{curly}for directions, plain text only for words you want sung.Match duration to the deliverable. Short stingers need only a few seconds; full songs benefit from 90 to 180 seconds on Music v2 or multiple sections on Advanced v2.
Lock results with
seed. The same prompt, settings, and seed reproduce the same track, which is ideal for A/B variations on one idea.
Known Limitations
Near-silent or minimal content can fail. Compositions built around sub-bass drones, long silences, or "fade to silence" sections sometimes fail to generate. Give every section real musical material.
No global prompt in Advanced v2. Styles are set per section via
positiveStyles, so define the identity in every section or later sections can drift.Genre blending is not the use case. Forcing opposite genres into one song, especially with
contextAdherence: low, produces disjointed output. Keep one identity per track.Per-section duration caps at 120 seconds in Advanced v2, while Music v2 allows a single track up to 180 seconds.
Platform audio auto-captions are unreliable. Ignore them as a QA signal; listen to the output instead.
Open the models: ElevenLabs Music v2 · ElevenLabs Music Advanced v2
Music (v1)
The ElevenLabs Music family generates original music from text descriptions. The standard model takes a single prompt and delivers a complete track in seconds. The Advanced model lets you define the song's internal structure section by section, giving you precise control over how a composition builds and changes over time. Both models output studio-ready audio in MP3 or Opus format.
Which Model Should I Use?
ModelID | Input | Best for | |
Simple |
| Text prompt, duration, vocal/instrumental toggle | Quick tracks, background music, game audio, rapid iteration |
Structured |
| Global styles, per-section styles and narrative lines, up to 20 sections | Full songs with distinct acts, cinematic scores, tracks where the mood must shift at a specific moment |
Use ElevenLabs Music when you need a track fast and the overall feel is what matters. Switch to ElevenLabs Music Advanced when the song needs to tell a story or when you require deliberate transitions between energy levels, moods, or musical phases.
Parameters
ElevenLabs Music
The only required input is the Prompt, a text description of the music you want, up to 2048 characters. Describe the genre, instruments, mood, tempo, and any other qualities that matter. The more specific the prompt, the closer the result will be to your intent.
Duration sets how long the track should be, from 3 seconds to 3 minutes. The default is 30 seconds. Longer tracks give the model room to develop themes and create a proper beginning, middle, and end. Cost scales with duration.
Enable Instrumental Mode to remove all vocals from the output. This is the recommended approach for background music, game audio, and any context where lyrics would be disruptive. Even when your prompt does not mention vocals, the model may add them if the genre typically includes singing, so this toggle is the only reliable way to guarantee an instrumental result.
Output Format controls the codec, sample rate, and bitrate of the file. The default is MP3 at 44.1 kHz and 128 kbps, which is a good balance of quality and size. See the format table below for all available options.
Seed is optional. Set it to any integer to lock the random seed. The same prompt and seed will produce the same output every time, which is useful when iterating on a result you want to refine.
ElevenLabs Music Advanced
The Advanced model requires two inputs: Global Styles and Sections.
Global Styles is a list of style tags that apply to the entire song, with a maximum of 10 tags. The default is "pop". Use tags to define the overall genre, tempo, instrumentation, and mood of the track. Pair it with Excluded Styles to steer the model away from unwanted elements that might otherwise bleed into the output.
Sections is an ordered list of song phases, with a maximum of 20. Each section has:
Section Name (required): a label such as "intro", "verse", "chorus", "bridge", or "outro".
Duration (required): the target length for this section in milliseconds. The model uses this as a guide and may vary the actual length slightly at transitions.
Section Styles (optional): style tags for this section only, such as "quiet", "building tension", or "full band". These refine or shift the global styles for this segment without affecting the rest of the song.
Excluded Section Styles (optional): style tags to suppress in this section only.
Lyrics (optional): the lyrics for this section. Each entry in the list is one line of singable text. Leave this empty for instrumental sections like intros and outros.
Output Format and Seed work the same as in ElevenLabs Music.
Output Format Options
The default is MP3 at 44.1 kHz and 128 kbps, a solid balance between file size and quality. For final deliverables, use MP3 at 192 kbps or Opus at 48 kHz / 192 kbps for the highest available quality, best for archiving and professional use. For the smallest possible file, the 22 kHz / 32 kbps MP3 option works but at a noticeable quality cost. Additional MP3 and Opus variants at lower bitrates are available for bandwidth-constrained use cases.
How ElevenLabs Music Works
ElevenLabs Music takes your text description and generates a complete audio track from scratch. The model interprets genre, instrumentation, mood, tempo, and any other qualities you describe, combining them into a coherent musical output. The more specific your prompt, the more accurately the result matches your intent.
Duration is controlled with musicLengthMs. Shorter tracks (under 30 seconds) tend to be complete musical ideas. Longer tracks (60 to 180 seconds) give the model room to introduce variation, develop themes, and provide a proper beginning, middle, and end. CU cost scales with duration at roughly 2.5 CU per second of output.
The forceInstrumental flag removes all vocals from the output. It works reliably and is the recommended approach for any background music, game audio, or ambient use case where lyrics would be intrusive.
Prompt example (instrumental):
"Retro 80s synthwave with pulsing arpeggios, gated reverb drums, neon-soaked atmosphere"
musicLengthMs: 45000
forceInstrumental: true
outputFormat: "mp3_44100_192"
How ElevenLabs Music Advanced Works
ElevenLabs Music Advanced structures a song as an ordered list of named sections. Each section has its own duration, local style tags, and narrative description. The model reads all sections together and composes a track where each phase transitions naturally into the next, while staying consistent with the global style tags applied to the whole song.
Global styles set the sonic identity for the entire track. Local styles refine or shift the character within a single section. This layered approach lets you build a song that starts quietly and ends with a full-band climax, or a cinematic score that moves through tension, action, and resolution without sounding like three separate tracks edited together.
The lines field inside each section is where you provide the actual lyrics for that part of the song. Each string in the array is one line of singable text. If you leave lines empty for a section, the model will generate vocals without fixed lyrics or produce an instrumental passage depending on the style tags.
Sections example (film score):
positiveGlobalStyles: ["cinematic", "orchestral", "epic"]
negativeGlobalStyles: ["electronic"]
sections: [
{
sectionName: "tension buildup",
durationMs: 15000,
positiveLocalStyles: ["suspenseful", "low strings", "mounting dread"],
lines: ["Slow string ostinato builds with rising dissonance"]
},
{
sectionName: "battle",
durationMs: 25000,
positiveLocalStyles: ["explosive", "full orchestra", "relentless"],
lines: ["Full orchestra erupts into fierce battle music with pounding drums"]
},
{
sectionName: "resolution",
durationMs: 10000,
positiveLocalStyles: ["victorious", "warm", "resolving"],
lines: ["Triumphant resolution with warm brass and final chord"]
}
]Use Cases
Game audio: Generate background music for menus, levels, and cutscenes. Use
forceInstrumental: trueand match the genre to the game's visual tone. Use the Advanced model to create tracks that shift from ambient exploration to intense action within a single asset.Video and film scoring: Score trailers, short films, and video ads. The Advanced model excels here, letting you sync musical sections to visual beats (quiet opening, rising action, climax, resolution) without needing a digital audio workstation.
Marketing and social content: Create branded background tracks that match campaign visuals. Quick iteration with the standard model allows testing different genres and moods before committing to a final direction.
Podcast and video production: Generate intros, outros, and transition stings. Short tracks (3 to 10 seconds) work well as bumpers. Longer tracks (60 to 90 seconds) serve as episode background music.
Prototyping and creative exploration: Use the standard model to explore a large space of genres and styles quickly. Once a direction is confirmed, switch to the Advanced model to build a polished, structured version.
Education and e-learning: Generate calm, non-distracting instrumental backgrounds for explainer videos, course content, and presentations. A 90 to 120-second loop at low bitrate keeps file sizes manageable.
Tips for Better Results
Be specific about instruments. Prompts that name specific instruments ("nylon string guitar", "upright bass", "brushed snare") produce more accurate results than genre labels alone. Genre tags set the overall feel, but instrument names define the texture.
Always set forceInstrumental for background tracks. Even when your prompt does not mention vocals, the model may add them if the described genre typically includes singing. Setting
forceInstrumental: trueis the only reliable way to guarantee an instrumental output.Use seed for reproducibility. If you generate a result you want to iterate on, note the seed value and reuse it with modified parameters. Changing only the duration or format with the same seed gives you a consistent base to work from.
For the Advanced model, keep local styles complementary to global ones. Local style tags work best when they refine, rather than contradict, the global styles. If your global style is "orchestral", a local tag of "electronic drop" will produce conflicting output. Use local tags to shift intensity or instrumentation within the same sonic universe.
Use the lines field for actual lyrics. In the Advanced model, the
linesarray inside each section is the lyrics field. Each string is one line of singable text the model will use as vocal content. Leave it empty on instrumental sections such as intros and outros.Scale duration to the number of sections. A two-section song with each section at 60 seconds gives the model enough time to develop each phase. Very short sections (under 5 seconds) may not render distinctly. Aim for at least 8 to 10 seconds per section for audible transitions.
Use opus_48000_192 for final deliverables. The default MP3 format is fine for drafts and iteration. For final assets going into productions, switch to
opus_48000_192for the highest available quality. The file size increase is modest compared to the quality gain.
Known Limitations
Maximum duration is 3 minutes (180 seconds). The Scenario implementation caps
musicLengthMsat 180,000 ms. The ElevenLabs provider API supports up to 600 seconds, but this extended range is not available through Scenario.Section durations are approximate. The
durationMsvalue in each section of the Advanced model is a guide, not a precise boundary. The model may extend or shorten a section slightly to produce a natural transition. Hard time-sync against external video requires post-production trim.No PCM output format. The ElevenLabs API supports PCM audio output (raw waveform). This format is not exposed in the Scenario implementation. Use
opus_48000_192as the highest-quality alternative.Music Finetunes not supported. The ElevenLabs API includes the ability to fine-tune music generation on custom audio. This feature is not available through Scenario.
No store-for-inpainting option. The provider API supports a flag to store audio for future inpainting (editing a specific segment of a generated track). This parameter is not exposed in Scenario.
Vocal content in instrumental outputs at low forceInstrumental precedence. In rare cases where the global or local style strongly implies a vocal genre, faint vocal artifacts may appear even with
forceInstrumental: true. Re-generating with a more strongly instrumental style description (e.g. adding "no vocals", "fully instrumental" to the prompt or style tags) reduces the occurrence.CU cost scales with total audio duration. For the Advanced model, CU cost is determined by the sum of all section durations. A five-section track totaling 90 seconds costs the same as a single 90-second track in the standard model. Budget accordingly when planning large batches.
Sound Effects (SFX)
1. Overview
ElevenLabs Sound Effects v2 represents a major advancement in AI-powered sound design, enabling creators to generate high-quality sound effects directly from text descriptions. Built on ElevenLabs' advanced audio AI, the model understands acoustic concepts, environmental contexts, and sonic characteristics to produce professional-grade audio.
Unlike traditional sound libraries limited to pre-recorded samples, Sound Effects v2 democratizes sound creation by translating natural language into custom audio effects. It captures complex acoustic relationships, from material properties to spatial environments, making it accessible for both professionals and beginners.
Key strength: Sound Effects v2 generates contextually appropriate audio for films, games, podcasts, and other media — always consistent, realistic, and tailored to your creative vision. The model excels at creating environmental sounds, mechanical effects, organic textures, impact and collision sounds, plus atmospheric and specialized audio effects for comprehensive sound design needs.
ElevenLabs Sound Effects v2 generates up to 30 seconds of professional-grade audio in multiple MP3 formats, suitable for direct integration across various media projects.
2. Getting Started
Model Selection
In Scenario, choose ElevenLabs Sound Effects v2 from the Generate Audio section. Use descriptive prompts specifying source, environment, and acoustic traits for best results.
Interface Overview
Sound Description Field – Enter your prompt
Duration Control – Set length (up to 30 seconds). If None, optimal duration will be determined from the prompt.
Guidance Control – Adjust model adherence to prompt
Loop Toggle – To create seamless audio loops
Output Format – Select audio quality and format
3. Crafting Effective Sound Descriptions
3-1. Basic Structure
Prompts could typically include:
Sound Source: footsteps, car engine, rain
Material/Surface (if relevant): gravel, wooden floor, metal
Environment: hall, outdoors, small room
Intensity/Quality: heavy, gentle, echoing
Example: “Heavy footsteps on wooden floorboards in an old creaky house.”
3-2. Advanced Techniques
You may want to add depth to your prompts, with:
Acoustic Properties: with reverb, dry, muffled
Temporal Elements: fading, sudden, continuous
Contextual Details: during a storm, underwater
Emotional Tone: ominous, peaceful, mysterious
Example: “Distant thunder rumbling across a vast open plain during a summer storm, gradually building in intensity.”
3-3. Categories & Examples
Environmental: “Gentle ocean waves lapping against a rocky shore.”
Mechanical: “Vintage typewriter keys clicking rhythmically.”
Impact: “Glass bottle breaking on concrete pavement.”
Organic: “Sizzling bacon in a hot cast iron pan.”
4. Generation Controls & Settings
Set duration between 1-30 seconds (3-10 seconds optimal for effects, longer for ambience).
Use high guidance (0.7-1.0) for precise results or low guidance (0.1-0.5) for creative variations.
Enable loop for continuous sounds like rain or hums.
Choose output format based on use: mp3_44100_128 for professional projects, mp3_44100_96 for balanced quality, mp3_44100_64 for web, or mp3_22050_32 for mobile applications.
5. Asset Management & Workflow
Generated sound effects can be previewed directly in Scenario, downloaded in your chosen format, pinned to favorites for easy access, tagged with custom labels for organization, and shared with team members using Scenario's collaboration tools.
6. Creative Applications
Use Eleven Labs Sound Effects v2 to custom foley, ambient soundscapes, UI sounds, and unique effects for:
Film & video post-production
Podcasts & voiceovers
Game design
Commercial & marketing content
7. Best Practices
Be specific about source, environment, and material
For mechanical sounds, describe condition and operation
For environmental sounds, specify weather, time, and space
Avoid combining unrelated sounds in one prompt
Use guidance and looping strategically
8. Example Prompts by Category
Environmental & Atmospheric
“Gentle rain falling on a tin roof during a quiet evening.”
“Peaceful forest ambience with chirping birds and rustling leaves.”
Mechanical & Industrial
“Vintage steam locomotive chugging slowly down railroad tracks.”
“Mechanical clock ticking steadily in a quiet room.”
Impact & Collision
“Heavy book dropping onto a hardwood table.”
“Ceramic mug shattering on kitchen tile floor.”
Organic & Natural
“Dry leaves crunching underfoot during an autumn walk.”
“Paper pages rustling as someone flips through a book.”
Specialized Effects
“Magical sparkles with ethereal chiming tones.”
“Sci-fi laser beam charging up and firing.”
9. Troubleshooting & Advanced Tips
If results are off, simplify prompts and refine descriptions
Add environmental or material details for accuracy
Adjust guidance for precision vs creativity
Use longer durations + looping for ambient effects
Pro tip: Build custom libraries by generating reusable effects, replacing stock samples, and creating unique audio tailored to your projects.