Gemini TTS: The Essentials
Last updated: September 24, 2026

Scenario hosts four generations of Google's Gemini text-to-speech models. The current pair is Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS: both read your script word for word in one of 30 preset voices and 24 language locales, both understand the new inline vocal tags like <laugh>, <sigh> and <long pause>, and both can voice a two-character scene in a single job. Flash is the quality tier, tuned for acting and nuance. Flash-Lite is the speed tier, tuned for fast, clean, high-volume reads.
Gemini 3.1 Flash TTS and Gemini 2.5 TTS remain available for existing workflows and are covered at the end of this article.
Which Model Should I Use?
Model | Generation | Best for |
Current, quality tier | Acting and performance: audiobooks, game dialogue, trailers, podcasts, character voiceover, heavy use of vocal tags | |
Current, speed tier | Volume and utility reads: voice agents, IVR menus, announcements, e-learning, navigation, read-aloud | |
Previous | Existing workflows built on square-bracket tags like [whispers] | |
Legacy | Existing workflows only |
Start with 3.8 Flash whenever the read is a performance: characters, emotion, narration, anything a listener will judge as acting. Use 3.8 Flash-Lite when the read is functional and you need a lot of it, fast. Both take exactly the same inputs, so a script written for one runs unchanged on the other.
How to Use Gemini 3.8 TTS
Write the script, not the stage directions
The biggest change in 3.8 is that the Text field is treated as a verbatim transcript. Every word you type is spoken. That means directions in parentheses or plain prose, like "(whispering nervously)" or "speak slowly and calmly", are read out loud as dialogue. We tested this directly: both were spoken word for word. Old 3.1-style square-bracket tags like [scared] are not spoken, but they are silently ignored, so they have no effect either.
Shape the performance with three tools instead: the words and punctuation of the script itself, the voice you pick, and inline vocal tags in angle brackets for specific moments. Here is a single-voice read with Flash, voice Charon, using a sigh, a chuckle and two pauses:
Day forty-one on the island. <sigh> The storm finally broke around four this morning, and for the first time in a week I could hear the gulls instead of the wind. <short pause> I walked the rocks at low tide and found a bottle wedged between two stones, green glass, cork still in it. <chuckle> Of course I opened it. Inside was a map of this very island, drawn by hand, with an X marked right where my lighthouse stands. <long pause> Whoever drew it has been here before me. And judging by the ink, they were here quite recently.Gemini 3.8 Flash, voice Charon, en-US: lighthouse keeper's log · Open on Scenario
The pause tags are real, measurable silences. In this clip the two <long pause> moments come out as gaps of about 2.1 and 2.5 seconds, against roughly half a second between ordinary sentences.
Inline vocal tags
Vocal tags mark a momentary, non-speech event at an exact point in the script. Place the tag where the sound should happen, usually between sentences. The supported set is:
<argh>, <breath>, <heavy breath>, <exhales>, <cackle>, <cheer>, <chuckle>, <cough>, <cry>, <gasp>, <giggle>, <groan>, <growl>, <grunt>, <grr>, <hiss>, <laugh>, <moan>, <pant>, <pff>, <phew>, <scream>, <shout>, <shriek>, <sigh>, <sneeze>, <snicker>, <snort>, <sob>, <throat-clearing>, <tsk>, <whimper>, <whispers>, <yawn>, <short pause>, <long pause>
Anything outside this list is treated as text. Tags describe a sound, not a mood, so there is no <angry> or <excited>: carry sustained emotion through the writing and the voice choice, and use tags for the breaths, laughs and gasps in between. A horror narration with a breath, a whisper and a gasp:
The house on Merrow Lane had been empty for eleven years, and yet every night at 3:17 the kitchen light came on. <breath> Clara told herself it was faulty wiring. She told herself that for six weeks. Then, on the first night of November, she crept down the stairs in her socks and pushed the kitchen door open. <whispers> The table was set for two. Two plates, two candles, two glasses of red wine, one of them half empty. <gasp> And on the chair facing her, a folded napkin with her name embroidered in thread the color of old blood.Gemini 3.8 Flash, voice Enceladus, en-US: horror audiobook · Open on Scenario
A villain monologue, where <cackle> and <growl> do the character work:
You thought the prophecy was about you? <cackle> Oh, little knight, how precious. I wrote the prophecy. I carved it into the temple walls myself, a thousand years ago, and I have waited every single one of those years for a fool brave enough to believe it. <growl> You climbed my mountain. You slew my guardians. You carried the Crown of Ash all the way to my throne, exactly as I planned. <long pause> Now kneel. <sigh> Or don't. It ends the same either way.Gemini 3.8 Flash, voice Fenrir, en-US: game villain · Open on Scenario
Tags work in every language. Set the Language parameter to match the script and keep the tag names in English. A Brazilian Portuguese football call with <shout> and <cheer>:
Bola rolando no Maracanã, quarenta e três do segundo tempo, e o placar ainda marca zero a zero. <breath> Lucas Moreira recebe pela direita, passa pelo primeiro, passa pelo segundo, olha a arrancada! <shout> Ele entra na área, corta pra perna esquerda, bate colocado... <short pause> GOOOOL! <cheer> É gol do Tricolor! A torcida vem abaixo, o estádio inteiro treme, e o camisa dez corre pro abraço com lágrimas nos olhos. Que momento, meus amigos, que momento!Gemini 3.8 Flash, voice Puck, pt-BR: football commentary · Open on Scenario
Two-speaker dialogue
To voice a scene, write each turn on its own line as Name: text, then open Multi-speaker and add one row per character, with the Speaker name matching the label in the script exactly and a voice for each. Up to two speakers are supported per job. For short listener reactions in the middle of a scene, wrap the word in pipes, like |hmm| or |no way|.
In our tests the two voices came out clearly distinct: speaker diarization on every dialogue example detected exactly two speakers with every turn attributed to the right character. A heist scene with Wren mapped to Leda and Brannock mapped to Algenib:
Wren: <whispers> Brannock, the forge is still warm. Someone was here less than an hour ago.
Brannock: |hmm| Dwarves don't leave a forge burning. Not unless they left in a hurry.
Wren: Then where's the crown? The pedestal is empty. <gasp> Wait, look at the floor, there are boot prints in the ash, heading toward the lower vault.
Brannock: <laugh> Heading toward the vault that floods at midnight. Whoever took it is either brave or very, very stupid.
Wren: Or both. Grab the lantern. We're going down.Gemini 3.8 Flash, two speakers (Leda, Algenib), en-US: fantasy game dialogue · Open on Scenario
A two-host podcast with Maya on Despina and Theo on Sadachbia:
Maya: Welcome back to Deep Cuts, the podcast about the weirdest corners of history. I'm Maya.
Theo: And I'm Theo, and today Maya is going to try to convince me that a war was once fought over a bucket.
Maya: Not try. It happened. In 1325, soldiers from Modena snuck into Bologna and stole a wooden bucket from the town well.
Theo: |no way| A bucket.
Maya: A bucket! <laugh> And Bologna declared war to get it back. Thousands of men fought. Bologna lost.
Theo: <chuckle> So where's the bucket now?
Maya: Still in Modena. They never gave it back.Gemini 3.8 Flash, two speakers (Despina, Sadachbia), en-US: history podcast · Open on Scenario
Flash-Lite handles the same format. A support call with the Agent on Autonoe and the Customer on Rasalgethi:
Agent: Thanks for calling Pixel Pantry support, this is Jess. How can I help you today?
Customer: Hi, yeah, <sigh> my order arrived this morning but the box was completely crushed and two of the jars were broken.
Agent: Oh no, I'm really sorry about that. Can I get your order number?
Customer: Sure, it's P P four eight two nine one.
Agent: Got it. I can see the order here. I've sent a replacement for both jars at no charge, and it'll ship out tomorrow with express delivery.
Customer: |oh great| That was quick, thank you.
Agent: My pleasure. Anything else I can do for you?Gemini 3.8 Flash-Lite, two speakers (Autonoe, Rasalgethi), en-US: customer support call · Open on Scenario
Flash-Lite for everyday reads
Flash-Lite shines on the reads that need to be clear, natural and plentiful rather than theatrical. It understands the same tags; they simply land a little lighter. A voice-agent style callback:
Hi, this is Nova from Brightline Travel. <short pause> I'm calling about your booking to Lisbon on the fourteenth of October. Good news first: your flight is confirmed, seat 12A, window, just like you asked. <breath> The not-so-good news is that your hotel moved your check-in to three in the afternoon, so if you land early, the front desk can hold your bags. Reply YES to keep everything as is, or call us back any time before Friday. Have a lovely trip!Gemini 3.8 Flash-Lite, voice Kore, en-US: travel booking voice agent · Open on Scenario
A museum audio guide:
Welcome to Gallery Seven: The Deep Ocean. <short pause> The creature in the glass case in front of you is a giant squid, and at just over nine meters long, it's one of the largest ever recovered intact. Look closely at its eyes. Each one is the size of a dinner plate, the biggest eyes of any animal on Earth, built to catch the faintest flicker of light two thousand meters below the surface. <breath> For centuries, sailors told stories of sea monsters dragging ships under. <chuckle> It turns out they weren't entirely making it up. When you're ready, continue to your left for the bioluminescence room.Gemini 3.8 Flash-Lite, voice Sulafat, en-US: museum audio guide · Open on Scenario
A Brazilian Portuguese news bulletin:
Boa tarde. Estas são as principais notícias desta quarta-feira. <short pause> Em São Paulo, a prefeitura inaugurou hoje a nova linha de ônibus elétricos que liga a Zona Leste ao centro da cidade em apenas quarenta minutos. <breath> No litoral de Santa Catarina, pesquisadores registraram a volta das baleias-francas, com mais de cem avistamentos só neste mês. E no esporte, a seleção feminina de vôlei garantiu vaga na final do campeonato sul-americano. <short pause> Voltamos às seis com mais informações. Até lá.Gemini 3.8 Flash-Lite, voice Schedar, pt-BR: news bulletin · Open on Scenario
And it still has enough character for lightweight game barks, like this shopkeeper greeting:
Ah, a customer! <chuckle> Come in, come in, mind the dragon egg, it's only mostly asleep. Welcome to Grimbolt's Curiosities, finest wares this side of the Sunken Road. <short pause> Looking for a sword? Got three. One of them only curses you on Tuesdays. <laugh> Potions? Top shelf. Red heals, blue burns, green, well, nobody's been brave enough to find out. <breath> Tell you what, friend. You look like the adventuring type. First purchase comes with a free map. Probably accurate.Gemini 3.8 Flash-Lite, voice Puck, en-US: fantasy shopkeeper NPC · Open on Scenario
Pushing the Limits
We also ran scripts built to stress the model: long chains of non-verbal sounds, emotional arcs inside a single read, and languages mixed mid-sentence. These held up well, and they show how far you can push a performance with nothing but the script.
A performance made mostly of sounds. A radio host with a cold, stacking <yawn>, <throat-clearing>, <cough>, <sneeze> and <groan> between lines. Every tag is performed, none is read aloud, and the host stays in character throughout:
<yawn> Good morning, early birds, you're listening to Sunrise FM, it's six o'clock, and <throat-clearing> yes, I sound like a foghorn today. <cough> Sorry. My kid brought home a cold from daycare and generously shared it with the whole family. <sneeze> Oh, excuse me. <groan> Okay. Traffic on the ring road is backed up past exit nine, so leave early or bring a podcast. <sneeze> And, uh, weather: rain until noon. <sigh> Which is perfect, because I'm going straight back to bed after this show. Here's your first song.Gemini 3.8 Flash, voice Sadachbia, en-US: radio host with a cold · Open on Scenario
An emotional arc in one take. A wedding toast that starts with nervous laughter, breaks into <sob> and <cry>, and recovers into a final laugh. The voice carries the tears into the words around the tags, not just the tags themselves:
I promised myself I wouldn't do this. <laugh> Okay. Hi everyone. For those who don't know me, I'm Lily's big sister, which means I've spent twenty-nine years telling her she was wrong about everything. <giggle> She wasn't wrong about Daniel. <breath> When our dad got sick last spring, Daniel drove four hours every single weekend just to sit with him and watch terrible westerns. <short pause> Dad called him the only man who laughed at the right parts. <sob> Sorry. <sigh> He would have loved today. <cry> So, please, raise your glasses. To Lily and Daniel. <laugh> And to Dad, who is definitely complaining about the music.Gemini 3.8 Flash, voice Kore, en-US: wedding toast, laughter to tears · Open on Scenario
Code-switching mid-sentence. English with Spanish phrases dropped in, Language left on en-US. The model switches pronunciation for each Spanish phrase and comes straight back to English without a seam:
Okay, so my abuela calls me at seven in the morning, right? And she goes, mija, ¿ya comiste? And I'm like, abuela, it's seven, I'm barely awake. <laugh> And she says, no importa, I'm sending your tío with tamales. <sigh> Twenty minutes later, tío Ramón is at my door with a cooler, like, a full cooler, and he says, tu abuela dice que estás muy flaca. <giggle> I live alone! Who is going to eat forty tamales? <short pause> Spoiler alert: me. I ate them. Todos. No regrets.Gemini 3.8 Flash, voice Callirrhoe, en-US with Spanish: abuela and the tamales · Open on Scenario
Flash-Lite under pressure. The lighter tier handles dense non-verbal work too: a scared child in a ghost train, moving through <whimper>, <gasp>, <shriek> and <pant> before relief lands on a <giggle>:
<whimper> Is someone there? <breath> Please, I got separated from my class, the lights went out in the ghost train and I can't find the exit. <gasp> What was that? <whimper> Something touched my arm. <short pause> I'm, I'm not scared. I'm not. <heavy breath> Okay, I'm a little scared. <shriek> It moved! The skeleton moved! <pant> Wait. <long pause> Oh. <giggle> It's just a mechanical one. It waves at everybody. <sigh> Can you walk with me to the exit? Please?Gemini 3.8 Flash-Lite, voice Pulcherrima, en-US: scared kid in a ghost train · Open on Scenario
Parameters
Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS expose the same four inputs. Only Text is required.
text
Required, 1 to 5,000 characters. The exact words to speak, plus inline vocal tags in angle brackets. It is read verbatim, so never put directions here: see the lighthouse example above for tags and pauses done right. In multi-speaker mode, every turn starts with Name: on its own line, and lines without a prefix continue the previous speaker's turn.
voice
Optional, default Puck. One of 30 preset voices, used for single-speaker jobs and ignored when Multi-speaker is set. The voice carries most of the character, so audition two or three before rewriting a script: compare the gravel of Fenrir in the villain monologue with the bright energy of Puck in the football call.
multiSpeakerConfig
Optional, up to 2 rows. Each row maps a Speaker name (the label before the colon in your script, 1 to 64 characters, must match exactly) to a Voice. Leave it empty for single-speaker reads. All four dialogue examples above use it.
language
Optional, default en-US. The locale for pronunciation and prosody. The model auto-detects when the script clearly differs from the setting, but set it explicitly for any non-English script, as in the pt-BR football and news examples. See Supported Languages below for the full list of 24 locales.
Supported Voices
All 30 voices are shared across Gemini 3.8 Flash, 3.8 Flash-Lite and 3.1 Flash, so a character keeps the same voice when you switch tiers. The Scenario interface labels each one:
Female: Achernar, Aoede, Autonoe, Callirrhoe, Despina, Erinome, Gacrux, Kore, Laomedeia, Leda, Pulcherrima, Sulafat, Vindemiatrix, Zephyr.
Male: Achird, Algenib, Algieba, Alnilam, Charon, Enceladus, Fenrir, Iapetus, Orus, Puck (default), Rasalgethi, Sadachbia, Sadaltager, Schedar, Umbriel, Zubenelgenubi.
Every voice was used at least once in our 30-clip test set. Starting points that worked well: Charon and Orus for narration and documentary, Fenrir and Algieba for villains and trailers, Enceladus for suspense, Puck and Umbriel for comedic and energetic characters, Kore, Sulafat and Autonoe for clear, friendly assistant and guide reads.
Supported Languages
24 locales are available in the Language parameter on all current Gemini TTS models:
Arabic (Egypt) ar-EG, Bengali (Bangladesh) bn-BD, Dutch nl-NL, English (India) en-IN, English (US) en-US, French fr-FR, German de-DE, Hindi hi-IN, Indonesian id-ID, Italian it-IT, Japanese ja-JP, Korean ko-KR, Marathi mr-IN, Polish pl-PL, Portuguese (Brazil) pt-BR, Romanian ro-RO, Russian ru-RU, Spanish (US) es-US, Tamil ta-IN, Telugu te-IN, Thai th-TH, Turkish tr-TR, Ukrainian uk-UA, Vietnamese vi-VN.
In testing, Flash and Flash-Lite produced correct, verbatim speech in all 14 locales we covered: en-US, en-IN, pt-BR, es-US, fr-FR, de-DE, it-IT, nl-NL, ru-RU, ar-EG, hi-IN, ja-JP, ko-KR and id-ID.
Use Cases
Game dialogue: villains, NPC barks, companions and two-character cutscenes with 3.8 Flash, using tags for the laughs, growls and gasps that make a line feel acted.
Audiobooks and narration: long-form chapters in up to 5,000 characters per job, with breaths and pauses placed exactly where the story needs them.
Trailers and advertising: dramatic trailer reads, luxury brand spots and localized ad variants across 24 locales.
Podcasts and explainers: two-host scripted episodes in a single job, with pipe backchannels for natural reactions.
Voice agents and IVR: callbacks, phone menus and assistant replies with 3.8 Flash-Lite, where clarity and turnaround matter more than theatre.
E-learning, guides and announcements: course narration, museum audio guides, station announcements and navigation prompts in the learner's own language.
Tips for Better Results
Only type what should be heard. Parentheses, prose directions and 3.1 square-bracket tags do not steer 3.8: the first two are spoken, the last is ignored.
Put vocal tags between sentences. Tags sitting at a sentence boundary, as in every example here, landed cleanly and never broke the flow of the line.
Use pause tags for timing.
<long pause>reliably produced about two seconds of silence and<short pause>a shorter beat, which is ideal for trailers and dramatic reveals.Stick to the supported tag list. A tag outside the list, like
<sniff>, is not just ignored: in our tests it made the tags after it get read aloud. Retaking the script without it fixed the read.Pick the voice before you polish the words. The same script changes character completely between voices, so audition two or three first.
Always set Language for non-English scripts. Every non-English example here had its locale set explicitly and came back with correct pronunciation.
Match speaker names exactly. The Speaker name in Multi-speaker must be identical to the label before the colon in the script, including capitalization.
Spell unusual names the way they sound. Invented names can drift: our "Brannock" was heard as "Brennok" by a transcriber. Spell tricky names phonetically if the exact sound matters.
Known Limitations
No separate style or direction field. Google's API accepts a per-turn delivery style alongside the transcript, but Scenario exposes only the Text field. Sustained tone comes from the voice and the writing; momentary sounds come from tags.
A tag can occasionally be spoken as words. In one trailer script, a
<short pause>was read aloud as "short pause" on two takes in a row, while the same tag worked fine everywhere else. If it happens, swap the tag (a<long pause>fixed it) or use punctuation such as a period or ellipsis instead.Two speakers maximum per job. For scenes with three or more characters, generate the scene in parts and join the files in an editor or in a Scenario workflow.
30 preset voices only. Google's voice design and voice replication features are not available on Scenario. For cloning, use a dedicated voice cloning model.
24 locales on Scenario. Google lists far more languages for 3.8 (130 for Flash, 101 for Flash-Lite), but the Scenario Language parameter currently offers 24.
No speed, pitch or SSML controls. Pacing comes from punctuation and pause tags; standard SSML is not recognized.
24 kHz mono output. Ideal for voiceover and dialogue; master it for music-heavy or stereo mixes.
Gemini 3.1 Flash TTS (previous generation)
Gemini 3.1 Flash TTS is still available for workflows built around it. It uses the same 30 voices, 24 locales and Multi-speaker setup, but a different tag system.
What Gemini 3.1 Flash TTS Does
Gemini 3.1 Flash TTS converts a text script into expressive speech. You provide a script, choose a voice, and the model returns a mono MP3 at 24kHz. What sets it apart is the audio tag system: by embedding tags like [determination], [whispers], or [enthusiasm] directly in the text, you control how specific lines are delivered without extra parameters or prompts. The direction lives in the script itself.
Audio Tags in 3.1
Audio tags appear in square brackets within the script and control delivery without altering the words themselves. Place them at the start of a sentence for consistent application:
[determination] We are not leaving without the artifact.
[whispers] The vault door was already open.
[enthusiasm] This is going to change everything.
[slow] Read each step carefully before you begin.
[laughs] I cannot believe that actually worked.
[soft] The people using the tools. That always brings it back to earth.Mid-sentence placement is supported but may produce uneven transitions in complex phrasing. The model supports over 200 audio tags covering emotions, interjections, pacing, and performance notes.
Multi-Speaker Dialogue in 3.1
Label each line with a speaker name and configure Multi-Speaker Config to map each name to a voice:
Scout: Someone followed us from the market. Do not look back.
Commander: [determination] How many?
Scout: Two. Maybe three. We take the alley on the left.
Commander: [low] I will handle the one on the right. You get clear.
Scout: Together or not at all.In this example, Multi-Speaker Config would be set to: Scout mapped to Kore, Commander mapped to Fenrir.
Known limitation in 3.1: voice differentiation between the two speakers is not reliably applied, and both may render with the same underlying timbre. If you need two clearly distinct voices in one job, use Gemini 3.8 Flash or Flash-Lite, or generate each speaker separately in 3.1 and combine the files.
Migrating to Gemini 3.8
Voices, locales, the Multi-speaker setup and the 5,000-character limit carry over unchanged from 3.1 and 2.5, so the migration is all about the script:
Replace square-bracket tags with the angle-bracket vocal tags from the 3.8 list, for example [laughs] becomes
<laugh>and [whispers] becomes<whispers>.Remove mood and pacing tags that have no 3.8 equivalent, such as [determination], [enthusiasm] or [slow], and carry that tone through the voice and the writing instead.
Remove any stage directions from the text, since 3.8 reads everything aloud.
Pick Flash for performance work and Flash-Lite for high-volume utility reads.
Gemini 3.1 Flash TTS itself uses the same voice names and Multi-Speaker Config format as Gemini 2.5 TTS, so 2.5 workflows can move to either generation with only the script changes above.