Gemini 3.8 Flash TTS lets creators do more than turn a script into speech: they can describe a character’s voice, audition different takes and direct how lines are performed. Google AI Studio brings those controls together in a playground for building voices and multi-speaker dialogue. Its companion model, Gemini 3.8 Flash-Lite TTS, also appears near the top of a public voice-quality ranking.

Where the models rank

On Hume AI’s public leaderboards, Gemini 3.8 Flash TTS ranked first in Voice Design, the task of creating a voice from a natural-language description, with a score of 71.4. Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS held the top two positions for overall voice quality. Those rankings offer a starting point; a creator still needs to audition a voice against the words and mood of a particular project.

Colored voice waveforms form vertical columns, with one rising highest against a dark gridded background.

▲ Voice design leaderboard ranking

Demonstration samples produced with the model span a wide range of performances: an intimate ASMR-style whisper, a forceful theatrical delivery, deep movie-trailer narration and speech in Spanish and Hindi. A voice replica was also among the generated examples. These performances illustrate why voice design involves more than choosing a pleasant sound. Accent, pacing, vocal texture and emotional direction all shape whether a line works.

Build and audition a custom voice

Google AI Studio offers two paths. Voice Design creates a voice from a written description. Voice Replication works from recorded speaker audio, although its availability depends on the user’s region. For a new character, Voice Design provides a way to specify the performance before writing a full dialogue.

  1. Open the Google AI Studio playground. In the model picker, look under audio models and select Gemini 3.8 Flash TTS.
  2. Choose Voice Design. Describe the speaker’s age, accent, timbre, pacing and emotional demeanor. One example brief calls for a sharp, theatrical British architecture critic in her 50s with dry wit.
  3. Use Improve My Prompt if the description needs more detail. It expands a short character brief into directions covering such qualities as accent, rhythm and vocal texture.
  4. Generate the voice. The design run produces three audition takes; listen to each for tone and pronunciation rather than treating them as interchangeable.
  5. Assign the preferred voice to a speech block, then write lines for a multi-speaker exchange.

Once the speakers are set, the script can carry acting directions too. Google AI Studio includes styles such as Whisper, Friendly, Narration, Promote and Podcast. Bracketed cues including [chuckle], [laugh], [cough] and [cheer] can go directly into dialogue lines. They provide a way to request vocal expressions without SSML, a markup format used to direct synthesized speech. Generated dialogue can include emotional inflection, breathing and laughter alongside the spoken words.

Hands type at a laptop as colorful sound waves arc from the screen toward a nearby microphone.

▲ Text-directed voice performance

The playground also exposes temperature, a setting that adjusts variation in generation, and a Filter Words toggle for conversational fillers and pauses. These controls are worth testing after the basic voice and script work: a more conversational delivery may suit one line, while a cleaner narration style may suit another.

Take the voice beyond the playground

The Get Code tab displays Python code using the google-genai library to generate audio programmatically. That gives developers a path from an auditioned dialogue to an application workflow. Google also offers more than 2,000 ready-made voices across more than 100 languages, so designing a voice from scratch is not the only option.

Potential uses include employee training videos, product walkthroughs and multilingual content localization. A practical first test is a short script with two contrasting speakers: define one voice precisely, compare its three takes, then add a style and one expression cue to see how the delivery changes. The main advantage of this workflow is control over performance at both the voice-design and dialogue stages; the next step is to hear whether that control serves the content you want to make.