4Geeks chosen to deliver AI education in the Bahamas alongside Harvard, Oxford, and Columbia.See more
7 min read

Gemini Text to Speech: Custom Voices Without a Studio

Gemini text to speech can now design custom voices from a description. Here's how to make your first one, direct the delivery, and what you can build.

Gemini text to speech is Google's AI voice tool that turns a written script into spoken audio. Since September 23, 2026, it can also design a new voice from a plain English description. Unlike older text to speech tools with one flat preset voice, the new Gemini 3.8 Flash TTS model takes direction like an actor. Whisper this line, pause here, sound excited there.

You can try it in your browser in Google AI Studio, and Google is also bringing it to everyday users in Google Vids. That makes it a realistic option for professionals and small business owners. Think voiceovers, a simple podcast, or a friendly voice for customers, all without a recording studio.

What Gemini text to speech does now

Text to speech (TTS) means software reads text out loud. The Gemini version goes further in three ways, according to Google's launch post:

  • You can design a voice. Describe the role, accent, and personality you want, and the model builds a voice to match.
  • You can direct each line. Tell it which lines to whisper, rush, or say with a sigh.
  • You can stage a conversation. One script can hold two speakers who take turns naturally, which is handy for podcasts.

Google released two models, and they do the same basic job:

ModelBuilt forPick it when
Gemini 3.8 Flash TTSCreative work and character voicesYou care most about expressive, carefully directed audio
Gemini 3.8 Flash-Lite TTSHigh volume at lower costYou need lots of everyday audio, dubbing, or a voice agent

If you're just starting, Flash TTS is the easier one to learn on, because it handles detailed direction best. You'll know which one fits once you've made a few clips.

Where you can use it

You don't need to be a developer to hear what it can do. Google lists these places, as of late September 2026:

  • Google AI Studio. A browser workspace where you can design voices and write a two speaker script line by line.
  • Google Vids. Google's video tool, rolling out to everyone, for adding voiceovers to videos.
  • Gemini API. The developer route, for putting Gemini voices inside your own app or website.
  • Gemini Enterprise. API access for companies, which Google says is coming soon.

So a marketer can make a voiceover in Vids, while a developer on the same team connects the API to the company's app. Start with the tool that matches how you already work.

How to make your first custom voice

Here's a simple first project: a 30 second intro for a podcast or a training video. Menu names in AI Studio may change, so follow the idea rather than exact buttons.

  1. Open Google AI Studio and sign in with your Google account. Look for the speech or audio area.
  2. Describe your voice in one or two sentences. For example: "A warm, confident narrator in her 40s with a light Mexican Spanish accent, speaking English." Google says voice design works across more than 100 languages and dialects.
  3. Paste your script. Write it the way people really talk. Short sentences sound more natural than long ones.
  4. Generate and listen. Most scripts sound fine on the first try with no extra direction at all.
  5. Add direction only where a line feels off. Keep it short, like "calm and reassuring."
  6. Save the voice so every future episode sounds like the same person. Then download the audio.

If you'd rather pick than design, there's a library too. Google says it holds more than 2,000 ready voices, including regional ones like Mexican Spanish and Quebec French.

How to direct the delivery

This part feels most like working with a real voice actor. Google's developer guide separates two kinds of direction:

  • Direction for a whole line. Mood, speed, and tone that last the full turn, such as "whispering" or "speaking slowly."
  • Moments inside a line. Short sounds placed exactly where they happen. You type the sound's name, like sigh, laugh, or short pause, inside angle brackets at that spot in the script.

Capital letters add stress to a word, as in "It was a VERY long day." Commas and ellipses slow things down.

One tip from the guide is worth stealing. Don't pile long instructions on every line, because extra text can make the voice drift. Build the character once with voice design, then keep your per line directions short.

Curious where this fits in your job? Voice tools are one piece of a bigger shift: people using AI assistants to take repetitive work off their plate. If you want to explore that with live classes and a mentor, and no coding, see how AI Fluency works.

What you can build with it

These are the jobs Google built the models for. Pick one that matches a task you already do.

Podcasts and audio versions of your content. Turn a blog post or newsletter into a two host conversation. Google says the model keeps voices steady across hours of audio, which matters for long episodes.

Dubbing and translated videos. Take one training video and voice it in several languages. Flash-Lite TTS is the model Google points to for high volume dubbing.

Voice agents. A voice agent is an AI assistant that talks back, like a bot that answers your business line. Here's the catch: text to speech is only the voice. Another AI model writes each reply, and Gemini speaks it. Building a full phone agent still takes code or a developer platform. Google names Agora, LiveKit, Pipecat, and Vercel as partners.

Explainer and training videos. Add consistent narration in Google Vids without booking voice talent for every small update.

Voice replication means copying a real person's voice. Google says it needs just a 30 second sample, and it comes with rules you should know before you try.

  • Consent is required. The voice owner has to record a spoken consent that matches the sample before the voice is created.
  • Every clip is watermarked. Google adds SynthID, an invisible watermark, to all audio from its Gemini audio models. That keeps AI speech detectable.
  • It isn't available everywhere. Voice replication in AI Studio isn't offered in Illinois, Texas, the European Economic Area (EEA), the UK, Switzerland, or India.

The practical rule is simple. Only clone your own voice or one you have clear rights to use.

What it's not built for

Gemini text to speech reads your exact words. It isn't the right tool for a free flowing live conversation. For that, Google points developers to a separate product, the Live API, which handles back and forth audio in real time.

A few more limits from Google's developer guide:

  • A single request can stage up to two speakers using the ready made voices.
  • Sound effects like applause don't work well. Stick to human sounds like laughs, sighs, and breaths.
  • It takes text in and gives audio out. It won't edit a recording you already have.

If your goal is narration, podcasts, dubbing, or the voice of an assistant, it fits. If you need to edit existing audio, look elsewhere.

Go deeper on AI at work

  • ChatGPT prompts for work: the same "give clear direction" skill, applied to writing.
  • AI Fluency: a 4 week program for professionals who want AI working inside their real job, no coding required.

Want AI working inside your actual job?

Explore AI Fluency, a 4 week program with live classes and 1:1 mentorship for professionals. No coding required. Get the details and compare it with other 4Geeks programs.

Frequently Asked Questions