Skip to main content
Back to insights

September 25, 2026

Gemini 3.8 Flash TTS: Price, Rankings and the Fine Print

Gemini 3.8 Flash TTS and Flash-Lite TTS bring prompt-based voice design at low prices. Here's what they cost, how they rank independently, and the catches.

By Tran Tien Van9 min read

Article focus

Google's Gemini 3.8 Flash TTS and Flash-Lite TTS let you design voices from a text prompt and direct performances line by line. Independent tests put Flash near the top for quality at a low price. But prices double in January, cloned voices rank mid-pack, and one headline benchmark needs context.

Gemini 3.8 Flash TTS is Google's new text-to-speech model, released on September 23, 2026, alongside a cheaper Flash-Lite TTS. It lets you design voices from a text prompt and direct delivery line by line, for about $0.81 per hour of audio. Independent tests rank it second for quality. The catches: prices double in January, cloned voices rank mid-pack, and Google's headline benchmark comes from a company it has close ties to.

Key Takeaways

  • Google released Gemini 3.8 Flash TTS and Flash-Lite TTS on September 23, 2026, in the Gemini API and Google AI Studio.
  • Flash TTS costs about $0.81 per hour of audio and Flash-Lite about $0.54, until prices double on January 1, 2027.
  • On Artificial Analysis's independent Speech Arena, Flash TTS ranks second and Flash-Lite sixth, at a third of the price of several rivals.
  • For cloned voices, they rank mid-pack, behind Inworld, Cartesia and ElevenLabs.
  • Google's claim of first place comes from Hume AI, which licenses technology to Google. Its founder now works at Google DeepMind.

What Did Google Release With Gemini 3.8 Flash TTS?

Two text-to-speech models with far more control over voices. Reported fact: Google announced both models on September 23, 2026.

  • Gemini 3.8 Flash TTS is built for creative work: designing new voices, directing performances and producing long-form audio such as audiobooks and podcasts.
  • Gemini 3.8 Flash-Lite TTS is built for scale: high-volume dubbing, audio content and voice agents, with control over tone and pacing.

The main new features, according to Google:

  • Voice design from a prompt. Describe a role, accent and voice characteristics in plain language, and the model creates a new voice.
  • Voice replication. Recreate a voice from about 30 seconds of audio of your own voice, or one you have the rights to use.
  • Line-by-line direction. Write stage directions, or add tags such as laughs, sighs and gasps, plus short listening cues like "mhm."
  • Two-speaker scenes. Stage a conversation between two voices from a single script.
  • Long-form stability. Google says voice quality holds across hours of audio with minimal drift.
  • Voice remixing. Adjust an existing voice's pitch, pace or accent by prompt. Google says this is coming soon.

Both are available as gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts in the Gemini API and Google AI Studio. Flash TTS also powers Gemini Notebook, and Flash-Lite TTS powers Google Vids. Enterprise API access through Gemini Enterprise is coming soon.

One detail to check: Google's blog and its developer docs describe the models a little differently. The blog cites more than 100 languages and 2,000 voices. The docs list more than 130 languages for Flash and more than 100 for Flash-Lite, with 30 named voices plus hundreds more in an extended library.

How Much Does Gemini 3.8 Flash TTS Cost?

Very little per hour today, and twice as much from January. Here's the pricing from Google's Gemini API pricing page, per million tokens:

ModelInput textOutput audioAudio per hour
Gemini 3.8 Flash TTS, through Dec 31, 2026$0.50$9.00About $0.81
Gemini 3.8 Flash TTS, from Jan 1, 2027$1.00$18.00About $1.62
Gemini 3.8 Flash-Lite TTS, through Dec 31, 2026$0.50$6.00About $0.54
Gemini 3.8 Flash-Lite TTS, from Jan 1, 2027$1.00$12.00About $1.08
Gemini 3.1 Flash TTS (previous model)$1.00$20.00About $1.80

Audio is billed at 25 tokens per second, so an hour of speech is 90,000 output tokens. The input text barely matters: an hour of speech is roughly 9,000 words, which costs a fraction of a cent. Batch requests cost half the standard price.

Even after January's increase, Flash TTS will cost less per hour than the previous 3.1 Flash TTS. But plan for the jump now. A voice agent that speaks 100,000 minutes a month on Flash-Lite costs about $900 today and about $1,800 from January. A 10-hour audiobook on Flash TTS costs about $8.10 today, or about $4.05 with batch requests, and about $16.20 at the new rate.

How Does Gemini 3.8 Flash TTS Rank on Independent Tests?

Near the top, and cheaper than most of the models around it. Artificial Analysis's Speech Arena ranks models by blind human votes between two voice samples. Here's where things stood when we checked:

RankModelEloPrice per million characters
1Cartesia Sonic 3.61279$49.00
2Gemini 3.8 Flash TTS1265$16.50
3Alibaba Qwen-Audio-3.0-TTS-Plus1259$27.60
4Inworld Realtime TTS-21246$20.80
5Speechify Simba 3.21239$6.60
6Gemini 3.8 Flash-Lite TTS1239$11.00
12ElevenLabs v3 Conversational1196$50.00
32OpenAI TTS-1 HD1101—

So Flash TTS is a close second, 14 points behind the leader, at about a third of Cartesia's and ElevenLabs' price per character. Flash-Lite ties Speechify Simba 3.2 on score, though Simba costs less. And both new models are well ahead of Google's previous Gemini 3.1 Flash TTS, which ranks tenth at 1202.

These prices reflect Google's rates before January's increase. After it, Flash TTS will cost roughly twice as much per character, still below several rivals near the top.

How Good Is Gemini 3.8 Flash TTS at Cloning Voices?

Less standout. Artificial Analysis also runs a cloned-voice leaderboard, where every model uses the same eight cloned voices, four American and four British.

  • Inworld Realtime TTS-2 leads at 1142, followed by Cartesia Sonic 3.6 at 1137.
  • ElevenLabs Eleven v3 ranks fifth at 1073.
  • Gemini 3.8 Flash-Lite TTS ranks eighth at 1055.
  • Gemini 3.8 Flash TTS ranks eleventh at 1046.

So Google's strengths depend on the job. It ranks near the top on its own voices and leads on prompt-based voice design by Hume AI's measure. But if your product depends on matching a specific person's voice closely, test the rivals too. Whichever you choose, record the sample in a quiet room with a good microphone, since background noise and echo tend to carry into the cloned voice.

Why Does the Hume AI Benchmark Need Context?

Because of the relationship between Google and Hume AI. Google says Flash TTS took first place on Hume AI's Voice Design Benchmark, with a score of 71.4, and led on accent modeling at 60.8. It also says Flash and Flash-Lite took first and second on Hume AI's Overall Quality Index.

Here's what readers should know:

  • The deal. In January 2026, Google DeepMind signed a licensing deal with Hume AI and hired its founder, Alan Cowen, and several engineers, as Wired reported.
  • The author. Alan Cowen is a named co-author of Google's TTS announcement.
  • The company. Hume AI still operates independently under a new chief executive.
  • The benchmark. Hume says it's built on more than a million human ratings.

None of this means the benchmark is wrong. It means you should read it alongside independent results, which here mostly agree that Google's models are among the best. Google's claims about top positions on Voice Arena in languages such as Vietnamese and Japanese were harder to check. The Voice Arena page we found showed an English leaderboard for conversational agents only.

What Are the Limits of Gemini 3.8 Flash TTS?

A few matter for real products. Most come from Google's own developer documentation:

  • Two-speaker scenes use built-in voices. The docs say native two-speaker requests work with up to two prebuilt voices. Custom voices need to be generated one speaker at a time.
  • Text in, audio out. The models take text only and read it as a verbatim script.
  • Voice storage limits. Saved custom voices are capped at 200 per project and kept for one year. Temporary voice keys expire after seven days.
  • Some features aren't here yet. Voice remixing and Gemini Enterprise API access are coming soon.
  • Prices double in January, so budgets built on today's rates will be off by half.

For live, two-way conversations, Google also offers separate Gemini 3.8 Live models, which handle speech in and out in real time.

How Do the Voice Cloning Safeguards Work?

Google has built in three main protections, according to its announcement:

  • Consent checks. To replicate a voice, you must upload a spoken consent recording from the voice owner, and it has to match the reference sample.
  • SynthID watermarks. Every audio clip from Gemini's audio models carries an inaudible watermark, so AI-generated speech can be detected.
  • Content credentials. Replicated voices also carry C2PA credentials, a standard for tracing where media came from.

These are sensible steps. But they don't replace your own policy. Before cloning any voice, get written permission, record how it will be used, and check the laws where your audience lives.

What Audio Formats and Streaming Options Are There?

Enough for most apps, including phone lines. According to Google's developer docs:

  • Default file output: WAV at 24 kHz, mono, 16-bit.
  • Streaming output: turn streaming on, and raw 16-bit audio at 24 kHz arrives while it's still being generated.
  • Phone formats: mu-law and A-law, the formats most phone systems expect.
  • Inline tags: breaths, coughs, sighs, laughs, gasps, and short or long pauses.

Streaming matters most for voice agents. Listeners notice even a one-second pause, so start playback on the first chunk instead of waiting for the whole reply. For phone agents, asking for mu-law or A-law directly saves a conversion step and a little delay.

Google says developer platforms including Agora, LiveKit, Pipecat and Vercel already support the new models through the Gemini API. If you build on one of them, switching voices may be a small change rather than a new integration.

Which Model Should You Use: Flash TTS or Flash-Lite TTS?

Match the model to the job. Google positions Flash TTS as its full voice studio, and Flash-Lite TTS as the cheaper option built for volume:

  • Audiobooks, podcasts and games: Flash TTS, for voice design, acting direction and long-form stability.
  • High-volume dubbing and narration: Flash-Lite TTS, for lower cost with strong quality.
  • Voice agents that read out responses: Flash-Lite TTS, with streaming output.
  • Real-time, two-way conversations: Gemini 3.8 Live instead, since it handles listening and speaking together.
  • Close matches to a specific person's voice: test Inworld, Cartesia and ElevenLabs alongside Google.

We covered how to run voice agents in production in our guide to serverless voice agents on AWS. The same design questions apply here: latency, cost per minute and fallbacks.

How Should You Test Gemini 3.8 Flash TTS?

With your own scripts, languages and listeners. Leaderboards use short, generic samples.

  • Use real scripts, such as your product's actual prompts, chapters or support replies.
  • Test each language you serve, including Vietnamese or any regional accent your users expect.
  • Run blind listening tests, where listeners pick between two unlabeled samples.
  • Measure time to first audio if you're building a voice agent, since delays feel longer in speech.
  • Price it at 2027 rates, so your budget holds after the increase.

Our guide to AI agent evaluation covers how to build test sets and blind comparisons that hold up.

How Van Data Team Helps Teams Build Voice Experiences

We help teams choose and ship voice models with evidence, not launch claims. That means blind listening tests in the languages you serve, latency and cost measured per minute, and fallbacks when a provider changes prices or slows down.

Gemini 3.8 Flash TTS makes expressive voices cheap, for now. If you want to build voice features that stay affordable after January, our work on AI agent evaluation and AI agent development cost is a good place to start.

Article FAQ

Questions readers usually ask next.

These short answers clarify the practical follow-up questions that often come after the main article.

Need a similar system?

If this article maps to a workflow your team already operates, the next step is usually a scoped review of the system, constraints, and rollout path.

Book your free workflow review here.