NewsTradingSentimentEventsCommunityBriefing
Tech

Google Adds Two New Voice Models to Gemini Suite

By Tech Desk · · 1 min read
A server rack with blinking status lights in a dark data center

Google launched two text-to-speech models for Gemini, offering custom voices and detailed performance controls for developers.

Key points

  • Google released Gemini 3.8 Flash TTS and Flash-Lite TTS for creating custom voices.
  • The Flash TTS model scored 71.4 on Hume AI's Voice Design Benchmark.
  • Users can replicate voices from a 30-second sample with built-in consent checks.

Google has released two new text-to-speech models for its Gemini platform. The announcement appeared today on the company's official blog. These tools allow users to generate speech from written text. They target creators and developers building audio products.

The new features expand the Gemini family of audio tools. They complement earlier releases like 3.5 Live Translate. The goal is to make voice generation more flexible. Users can now design voices rather than just picking presets.

Two models serve different needs

One model is Gemini 3.8 Flash TTS. It is built for deep creative direction. It lets users create new voices from scratch. You can use natural language prompts to define character traits. This helps in gaming and audiobook production.

The other is Gemini 3.8 Flash-Lite TTS. It focuses on high-volume and cost-efficient scaling. It is optimized for dubbing and voice agents. It offers fine-grained control over tone and pacing. This makes it suitable for large-scale content creation.

Custom voices and performance controls

Users can access a library of over 2,000 voices. These cover many languages and regional dialects. The system also supports voice replication from a 30-second sample. Built-in consent verification and SynthID watermarking protect users. This ensures safety and ownership of vocal profiles.

Both models allow line-by-line performance direction. Users can add stage directions to scripts. The system handles non-verbal cues like laughs or sighs. It also manages two-speaker conversations naturally. This reduces the need for manual editing.

Performance metrics and trade-offs

Google reports the Flash TTS model leads in benchmarks. It secured the top spot on Hume AI’s Voice Design Benchmark. Its score was 71.4 overall. It also led in accent modeling with a score of 60.8. These figures indicate strong customization capabilities.

However, the advanced model likely costs more to run. The Lite version prioritizes efficiency over maximum creative depth. Users must choose based on their volume and quality needs. The trade-off is between creative control and processing cost.

Based on reporting by blog.google, compiled by the Tradingbird desk.

Read next

More in Tech

More from the Tech desk

All desk stories