AI Weekly Malaysia

Back to items Summaries

Gemini 3.8 TTS Playground

ID
27880
Status
summarized
Published
24 Sep 2026, 1:12 AM
Fetched
24 Sep 2026, 5:26 AM
Provider
Simon Willison
Category
developer-ai
Original URL
https://simonwillison.net/2026/Sep/23/gemini-tts-playground/
Source URL
https://simonwillison.net/atom/everything/

Summary

Score
6.5
Created
24 Sep 2026, 5:27 AM
Tags
Audience
developersvibe_codersai_agent_users

What happened

Google released two new Gemini text-to-speech models—gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts—offering 2,000+ voices, custom voice cloning from a 30-second sample, and native multi-speaker conversation support. Simon Willison built a bring-your-own-key playground (vibe-coded with GPT-6 Astra) that exploits the API's open CORS policy to let users compose, preview, and share bookmarkable TTS sessions.

Why it matters

The open CORS policy means you can call Gemini TTS directly from browser-side JavaScript without a proxy—useful for shipping lightweight voice apps. At ~2.74 cents for 78 seconds of audio on Flash (not Flash-Lite), you should benchmark cost against your expected usage before committing. The multi-speaker conversation API is worth testing if you build agent voice interfaces or narration tools.

Discussion angle

Compare Gemini 3.8 TTS pricing and multi-speaker API design against existing options like OpenAI TTS or ElevenLabs—does the CORS-friendly browser-callable approach change how you'd architect a voice feature?

Top