Gemini 3.8 TTS Playground
- ID
- 27880
- Status
- summarized
- Published
- 24 Sep 2026, 1:12 AM
- Fetched
- 24 Sep 2026, 5:26 AM
- Provider
- Simon Willison
- Category
- developer-ai
- Original URL
- https://simonwillison.net/2026/Sep/23/gemini-tts-playground/
- Source URL
- https://simonwillison.net/atom/everything/
Summary
- Score
- 6.5
- Created
- 24 Sep 2026, 5:27 AM
- Tags
- Audience
- developersvibe_codersai_agent_users
What happened
Google released two new Gemini text-to-speech models—gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts—offering 2,000+ voices, custom voice cloning from a 30-second sample, and native multi-speaker conversation support. Simon Willison built a bring-your-own-key playground (vibe-coded with GPT-6 Astra) that exploits the API's open CORS policy to let users compose, preview, and share bookmarkable TTS sessions.
Why it matters
The open CORS policy means you can call Gemini TTS directly from browser-side JavaScript without a proxy—useful for shipping lightweight voice apps. At ~2.74 cents for 78 seconds of audio on Flash (not Flash-Lite), you should benchmark cost against your expected usage before committing. The multi-speaker conversation API is worth testing if you build agent voice interfaces or narration tools.
Discussion angle
Compare Gemini 3.8 TTS pricing and multi-speaker API design against existing options like OpenAI TTS or ElevenLabs—does the CORS-friendly browser-callable approach change how you'd architect a voice feature?