Gemini Live audio
- ID
- 24905
- Status
- summarized
- Published
- 16 Sep 2026, 6:47 AM
- Fetched
- 16 Sep 2026, 8:04 AM
- Provider
- Simon Willison
- Category
- developer-ai
- Original URL
- https://simonwillison.net/2026/Sep/15/gemini-live/
- Source URL
- https://simonwillison.net/atom/everything/
Summary
- Score
- 7.0
- Created
- 16 Sep 2026, 8:04 AM
- Tags
- Audience
- developersai_agent_usersai_ml_learners
What happened
Google released Gemini 3.8 Live and 3.8 Live Extended Thinking, two new speech-to-speech models comparable to OpenAI's GPT-Live family. Simon Willison built a zero-library browser UI that connects directly to Google's WebSocket endpoint (wss://generativelanguage.googleapis.com/ws/...BidiGenerateContent) using the Web Audio API for bidirectional voice conversation, including mid-speech interruption.
Why it matters
The WebSocket API and tutorial mean you can prototype real-time voice agents in a browser with no SDK dependencies—useful if you're building voice-first AI agent interfaces and want to evaluate Gemini's speech-to-speech quality against OpenAI's equivalent before committing to a provider.
Discussion angle
Compare Gemini 3.8 Live's WebSocket-only integration path against OpenAI's Realtime API—what does a no-SDK, direct-WebSocket approach mean for latency, cost, and deployment complexity in production voice agents?