AI Weekly Malaysia

Back to items Summaries

Gemini Live audio

ID
24905
Status
summarized
Published
16 Sep 2026, 6:47 AM
Fetched
16 Sep 2026, 8:04 AM
Provider
Simon Willison
Category
developer-ai
Original URL
https://simonwillison.net/2026/Sep/15/gemini-live/
Source URL
https://simonwillison.net/atom/everything/

Summary

Score
7.0
Created
16 Sep 2026, 8:04 AM
Tags
Audience
developersai_agent_usersai_ml_learners

What happened

Google released Gemini 3.8 Live and 3.8 Live Extended Thinking, two new speech-to-speech models comparable to OpenAI's GPT-Live family. Simon Willison built a zero-library browser UI that connects directly to Google's WebSocket endpoint (wss://generativelanguage.googleapis.com/ws/...BidiGenerateContent) using the Web Audio API for bidirectional voice conversation, including mid-speech interruption.

Why it matters

The WebSocket API and tutorial mean you can prototype real-time voice agents in a browser with no SDK dependencies—useful if you're building voice-first AI agent interfaces and want to evaluate Gemini's speech-to-speech quality against OpenAI's equivalent before committing to a provider.

Discussion angle

Compare Gemini 3.8 Live's WebSocket-only integration path against OpenAI's Realtime API—what does a no-SDK, direct-WebSocket approach mean for latency, cost, and deployment complexity in production voice agents?

Top