ElevenLabs’ new v4 speech model supports more expression control and 90 languages
- ID
- 29297
- Status
- summarized
- Published
- 28 Sep 2026, 10:00 PM
- Fetched
- 28 Sep 2026, 10:38 PM
- Provider
- TechCrunch
- Category
- technology
- Original URL
- https://techcrunch.com/2026/09/28/elevenlabs-new-v4-speech-model-supports-more-expression-control-and-90-languages/
- Source URL
- https://techcrunch.com/feed/
Summary
- Score
- 6.0
- Created
- 28 Sep 2026, 10:40 PM
- Tags
- Audience
- developersai_ml_learnersai_agent_userssaas_founders
What happened
ElevenLabs launched two new speech models, v4 and v4 Turbo, on September 28, 2026, built on a new architecture that the company says enables voice cloning from just 10 seconds of audio. Language support rises from 70 in v3 to more than 90, with ElevenLabs reporting the largest quality gains in Japanese, Brazilian Portuguese, Mandarin and Cantonese. The models also expand v3's inline expression tags (now stackable in sequence), cut latency for voice agents, and can begin generating audio as soon as the backing LLM starts producing an answer, with handling for confrontations, escalations and holds.
Why it matters
If you run or plan a customer-facing voice agent, the concrete specs to test are the 10-second cloning requirement and the streaming behaviour where audio starts before the LLM finishes, since that is what determines whether a conversation feels turn-based or fluid. The Mandarin and Cantonese quality jump is directly relevant to Malaysian support lines and IVRs, which often serve those languages alongside English and Bahasa Malaysia, and the stacked inline tags are worth re-testing against your existing v3 prompts since tag sequencing behaviour changed. Note this is a vendor announcement with no independent benchmarks or pricing in the text, so treat the quality claims as unverified until you run your own samples.
Discussion angle
Build a small A/B: same script through v3 and v4 with stacked expression tags, in Cantonese or Mandarin, and compare whether the streaming start (audio before LLM completes) actually reduces perceived latency in a real call flow.