AI Weekly Malaysia

Back to items Summaries

ElevenLabs’ new v4 speech model supports more expression control and 90 languages

ID
29297
Status
summarized
Published
28 Sep 2026, 10:00 PM
Fetched
28 Sep 2026, 10:38 PM
Provider
TechCrunch
Category
technology
Original URL
https://techcrunch.com/2026/09/28/elevenlabs-new-v4-speech-model-supports-more-expression-control-and-90-languages/
Source URL
https://techcrunch.com/feed/

Summary

Score
6.0
Created
28 Sep 2026, 10:40 PM
Tags
Audience
developersai_ml_learnersai_agent_userssaas_founders

What happened

ElevenLabs launched two new speech models, v4 and v4 Turbo, on September 28, 2026, built on a new architecture that the company says enables voice cloning from just 10 seconds of audio. Language support rises from 70 in v3 to more than 90, with ElevenLabs reporting the largest quality gains in Japanese, Brazilian Portuguese, Mandarin and Cantonese. The models also expand v3's inline expression tags (now stackable in sequence), cut latency for voice agents, and can begin generating audio as soon as the backing LLM starts producing an answer, with handling for confrontations, escalations and holds.

Why it matters

If you run or plan a customer-facing voice agent, the concrete specs to test are the 10-second cloning requirement and the streaming behaviour where audio starts before the LLM finishes, since that is what determines whether a conversation feels turn-based or fluid. The Mandarin and Cantonese quality jump is directly relevant to Malaysian support lines and IVRs, which often serve those languages alongside English and Bahasa Malaysia, and the stacked inline tags are worth re-testing against your existing v3 prompts since tag sequencing behaviour changed. Note this is a vendor announcement with no independent benchmarks or pricing in the text, so treat the quality claims as unverified until you run your own samples.

Discussion angle

Build a small A/B: same script through v3 and v4 with stacked expression tags, in Cantonese or Mandarin, and compare whether the streaming start (audio before LLM completes) actually reduces perceived latency in a real call flow.

Top