ElevenLabs released Eleven v4 on Monday, a speech model that reads tone and emotion from plain text. A voice cloned from ten seconds of audio can speak more than 90 languages and sound local. A Turbo version starts talking about 150 milliseconds after the request, for live voice agents.