Tools
AI Minute Newsroom
2026-08-22
A voice model that starts speaking in under 50 milliseconds, and the serving code is open
Nari Labs published a write-up on 19 August of a serving stack it built for Qwen3-TTS 1.7B CustomVoice, and open-sourced both the implementation and the benchmark. On a single H100 SXM it holds a 95th-percentile time-to-first-audio under 50 milliseconds at 10 requests per second, and stays below 100 milliseconds even at 20. At full load the cost works out at roughly $2 per million characters of speech. The gains come from scheduling rather than from a new model: the three components of the system, the Talker, the Code Predictor and the codec, are scheduled as one unit instead of three separate queues, requests are prioritised by deadline so streaming work does not stall behind batch work, and decoding is incremental with cached state.
Why it mattersLatency is what decides whether a spoken interface feels like a conversation or a phone menu. Under roughly 200 milliseconds a person reads the reply as an answer; past half a second they start talking over it. Getting under 50 has generally meant a proprietary stack from a voice API company, priced accordingly. This is the trick that has been reshaping text inference for two years, that most of the speed was sitting in the serving layer rather than the weights, arriving in audio, in public, with the code attached. If you are building anything that talks, the floor for acceptable just moved.
✓ Verified · 2 sources
Read in the app — free, in 9 languages
Related stories
TikTok put a shopping chatbot inside the video you are watching.
2026-10-06Microsoft's new transcriber starts writing before you finish speaking.
2026-10-05ElevenLabs is giving students a free year of its reading voice.
2026-10-05Meta's assistant keeps an hourly file on everyone in your life.
2026-10-05One command restores the Apple Intelligence off switch Apple deleted.
2026-10-05