Microsoft released MAI-Transcribe-2-Streaming on 1 October, its first live speech-to-text model. It returns draft words about 100 milliseconds after hearing audio and finishes 0.13 seconds after you stop. Artificial Analysis ranks it first for streaming accuracy at a 2.5% error rate, and it covers 60 languages.