aiminute. ← All AI news
New Models AI Minute Newsroom 2026-08-28

Google's new transcription model does not write down what you said — it writes down what you meant, in 85 languages

Google's new transcription model does not write down what you said — it writes down what you meant, in 85 languages

Google has released Gemini 3.5 Transcribe, a speech-to-text model built on Gemini's audio understanding and now in public preview. It automatically detects more than 85 languages, follows speakers who switch language mid-sentence, separates up to three voices, and emits word-level timestamps. The part that makes it different from ordinary dictation is what Google calls smart transcription: it strips filler words, repairs self-corrections and formats the text as it goes, so the output reads as clean prose rather than a literal record of the audio. It can also be biased towards a custom vocabulary for domain-specific terms, and can call functions to hand off tasks mid-transcription. Google reports a word error rate of 2.6 percent on recorded audio and 4.0 percent streaming, and on the multilingual FLEURS benchmark 5.04 and 5.50 percent respectively, with latency about 70 percent lower than its predecessor Chirp 3. It is available through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform, in Gboard's Rambler on Android and the Gemini app on macOS, with Chrome support said to be coming.

Why it mattersTranscription is the plumbing under a lot of ordinary work — lecture notes, medical dictation, subtitles, meeting minutes, the interview a journalist has to type up. Support for 85 languages with code-switching is the line that matters outside English: most people do not speak one language cleanly in one sitting, and until now the models handled that badly. The cleanup is a genuine trade, though. A transcript that quietly fixes your sentences is easier to read and no longer a faithful record of what was said, which is exactly the wrong property for a deposition or a disciplinary hearing.
#Audio & Music

✓ Verified · 4 sources

WhatsApp X Telegram
Read in the app — free, in 9 languages

Related stories

Google's video model can now be told where a shot starts and where it ends — and drafts cost a third as much
2026-08-28
Australia's charts now require a human to have written the song and sung the lead — the rule starts Friday
2026-08-27
The stealth model developers spent weeks probing is now MIT-licensed and downloadable — and Z.ai says the entire preview ran on Chinese chips
2026-08-27
IBM's new open models were not taught to describe using a terminal. They were trained by using one.
2026-08-27
The record labels just bought into the AI image company: Universal, Sony, Warner and EA are all in Stability's $76 million round
2026-08-26