Google ships Gemini 3.5 Transcribe for developers
Archive item — written before sources were shown.
Google's new speech-to-text model hits a 2.6% word error rate on recorded audio, cleans up filler words, and is now live in the Gemini API.
Google DeepMind introduced Gemini 3.5 Transcribe on August 26, a speech-to-text model built to convert raw audio directly into polished, formatted text rather than a literal word-for-word transcript. It automatically cleans up self-corrections and filler words, recognizes custom vocabulary and specialized jargon, and can attribute speech to up to three separate speakers with word-level timestamps in recorded audio. The model detects and transcribes over 85 languages and handles regional accents.
Google says the model measures a 4.0% word error rate for real-time streaming and 2.6% for pre-recorded audio, as scored by Artificial Analysis, with time-to-final-transcription improving 70% over Chirp 3, its previous transcription model. It ships as two separate endpoints: a real-time streaming version (gemini-3.5-transcribe-live) for sub-second-latency voice apps via the Live API, and a version for recorded audio, meetings, and call logs (gemini-3.5-transcribe) via the Interactions API, both available now in Google AI Studio and the Gemini Enterprise Agent Platform.
What it means for you
The two-endpoint split matters if you’re actually building with this: a live voice agent needs the streaming model, while a call-recording or meeting-notes pipeline wants the batch one, and pricing and latency characteristics differ between them. The custom-vocabulary support is the detail worth testing first if you work with technical jargon, product names, or acronyms that generic transcription models mangle, since Google is specifically claiming an improvement there over its own prior model, not just a marginal benchmark gain.
- 01Intelligent transcription with Gemini 3.5 Transcribedeepmind.google · primary
- 02Google's new AI transcription edits out your 'ums' and 'ahs'theverge.com · reporting
