At present, we’re introducing Gemini 3.5 Transcribe, our most exact speech-to-text mannequin but, designed for clever voice interactions. In contrast to standard speech recognition fashions that wrestle with background noise, advanced jargon, and disfluency cleanup, Gemini 3.5 Transcribe converts uncooked audio instantly into correct, polished, formatted textual content.
Throughout our merchandise just like the Gemini app and on Android, we’ve seen shoppers already benefiting from this transcription mannequin with new voice capabilities like Rambler on Android and within the Gemini app on macOS. Now, builders can construct related capabilities with Gemini 3.5 Transcribe within the Gemini API in Google AI Studio and Gemini Enterprise Agent Platform.
We have constructed 3.5 Transcribe to plug seamlessly into your developer workflows, whether or not you’re constructing voice brokers, real-time captioning instruments, or post-call analytics pipelines. The mannequin is offered throughout two separate APIs:
- Actual-time streaming: Delivers steady, bidirectional streaming with sub-second latency for interactive voice apps by way of the Stay API utilizing
gemini-3.5-transcribe-live. - Pre-recorded audio processing: Transcribes recorded audio, conferences, name logs, and extra with speaker attribution and word-level timestamps by way of the Interactions API utilizing
gemini-3.5-transcribe.
Get extra exact and clever transcription
Gemini 3.5 Transcribe is designed to seize your pure talking model to raised perceive your intent and acknowledge customized vocabulary, so you’ll be able to execute duties along with your voice.
- Sensible transcription: Seamlessly handles self-corrections (like “let’s meet Tuesday—no, Wednesday”), removes filler phrases (“ums” and ‘“ahs”), auto-formats your textual content.
- Operate calling: The mannequin can delegate advanced duties (akin to picture technology and file evaluation) to different Gemini fashions by way of operate calls. At present accessible within the Gemini macOS app.
- Extra exact transcription: As measured by Synthetic Evaluation, achieves a mean Phrase Error Price (WER) of 4.0% for streaming and a couple of.6% for non-streaming use-cases. It reveals sturdy efficiency throughout noisy, real-world environments, precisely capturing alphanumeric entities like postal codes and order IDs.
- Customized vocabulary: Acknowledges specialised jargon and distinctive spellings by seamlessly adapting transcriptions to your offered customized vocabulary.
- World language help: Robotically detects and transcribes over 85 languages, seamlessly dealing with regional accents and numerous dialects.
- Multi-speaker identification: Precisely attributes speech in pre-recorded audio with timestamps for as much as three audio system (help for 3+ audio system is experimental).









