Google Launches Gemini 3.5 Transcribe in Preview
Google’s new speech model handles live multilingual audio, cleans disfluencies and powers voice-driven workflows in products such as Gemini for macOS.
Two models for live and recorded speech
Google introduced Gemini 3.5 Transcribe, a speech-to-text model offered in separate preview variants for real-time streaming and prerecorded audio. Developers can access it through the Gemini API and Google AI Studio, while Google is also using the technology in products including Gboard’s Rambler feature on Android and the Gemini app for macOS.
The prerecorded-audio model supports more than 85 languages, a 96,000-token input window and speaker attribution with timestamps for up to three speakers. Google reports word error rates of 5.04% on a selected group of FLEURS languages in non-streaming evaluation and 5.50% in streaming mode. Those results are Google’s own benchmark measurements and do not establish performance across every accent, acoustic setting or domain vocabulary.
Transcription becomes an interface
The model is designed to produce usable prose rather than a strictly literal transcript. It can remove filler words, add punctuation and formatting, interpret spoken corrections, and improve recognition of phone numbers, postal codes and order identifiers. It also detects language changes during a conversation.
In the Gemini app for macOS, Google combines transcription with other Gemini models so spoken commands can initiate tasks such as image generation, information lookup, document summarization and file analysis. Google describes this product experience as function calling, but the current Gemini API documentation marks function calling as unsupported by the Transcribe endpoint itself. Developers therefore should not treat tool use as a native capability of either transcription API.
This product-level integration still turns transcription into an input layer for agents: a user can dictate, revise wording verbally and initiate a task without switching interfaces. The same flexibility raises accuracy and governance questions because removing disfluencies or interpreting intent can alter the evidentiary value of a transcript.
Why it matters
Speech recognition is becoming infrastructure for agentic systems rather than a standalone conversion utility. Google can distribute this model across consumer software, enterprise tools and developer APIs, giving it unusually broad reach at launch. Its handling of alphanumeric strings addresses a persistent weakness in voice interfaces, while integrations such as Gemini for macOS connect speech to actions. Buyers will still need to distinguish polished dictation from faithful transcription—especially in medical, legal and compliance settings where an omitted hesitation or normalized phrase may carry meaning.