NewsStocksGoogle DeepMind introduces Gemini 3.5 Transcribe for real-time speech-to-text

Google DeepMind introduces Gemini 3.5 Transcribe for real-time speech-to-text

Author: Google DeepMind Blog·

Key Takeaways

  • Google said Gemini 3.5 Transcribe is its most accurate speech-to-text model so far and is built for voice-driven workflows.
  • The model offers real-time streaming transcription through the Live API and recorded-audio processing with speaker attribution and timestamps through the Interactions API.
  • Google said the model supports more than 85 languages, custom vocabulary, live language switching and multi-speaker identification for up to three speakers in pre-recorded audio.
  • The company said Gemini 3.5 Transcribe improves on Chirp 3 with lower latency and stronger word error rates, including a 70% faster time to final transcription.
  • Google said the model is available in public preview for developers, enterprises and consumers across Google products, including the Gemini app on macOS and Rambler on Android.
Google DeepMind introduces Gemini 3.5 Transcribe for real-time speech-to-text

Google DeepMind introduces Gemini 3.5 Transcribe for real-time speech-to-text

Aug. 26, 2026

Diego Melendo Casado, Senior Director of Engineering, Gemini Audio, and Luke Leonhard, Chief of Staff, Gemini Audio, on behalf of the Gemini Audio team, said Google is introducing Gemini 3.5 Transcribe, its most precise speech-to-text model to date, designed for intelligent voice interactions.

According to the company, Gemini 3.5 Transcribe differs from conventional speech recognition systems that can struggle with background noise, complex jargon and cleanup of disfluencies. Google said the model converts raw audio directly into accurate, polished and formatted text, which matters for teams building voice tools that need usable output without heavy post-processing.

Google said consumers are already using the transcription model across products such as the Gemini app and Android, including new voice capabilities such as Rambler on Android and in the Gemini app on macOS. Developers can now build similar capabilities with Gemini 3.5 Transcribe through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform.

The model is designed to fit into workflows for voice agents, real-time captioning tools and post-call analytics pipelines, areas where latency, formatting and speaker attribution can affect how quickly transcription can be turned into action. Google said it is available through two separate APIs:

  • Real-time streaming: Continuous, bidirectional streaming with sub-second latency for interactive voice applications via the Live API using gemini-3.5-transcribe-live.
  • Pre-recorded audio processing: Transcription for recorded audio, meetings, call logs and more with speaker attribution and word-level timestamps via the Interactions API using gemini-3.5-transcribe.

Features and performance

Google said Gemini 3.5 Transcribe is designed to capture natural speaking style, understand user intent and recognize custom vocabulary so users can carry out tasks with their voice.

The company highlighted several capabilities:

  • Smart transcription: Handles self-corrections, such as “let’s meet Tuesday—no, Wednesday,” removes filler words such as “ums” and “ahs,” and automatically formats text.
  • Function calling: Can delegate complex tasks, such as image generation and file analysis, to other Gemini models via function calls. Google said this is currently available in the Gemini macOS app.
  • More precise transcription: As measured by Artificial Analysis, the model reaches an average Word Error Rate of 4.0% for streaming and 2.6% for non-streaming use cases. Google said it performs strongly in noisy real-world environments and can accurately capture alphanumeric entities such as postal codes and order IDs.
  • Custom vocabulary: Recognizes specialized jargon and unique spellings by adapting transcriptions to user-provided custom vocabulary.
  • Global language support: Automatically detects and transcribes more than 85 languages, including regional accents and diverse dialects.
  • Multi-speaker identification: In pre-recorded audio, it can attribute speech with timestamps for up to three speakers, while support for more than three speakers is experimental.

Google said Gemini 3.5 Transcribe can handle live language switches and seamless streaming transcription. The company also said the model represents a major advance over its previous transcription model, Chirp 3, with new capabilities, better word error rates and lower latency. As measured by Artificial Analysis, time to final transcription improves by 70%. On the FLEURS benchmark across a set of top languages and locales, Google said the model achieves a 5.50% WER in streaming mode and 5.04% WER in non-streaming use cases.

Smart transcription across Google products

Google said Gemini 3.5 Transcribe goes beyond standard speech-to-text in Google products by adding context-aware understanding to surfaces including Gboard, Antigravity, the Gemini app and Chrome.

On Gboard on Android, the new Rambler feature uses Gemini 3.5 Transcribe to turn spoken thoughts into formatted text while filtering out filler words. Users can also use voice to make edits, correct misspellings and change writing style.

On Google Antigravity, the model pairs screen context and chat history, with user permission, to improve transcription accuracy across file names, agent thoughts and active documents.

In Google AI Studio, users can access Gemini 3.5 Transcribe in Build mode to create apps with voice input.

In the Gemini app on macOS, the model transcribes natural speech into clean, formatted text and supports voice commands that can use screen context to power more complex workflows. Google said the model can call on other Gemini models in the background to summarize local files, repurpose text across apps or generate images at the cursor using only voice.

Google said Chrome support is coming soon, allowing users to talk to type in any web field to dictate replies, draft posts or prompt Gemini in Chrome more naturally.

The company also said the model can analyze files, generate images and search in the Gemini app on macOS using only voice.

Early reviews and ecosystem support

Google said developer platforms including Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel and Vision Agents are using the Gemini Live API to build and deploy voice-driven interfaces. These platforms manage real-time media streaming infrastructure so developers can focus on user experience, which can shorten the path from model access to production voice applications.

The company also said Vivo, Intellitek Health and Lingopal have shared positive feedback on Gemini 3.5 Transcribe, pointing to its latency, accuracy and language support.

Availability

Google said Gemini 3.5 Transcribe is available in public preview:

  • For developers: Through the Gemini API in Google AI Studio and Google Antigravity.
  • For enterprises: Through Gemini Enterprise Agent Platform and coming soon to Gemini Enterprise for Customer Experience.
  • For everyone: In the Gemini app on macOS in English, in Rambler on Android in select countries and languages, and coming soon to Chrome.