Google Releases Gemini 3.5 Transcribe, Can Convert Voice to Text With Precision

Technology-Portfolio.Net - Google launches Gemini 3.5 Transcribe, an AI model designed to produce more accurate speech transcriptions while also understanding conversational context.
Gemini 3.5 Transcribe was developed to handle everyday speech and produce cleaner text. Google calls it the most precise speech-to-text model they have ever made.
This model does not simply transcribe users' words verbatim. Gemini 3.5 Transcribe can remove filler words and correct speech when someone corrects themselves in the middle of a sentence.
This capability makes the transcription results more like edited text. Users do not need to manually correct the recording results too much after the transcription process is complete.
Gemini 3.5 Transcribe can also be used for real-time needs as well as processing existing recordings.
For real-time use, the AI model works by converting speech into text in less than one second. Meanwhile, recorded audio can be processed with additional information such as who is speaking and when the words were spoken.
Google says the AI can recognize more than 85 languages. The system is also designed to handle differences in accents, dialects, and specialized terminology.
The model can even recognize up to three speakers in recorded audio. This capability makes it potentially useful for transcribing meetings, interviews, and conversations involving several people.
Transcription Available in Other Google Products
Google is also bringing this transcription capability to several of its products. On Android, Gemini 3.5 Transcribe is used in the Rambler feature to generate text from speech and help users edit it using their voice.
In the Gemini app for macOS, users can speak to create, edit, and summarize text directly on the screen they are using. Google is also preparing integration with Chrome so users can type by voice in various fields on websites.
For developers, Gemini 3.5 Transcribe is available in public preview through the Gemini API, Google AI Studio, and Google Antigravity. Meanwhile, companies can access it through the Gemini Enterprise Agent Platform.
Only a 4 Percent Error Rate
Google also claims that, based on measurements by Artificial Analysis, this model has an average word error rate of 4% for streaming use and 2.6% for audio that is not live.
With this capability, Google does not just want to create AI that can hear.
Gemini 3.5 Transcribe is designed to understand how humans speak and turn it into text that is more ready to use.