Preferred Sources
X
Mail

Announcing “Gemini 3.5 Transcribe”: The New Speech Recognition AI Model

Gemini 3.5 Transcribe
この記事のポイント
  • Google officially announced “Gemini 3.5 Transcribe,” a new speech recognition AI model for Google AI “Gemini.”
  • Achieves the highest-precision speech recognition AI model capable of accurate text conversion while correcting background noise, complex technical terms, and unclear pronunciations.
  • New features utilizing “Gemini 3.5 Transcribe” are already rolling out, such as the “Rambler” (AI voice typing) feature in the Android “Gboard” app and voice input features in the macOS “Gemini” app.

On Wednesday, August 26, 2026 local time, Google officially announced “Gemini 3.5 Transcribe,” a new speech recognition AI model for Google AI “Gemini.”

“Gemini 3.5 Transcribe” is the highest-precision speech recognition AI model that enables accurate text conversion while precisely correcting background noise, complex technical terms, and unclear pronunciation. Even with noisy audio or conversations mixed with technical terms that were difficult to understand with conventional speech recognition, the AI understands the context and converts it into well-formed sentences.

“Gemini 3.5 Transcribe”

New features utilizing “Gemini 3.5 Transcribe” have already begun rolling out in the “Gemini” app and on Android. Examples include the “Rambler” (AI voice typing) feature in the Android “Gboard” app and the voice input feature in the macOS “Gemini” app, which automatically remove verbal fillers like “um” and “uh” from your speech and format it into easy-to-read text.

It also supports corrections during speech, such as saying “Let’s make the meeting on Tuesday—no, Wednesday,” capturing the speaker’s intent and leaving only the correct information in the text.

“Gemini 3.5 Transcribe”

“Gemini 3.5 Transcribe” supports a wide range of over 85 languages and automatically adapts to different dialects and accents. Furthermore, for recorded audio, it can automatically distinguish between up to three speakers and record who said what and when with timestamps.

For developers, it is now available to integrate into applications and services through “Google AI Studio” and the “Gemini Enterprise Agent Platform.” “Gemini 3.5 Transcribe” comes in two types: the “Live API” for real-time conversations and the “Interactions API” suited for analyzing pre-recorded audio, allowing users to choose according to their needs.

For general users, it is already available via the “Rambler” (AI voice typing) feature built into the Android “Gboard” app and the macOS “Gemini” app, and is scheduled to roll out to “Chrome” soon, enabling voice input in any text field on the web. By simply speaking, users will be able to create and edit text or even give instructions for image generation, raising expectations for its future expansion as a new alternative to keyboard typing.

Source: Google (https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe/)


*This site uses affiliate advertising.