Gemini 3.5 Transcribe speech-to-text launches with streaming

APIs Tools

TL;DR: Google introduced Gemini 3.5 Transcribe, a speech-to-text model with lower WER, realtime streaming, and 85+ language support.

Summary: Google announced Gemini 3.5 Transcribe, a new speech-to-text model featuring smart transcription, function calling, and lowered word error rates. It adds custom vocabulary support, multi-speaker identification, realtime streaming, and coverage of over 85 languages.

Why it matters: This gives AI builders a production-ready STT option with multilingual and streaming capabilities. Watch for API availability and test WER on your specific domain to decide whether to adopt it.

Source: x_com