Skip to content

Gemini 3.5 Transcribe

gemini-3.5-transcribe
Google · chat · api
GA Alert me on changes
Context
98.3K
Max output
32.8K
Input $/1M
$2.00
Output $/1M
$12.00
Modalities
text
Released
26 Aug 2026
Download image Share on X Share on LinkedIn
AI summary
● machine-written

Google releases Gemini 3.5 Transcribe speech-to-text model

Gemini 3.5 Transcribe is Google's speech-to-text model supporting 85+ languages with intelligent transcription, speaker diarization, and word-level timestamps. Available via API in unary and live streaming modes, it processes audio files up to 1 hour and offers pricing starting at $2.00 per million input tokens and $12.00 per million output tokens.

What's new
  • Supports 85+ languages with automatic detection and mid-session code-switching
  • Smart transcription mode removes filler words and resolves spoken corrections
  • Speaker diarization supported for up to 8 speakers in file mode
  • Custom vocabulary biasing with up to 1,000 terms
  • Live API variant with 10-minute session limit for real-time transcription
Best for
Meeting transcription and note-takingCall center audio analysis with speaker attributionMultilingual content processing with auto-detectionReal-time voice applications via Live API
Sources

Source: https://ai.google.dev/gemini-api/docs/models/gemini-3.5-transcribe