Speech in, text out: up to an hour of audio per request, or 30 minutes with diarization or
timestamps. A live variant, gemini-3.5-transcribe-live, streams. Its "smart transcription"
mode handles self-corrections and filler words, the same problem the dictation apps
Wispr Flow and Superwhisper solve.