Topic: #speech-to-text
-
Google's Gemini 3.5 Transcribe adds emotion detection and speaker ID
Google unveiled Gemini 3.5 Transcribe at I/O in May 2026, a multimodal audio model that transcribes speech while detecting emotion, identifying speakers, and summarizing content—capabilities that rival OpenAI's Whisper and reshape how professionals handle recorded audio.