Google’s new AI transcription edits out your ‘ums’ and ‘ahs’
Google says that 3.5 Transcribe “represents a major advancement from our previous transcription model, Chirp 3,” especially regarding multilingual performance and wording error rates. The transcription model allows users to “edit naturally with just your voice,” according to Google, and can automatically format text and remove filler words like “um” and “uh.”
Users can provide a customized vocabulary to the model, allowing 3.5 Transcribe to automatically adapt transcription to unique spelling requirements and specialized jargon to prevent those words from being edited manually. It can also attribute speech for up to three speakers in pre-recorded audio, alongside providing word-level timestamps.
Alongside 3.5 Transcribe, Google also said that 3.5 Live and 3.5 Live Experimental updates will be coming to Gemini Audio today that build on the existing speech recognition tech powering Gemini’s voice chat mode. After we published this story, Google then reached out to say these additional models arent being launched yet, and didn’t provide a new launch date. According to the information Google previously provided, Gemini 3.5 Live is better at handling mid-sentence interruptions, language recognition, and live visual processing, while Gemini 3.5 Live Experimental goes further by narrating its progress step by step in real time while it tackles reasoning on more complex tasks.
Gemini 3.5 Transcribe is rolling out starting today in English for all macOS Gemini app users, and the Rambler dictation feature on Android in select countries and languages. It’s also available for developers in public preview in the Gemini API via AI Studio and Antigravity. Google says that Chrome support is coming soon.
Update, August 26th: Google also mentioned two Gemini 3.5 Live and 3.5 Live Experimental in information provided to The Verge prior to publication, but now says that only 3.5 Transcribe is being announced today.
Google says that 3.5 Transcribe “represents a major advancement from our previous transcription model, Chirp 3,” especially regarding multilingual performance and wording error rates. The transcription model allows users to “edit naturally with just your voice,” according to Google, and can automatically format text and remove filler words like “um”…
Recent Posts
- Asobo Studio used the soundtrack from A Plague Tale: Requiem during Resonance: A Plague Tale Legacy mocap to help the cast ‘get into the sadder moments, the darker moments’
- OpenAI, Google and dozens of other companies publish open letter calling for collective action on cyber defense
- GTA VI looks just as great as we could hope for
- Google’s AI note-taking app now allows you to interact with books
- Ubisoft apologizes for accidentally selling Heroes of Might and Magic 3 without the actual game files
Archives
- August 2026
- July 2026
- June 2026
- May 2026
- April 2026
- March 2026
- February 2026
- January 2026
- December 2025
- November 2025
- October 2025
- September 2025
- August 2025
- July 2025
- June 2025
- May 2025
- April 2025
- March 2025
- February 2025
- January 2025
- December 2024
- November 2024
- October 2024
- September 2024
- August 2024
- July 2024
- June 2024
- May 2024
- April 2024
- March 2024
- February 2024
- January 2024
- December 2023
- November 2023
- October 2023
- September 2023
- August 2023