Google unveiled a new model for its Gemini Live conversational voice assistant. [Photo provided by Google]

Google has overhauled its voice recognition artificial intelligence (AI) lineup on a large scale.

It rolled out Gemini 3.5 Live for its Gemini Live conversational voice assistant. It also unveiled Gemini 3.5 Live Experimental for developers and Gemini 3.5 Transcribe, a speech-to-text conversion model, at the same time.

On Aug. 26, local time, a number of foreign media outlets including TechRadar reported that Gemini 3.5 Live has greatly improved its ability to handle user interruptions during conversations compared with Gemini 3.1 Live. It newly supports real-time visual information processing, mixed-language recognition and background tool-calling. Until now, Gemini Live did not have a function to access and use other tools, but the update greatly broadens its scope of use.

The existing Gemini Live often made errors in speech recognition and conversation flow. In some situations, it misrecognized a user's speech as another language and switched languages. If an interruption occurred during a conversation, it would also cause confusion and sometimes revert to the default voice rather than the voice selected by the user. Google expects Gemini 3.5 Live to resolve these issues.

Gemini 3.5 Transcribe, also unveiled on the day, is a new model that replaces the existing speech-to-text model Chirp 3. Google said the time from voice input to final text conversion is about 70 percent shorter than Chirp 3, and the real-time speech recognition error rate has fallen to 5.5 percent. The improvement is not large compared with Chirp 3's 7.32 percent error rate, but it is expected to reduce the inconvenience of having to correct typos one by one during voice input.

Gemini 3.5 Transcribe does more than simply understand speech more accurately. While a user is speaking, it automatically filters out filler expressions such as "um" and "uh" and revises sentences in real time if the user corrects themselves. It also handles technical terms by referring to custom vocabulary registered in advance by the user. All of these functions work in 85 languages, and with pre-recorded audio it can distinguish and recognize up to 3 speakers.

Still, because the AI refines sentences by making its own judgments about a user's intent, a limitation is that differences can arise between the actual spoken content and the final text. It smooths awkward phrasing or stuttering in short sentences, but it is hard to say it is suitable for all situations because it changes the original expression itself.

Gemini 3.5 Transcribe has already been applied to some products. A representative example is the Rambler feature in Android Gboard, which turns long-winded speech into organized text and automatically removes filler expressions. For now, it is supported only on the Pixel 11 model, and Google said it plans to expand it within the year to "more devices equipped with Gemini intelligence." The Gemini application for macOS has also applied the model to voice input from the day, and combined with screen recognition it can handle local file summaries, moving text between apps and generating images at the cursor location using voice commands alone.

Developer support is also expanding. The coding tool Antigravity will support Gemini 3.5 Transcribe from the day, allowing it to reference screen context and conversation history with user consent, and the model is also being applied to AI Studio's build model, making it possible to develop apps with voice alone through "vibe coding." Developers can also access the model through the Gemini API.

Google said it plans to introduce the function to the Chrome browser soon. If applied, users will be able to write by voice in any text input field on the web, including email replies, social media posts and AI chatbot prompt input.

Gemini 3.5 Live, Gemini 3.5 Live Experimental and Gemini 3.5 Transcribe were all announced together on the day. Among them, Gemini 3.5 Live Experimental is characterized by explaining the process step by step in real time when handling complex reasoning tasks. The three models will be provided first in English to all users of the Gemini app for macOS and to Android Rambler users in some countries and language regions, while developers will get them as a public preview through AI Studio and Antigravity's Gemini API. Google did not separately disclose when it will roll them out to all users, but distribution is expected to take place within a few days.

The announcement comes as Google has yet to release the Gemini 3.5 Pro model, which it had said it would launch in June.

Keyword

#Google #Gemini Live #Gemini 3.5 Live #Gemini 3.5 Transcribe #Chrome
Copyright © DigitalToday. All rights reserved. Unauthorized reproduction and redistribution are prohibited.