[Photo: Shutterstock]

[Digital Today reporter Chi-gyu Hwang] PolyAI has launched a real-time voice conversational AI model, Dialogue-RSN-1, that recognises speech directly and responds. SiliconANGLE reported on Wednesday that the model focuses on reducing response delays in phone calls and delivering a more human-like conversational flow.

Conventional voice AI typically goes through speech recognition and processing, text generation and voice output. Dialogue-RSN-1 performs speech recognition and processing together inside the model, while separating only voice output into a separate module. It is designed to better capture intonation, speech rhythm and conversational context that a separate preprocessing model can easily miss.

Dialogue-RSN-1 interprets raw audio itself rather than transcripts. This means it can better handle words with unclear pronunciation and expressions that must be distinguished by context, the company said. For example, it can distinguish based on context cases where a speaker spells part of a word, such as "Matthew with two T's", and brand names such as "Audibel" that can be confused with common words.

Response speed was also presented as a strength. Dialogue-RSN operates with a delay of 280 to 500 milliseconds, and can respond within 300 milliseconds, PolyAI said. OpenAI's GPT Realtime 2.1, offered as a comparison, is at 860 to 1,900 milliseconds.

Keyword

#PolyAI #Dialogue-RSN-1 #OpenAI #GPT Realtime 2.1 #SiliconANGLE
Copyright © DigitalToday. All rights reserved. Unauthorized reproduction and redistribution are prohibited.