AI & Enterprise
PolyAI launches speech model that does not go through transcription, reflects intonation and emotions
PolyAI has launched a real-time voice conversational AI model, Dialogue-RSN-1, designed to recognise speech directly and respond with less delay in phone calls. The model performs speech recognition and processing inside the model and separates only voice output into a separate module. It interprets raw audio rather than transcripts, aiming to better capture intonation, rhythm and context. PolyAI said it can respond within 300 milliseconds.