Microsoft (MS) [Photo: Shutterstock]

Microsoft has unveiled three voice AI models, including 1 real-time transcription model and 2 speech synthesis models.

SiliconANGLE reported on Oct. 2 that the centerpiece is MAI-Transcribe-2-Streaming, its first streaming transcription model. It converts speech to text in real time and keeps updating the transcript as the user continues speaking. It finalises the transcript when the utterance ends. It can be used for applications that display live captions or begin processing requests before a user finishes speaking.

MAI-Transcribe-2-Streaming produces its first transcription candidate in an average of 320 milliseconds. It supports more than 60 languages and automatically detects the language being used. It costs $0.54 per audio hour.

Of the 2 speech synthesis models, MAI-Voice-2.1 is designed to improve expressiveness and audio quality. MAI-Voice-2.1-Flash focuses on faster responses and lower costs by reducing some performance. Both models support 23 languages.

Microsoft last month also unveiled MAI-Transcribe-2, a non-streaming model. It costs $0.10 per audio hour, more than 5 times cheaper than the streaming model. The streaming model has a heavier processing load because it must keep returning interim results while audio is coming in.

The launch also aligns with Microsoft's strategy to expand its in-house models. Microsoft AI Chief Executive Officer Mustafa Suleyman (무스타파 술레이만) in July stressed strengthening the MAI family amid concerns about the costs of high-performance models from OpenAI and Anthropic. SiliconANGLE reported that Microsoft aims to apply its in-house models to Copilot agents on platforms such as Excel and Outlook.

Keyword

#Microsoft #MAI-Transcribe-2-Streaming #MAI-Voice-2.1 #SiliconANGLE #Mustafa Suleyman
Copyright © DigitalToday. All rights reserved. Unauthorized reproduction and redistribution are prohibited.