[DigitalToday reporter Jinju Hong (홍진주)] OpenAI has launched the GPT-Live-1 API, a real-time voice conversation model. Unlike existing voice AI that waits for users to finish speaking before responding, it supports development of voice agents that can handle interruptions and turn-taking naturally during conversation.
On Sept. 11 (local time), online media outlet Gigazine reported that OpenAI began providing the GPT-Live-1 API to developers from Sept. 10. Developers can use it to fine-tune a voice agent’s tone, speed, conversation style and response flow.
The core of GPT-Live-1 is a full-duplex method. Unlike the existing structure that converts speech to text, passes it through a language model and then synthesises it back into speech, it handles input and output audio together in a single model. It is designed to keep conversations flowing even if users cut off the AI mid-answer or add new information.
OpenAI explained that this reduces voice-conversation latency and instability that occurs during information transfer between models. It also handles ambient noise and silent segments naturally and is improved to maintain context even in long conversations.
Another feature for developers is the ability to design a voice agent’s character directly. Using system prompts, developers can specify speaking speed, intonation and conversation style, and voice options are provided to support various languages and dialects.
Complex tasks can be handed off to a separate backend model. GPT-Live-1 handles voice conversations, while reasoning or tool calls are left to a text model such as GPT-6 Astra or an external model. This is expected to make it possible to build AI agents for real-world phone responses such as restaurant reservations and customer service centres.
OpenAI said performance is also higher than existing models. The company said GPT-Live-1 improved full-duplex benchmark performance by 30 percent over its existing voice model GPT-Realtime-2.1, and improved turn-switching latency and interaction quality. It added that it ranked No. 1 by the 'Tau3' standard in a voice-agent task intelligence evaluation combined with GPT-6 Astra.
Adoption by companies has also begun. EliseAI, an AI solutions provider for housing and medical institutions, participated as an early design partner and applied GPT-Live-1 in a real operating environment.
Pricing is $0.05 per minute for the voice front end. Developers can combine it with backend models and required agent functions to build voice AI for each service.
This API launch is seen as a move by OpenAI to expand real-time voice AI beyond ChatGPT’s internal features to external services such as customer service, reservations and phone consultations. Attention is focused on whether natural conversational AI that can exchange speech and interrupt will take hold as a new interface for real services, moving beyond voice AI that waits for a person to finish speaking.
GPT-Live-1 is now available in the API. Bring ChatGPT’s natural back-and-forth to your app, with voice agents that listen while they speak and work with the models and harness you choose. https://t.co/gIl1gwsBDV