Kakao said on Aug. 4 it has upgraded the voice generation technology of its in-house omni artificial intelligence (AI) model, Kanana-o.
According to details Kakao released on its tech blog the same day, when users instruct the speaking style in natural language, the AI generates speech reflecting speed, volume, pitch, emotion, intonation and intensity.
Beyond simply reading text, it can carry out complex instructions such as "Read it quickly in a sad voice" and "Like a sports broadcast," as well as role-based instructions. It was trained mainly on Korean data, but it also carries out the same speaking instructions in English voice generation.
The model scored 94.50 on the Korean benchmark InstructTTSEval, which evaluates the ability to follow speaking instructions. It was higher than OpenAI's GPT-4o-mini-tts (91.10) and at a similar level to Google's Gemini-2.5-Flash-Preview-tts (95.38).
Kakao also newly applied its in-house voice tokenizer LM-SPT to improve voice generation speed and efficiency. It compresses and represents speech with a smaller number of tokens by reducing the amount of data processed by the AI.
Kakao said LM-SPT posted superior indicators compared with models that apply the latest global technologies, including Mimi, DualCodec and CosyVoice2. Kakao plans to conduct research to handle voice understanding and generation capabilities in an integrated structure and to further develop technology that generates non-verbal expressions such as laughter, sighs and exclamations.
Byeong-seok Noh (노병석), performance leader for Kakao's Unified Foundation Model, said, "This upgrade of Kanana-o voice technology focused on building the capability to express not only natural, human-like speech generation but also the speaking style, emotion and intonation users want, following natural-language instructions." He added, "We will apply the Kanana-o model to various services to provide a more natural and convenient AI voice experience."