Fluent in Gyeongsang and Jeolla Dialects… Kakao Unveils “Kanana-O” Voice Technology
Advancements in ‘AI That Speaks in Your Preferred Tone’
Adjusting Tone, Speed, Emotion, and Intonation Using Natural Language Commands
Implementation of In-House Speech Tokenizer…Improved Generation Speed and Efficiency
Outperforms GPT-4o-mini-tts on Korean Benchmark
[Edaily Reporter Lee So-Hyun ] Kakao(035720)has enhanced the voice generation technology of its in-house developed Omni AI model, “Kanana-o.” A key feature is that it goes beyond simply reading text naturally; users can adjust the desired tone, emotion, and intonation using only natural language commands.
Benchmark results for the voice generation technology of Kakao’s in-house Omni AI model “Kanana-o” (Photo: Kakao)
Kakao announced on the 4th via its tech blog that it had unveiled details regarding the enhancements to Kanana-o’s voice generation technology.
The newly improved Kanana-o can control the voice expression itself according to the user’s instructions. When users input their desired speech style using natural language—such as “Read it very quickly,” “Read it in a low voice,” “Read it in a sad voice,” or “Read it in the Gyeongsang or Jeolla dialects”—the AI generates speech that reflects not only speed, volume, and pitch but also emotion, intonation, and intensity.
Role-based instructions are also supported. When users specify a particular situation or role—such as “like a sports broadcast,” “like a news anchor,” or “like reading a children’s story”—the system generates a voice tailored to that scenario. It can also execute complex instructions that combine multiple conditions simultaneously, such as “Read it quickly in a sad voice with a lower tone.”
Another key feature is that, although it was trained primarily on Korean data, it can execute the same speech instructions for English voice generation as well.
Kanana-O scored 94.50 on the “InstructTTSEval” Korean benchmark, which evaluates the ability to follow speech instructions. According to Kakao, this score surpasses the 91.10 achieved by GPT-4o Mini TTS and is comparable to the 95.38 scored by Google Gemini 2.5 Flash Preview TTS.
Improvements were made not only in speech generation quality but also in speed and efficiency. Kakao has newly implemented its proprietary speech tokenizer, “LM-SPT.” LM-SPT is a technology that enables AI to compress and represent speech using fewer tokens. By reducing the amount of data the AI must process, it helps generate speech more quickly and efficiently.
Kakao explained that LM-SPT demonstrated superior performance compared to models utilizing the latest global technologies—such as Mimi, DualCodec, and CosyVoice2—in evaluations of Korean and English speech understanding and generation capabilities across various speech language models.
It also received the highest scores in expert evaluations, which quantify the naturalness of synthesized speech and speaker similarity based on listening tests.
Kakao plans to continue advancing its Kanana-O voice technology. Through research that integrates speech understanding and generation capabilities into a single unified architecture, the company will focus on providing a more natural and seamless voice experience. It also plans to develop sophisticated voice control technologies, such as those capable of naturally generating nonverbal expressions like laughter, sighs, and exclamations.
Noh Byung-seok, Performance Leader for Kakao’s Unified Foundation Model, stated, “This enhancement of the Kanana-O voice technology focused not only on generating speech as natural as that of a human but also on enabling the system to express the tone, emotions, and intonation desired by users based on natural language instructions.” He added, “Going forward, we will apply the Kanana-O model to various services to provide an even more natural and convenient AI voice experience.”
SV INVESTMENT CORPORATION(289080)is expected to reap the full benefits of its investments in the second half of the year. This is because five of its domestic portfolio companies are preparing to list…
LOTTE Himart continued to post weak results in the second quarter due to a slump in the home appliance market. LOTTE Himart plans to expand its business and improve performance by focusing on its “Ans…
The virtual asset WEMIX has once again been embroiled in a major security incident. With this hacking incident recurring one year and five months after last year’s massive theft, “restoring trust in t…