Bringing a Voice AI Model to Production: The Journey of Optimizing Kanana-O Serving (opens in new tab)
Kanana-O is a multimodal model that understands text, images, and audio, then responds with text and speech. Deploying it for real-time voice conversations required solving problems that do not arise during model training, including low first-response latency, concurrent users, streaming across multiple models, and uneven GPU memory demands. Kakao built the specialized Kanana-Omni Server, achieving 1.6× the throughput of vLLM-Omni at 64 concurrent users.
Kanana-O’s Three-Stage Architecture
- Thinker processes multimodal inputs and generates text.
- Talker converts Thinker’s text embeddings into sequential speech tokens.
- VoiceBox combines speech tokens into audible audio waveforms.
- In production, these components must operate concurrently rather than sequentially to deliver audio within hundreds of milliseconds.
Why a Specialized Serving Server Was Needed
- Thinker passes hidden-state embeddings directly to Talker rather than ordinary token IDs.
- These high-dimensional tensors must be transferred continuously, making serialization or CPU copies too expensive.
- Talker produces speech tokens step by step, while VoiceBox waits for enough tokens to form larger audio chunks.
- Talker also combines speaker embeddings, Thinker outputs, and its own accumulated audio embeddings, creating an input structure unlike standard autoregressive decoding.
- These constraints made a custom server more suitable than general-purpose frameworks.
Zero-Copy Data Transfer
- The server preallocates shared-memory blocks during startup.
- Thinker writes tensors into an available block, while Talker receives only metadata such as the block identifier and byte size.
- This avoids repeated allocation, copying, and serialization.
- For GPU tensors on the same node, CUDA IPC transfers data directly between GPU processes, avoiding Device→Host→Device movement.
Cascaded Streaming Pipeline
- Thinker, Talker, and VoiceBox run as overlapping asynchronous stages.
- Thinker can send its first output chunk while Talker processes earlier chunks and VoiceBox synthesizes audio from still earlier ones.
- Talker buffers speech tokens until VoiceBox has enough data to create an audio chunk.
- This pipelining significantly reduces the time before the user hears the first response.
Process Isolation and Fault Containment
- Thinker and Talker each run their own vLLM engine in separate processes.
- This avoids conflicts between CUDA contexts, model memory, KV caches, and schedulers.
- Processes are started with
spawnrather thanfork, preventing inherited CUDA state from causing corruption. - If one component fails, such as Thinker running out of memory, the other components and the API server can continue operating and be restarted independently.
Continuous Batching with vLLM
- Manually batching requests is difficult because multimodal inputs and accumulated Talker embeddings vary in size.
- The server submits requests rapidly and delegates batch construction to vLLM’s continuous-batching scheduler.
- Each request runs as an independent asynchronous generation task.
- vLLM combines requests internally during forward passes, while request IDs ensure each task receives only its own streamed output.
- This improves GPU utilization without requiring custom synchronization and padding logic.
Single FastAPI Worker and Asynchronous Execution
- Multiple Uvicorn workers would load separate copies of the vLLM engines, multiplying GPU memory usage and model-loading costs.
- Therefore, the server uses
workers=1. - Since a blocking operation would otherwise stall every connected user, the entire request path—from the API endpoint through final audio generation—is designed around
async/await. - Keeping the pipeline non-blocking allows one worker to accept and progress many concurrent requests.
Kakao’s main recommendation is to design serving infrastructure around the model’s actual dataflow rather than forcing it into a generic framework. For complex multimodal pipelines, zero-copy transfers, asynchronous cascaded streaming, process isolation, and engine-level continuous batching can be more important than simply scaling API workers.