TL;DR
Our Nemotron-VoiceChat implementation runs 64 concurrent conversations on a single H100. At full utilization, this works out to about 6.7 cents in GPU cost per conversation hour.
Nemotron-VoiceChat listens and responds in audio while also calling external tools to retrieve information or perform tasks. We extended our PersonaPlex serving architecture to keep live audio on schedule while the model generates tool calls and processes their results quickly.
Hear it in action
Nemotron checks availability and books the caller’s chosen dental appointment using tools.
Unedited conversation between an OpenAI voice model as the caller and Nemotron-VoiceChat on Nari as the agent.
What We Learned from PersonaPlex
In our previous PersonaPlex post, we made all conversations advance on the same 80 ms clock. We ran sessions in fixed batches and kept conversation state on the GPU. Networking and audio transport ran outside the model execution loop.
We used the same approach for Nemotron. By connecting the path from incoming speech to generated audio on the GPU, we reduced redundant computation and memory traffic. We also wrote a persistent kernel specifically for the model to reduce execution overhead between operations.
Scheduling Audio and Tool Calls
Tool use adds another scheduling problem. Audio processing has to keep up with incoming speech, but generating a tool-call string or reading a result that has already arrived does not need to wait for the next audio frame.
If we generate a tool call at one token every 80 ms, even a short call can take several seconds. Running all the tool work in one go can also hold up the other conversations sharing the GPU.
We first reserve GPU time for live conversations, then fit tool work into the time left before the next audio deadline. We batch sessions that are generating tool calls together. When results arrive, we use chunked prefill to process several known tokens at once. This reduces the wait before the conversation resumes.
While a session processes a tool call or result, its conversational model is paused. The other live conversations continue on their regular cadence. Prerecorded acknowledgments such as “Working on it” can play while the user waits.
Results
We ran 64 sessions with 16 using tools on a single H100. The live conversations had no missed 80 ms processing deadlines or dropped frames.
| Metric | Result |
|---|---|
| Max sessions | 64 |
| GPU cost per conversation hour at full utilization | $0.067 |
| Tool-call generation time | 0.877 s |
| Tool result arrival → live audio processing resumes | 0.781 s |
H100 SXM 80GB, 1,500-position context, BF16-based mixed precision.
How Do Other Engines Compare?
We tested vLLM-Omni, NVIDIA NIM, and NVIDIA’s C++ engine on the same H100 for concurrent capacity and single-session tool-processing latency. For the capacity tests, we used tool-using traffic and required uninterrupted playback with a 160 ms playback buffer.
| Engine | Max sessions | GPU cost per session-hour1 | Tool call: 32 tokens |
Tool result: 128 tokens |
|---|---|---|---|---|
| vLLM-Omni | 1 | $4.290 | 2.508 s | 10.189 s |
| NVIDIA NIM | 4 | $1.073 | 0.474 s | 1.919 s |
| NVIDIA C++ | 6 | $0.715 | 0.243 s | 0.953 s |
| Nari | 64 | $0.067 | 0.351 s | 0.611 s |
1 At full utilization, using an H100 instance price of $4.29/hour. The C++ engine uses Q4_K_M quantization.
The quantized C++ engine was fastest at generating individual tool calls. Our implementation generated tool calls faster than NVIDIA NIM and processed tool results faster than the quantized C++ engine.
In the vLLM-Omni path we tested, tool processing still followed the 80 ms audio cadence. NIM and C++ have fast tool-processing paths. We combine fast tool processing with batching and scheduling so that many concurrent conversations can keep meeting their audio deadlines.
Voice agents need to retrieve information and take action as part of a conversation. We continue to work on faster real-time multimodal inference so more users can use these capabilities at a lower cost.
Interested in deploying a model or optimizing your inference workload? Let’s chat.