Research

Nemotron-VoiceChat: 64 Tool-Using S2S Agents on 1 GPU

TL;DR

Our Nemotron-VoiceChat implementation runs 64 concurrent conversations on a single H100. At full utilization, this works out to about 6.7 cents in GPU cost per conversation hour.

Nemotron-VoiceChat listens and responds in audio while also calling external tools to retrieve information or perform tasks. We extended our PersonaPlex serving architecture to keep live audio on schedule while the model generates tool calls and processes their results quickly.

Hear it in action

Nemotron checks availability and books the caller’s chosen dental appointment using tools.

Dental appointment

Unedited conversation between an OpenAI voice model as the caller and Nemotron-VoiceChat on Nari as the agent.

What We Learned from PersonaPlex

In our previous PersonaPlex post, we made all conversations advance on the same 80 ms clock. We ran sessions in fixed batches and kept conversation state on the GPU. Networking and audio transport ran outside the model execution loop.

We used the same approach for Nemotron. By connecting the path from incoming speech to generated audio on the GPU, we reduced redundant computation and memory traffic. We also wrote a persistent kernel specifically for the model to reduce execution overhead between operations.

Scheduling Audio and Tool Calls

Tool use adds another scheduling problem. Audio processing has to keep up with incoming speech, but generating a tool-call string or reading a result that has already arrived does not need to wait for the next audio frame.

If we generate a tool call at one token every 80 ms, even a short call can take several seconds. Running all the tool work in one go can also hold up the other conversations sharing the GPU.

We first reserve GPU time for live conversations, then fit tool work into the time left before the next audio deadline. We batch sessions that are generating tool calls together. When results arrive, we use chunked prefill to process several known tokens at once. This reduces the wait before the conversation resumes.

While a session processes a tool call or result, its conversational model is paused. The other live conversations continue on their regular cadence. Prerecorded acknowledgments such as “Working on it” can play while the user waits.

Results

We ran 64 sessions with 16 using tools on a single H100. The live conversations had no missed 80 ms processing deadlines or dropped frames.

MetricResult
Max sessions64
GPU cost per conversation hour at full utilization$0.067
Tool-call generation time0.877 s
Tool result arrival → live audio processing resumes0.781 s

H100 SXM 80GB, 1,500-position context, BF16-based mixed precision.

How Do Other Engines Compare?

We tested vLLM-Omni, NVIDIA NIM, and NVIDIA’s C++ engine on the same H100 for concurrent capacity and single-session tool-processing latency. For the capacity tests, we used tool-using traffic and required uninterrupted playback with a 160 ms playback buffer.

Engine Max sessions GPU cost per session-hour1 Tool call:
32 tokens
Tool result:
128 tokens
vLLM-Omni1$4.2902.508 s10.189 s
NVIDIA NIM4$1.0730.474 s1.919 s
NVIDIA C++6$0.7150.243 s0.953 s
Nari64$0.0670.351 s0.611 s

1 At full utilization, using an H100 instance price of $4.29/hour. The C++ engine uses Q4_K_M quantization.

The quantized C++ engine was fastest at generating individual tool calls. Our implementation generated tool calls faster than NVIDIA NIM and processed tool results faster than the quantized C++ engine.

In the vLLM-Omni path we tested, tool processing still followed the 80 ms audio cadence. NIM and C++ have fast tool-processing paths. We combine fast tool processing with batching and scheduling so that many concurrent conversations can keep meeting their audio deadlines.


Voice agents need to retrieve information and take action as part of a conversation. We continue to work on faster real-time multimodal inference so more users can use these capabilities at a lower cost.

Interested in deploying a model or optimizing your inference workload? Let’s chat.