Nari Model APIs Are Generally Available

Multimodal inferenceat the speed of light

Ship multimodal models with lower latency, higher efficiency and less infrastructure work.

Model APIs

Optimized model APIs

Optimized multimodal model endpoints built to be fastest in production: starting with speech.

Learn More

Realtime TTS

LatencyElevenLabs vs Nari FastTTFA

CostElevenLabs vs Nari StandardUSD / 1M characters

Streaming STT

LatencyDeepgram vs Nari FastTTFS

CostDeepgram vs Nari StandardUSD / hour of audio

Based on preliminary independent measurements of publicly available TTS and STT endpoints from US East. Results may vary by region, connection, and workload. Pricing checked Sep 8, 2026 from Artificial Analysis.

Inference

Dedicated Multimodal Inference for Enterprise

Model and workload-specific runtimes deliver predictable latency and production-scale serving across managed, dedicated, and private infrastructure.

Learn More

Model-specific optimizations

Kernels, batching, quantization, and routing tuned to your workload.

Proven on our stack

The same technology we used to achieve the Pareto frontier and #1 on Benchmarks.

Deployment flexibility

Managed, dedicated, or private infrastructure.

Training

Finetune a model for your use-case

Fine-tune multimodal models on data that matches your use-case and deploy using our optimized stack. Work with the team behind Dia and Narvatar.

Learn More
Two speaker profiles combine into one dialogue audio file

Open-source dialogue TTS

Dia

Our dialogue TTS model series with 2M+ downloads and 20K+ GitHub stars.

In-house avatar system

Narvatar

Our best-in-class human-centric avatar model under 5B, available in both streaming and non-streaming variants. Optimized for realtime performance and low costs.

Coming next

Your Model

Expand the frontier by working with us

What's new

Latest news from Nari Labs

Launch

Nari Model APIs Are Now Generally Available

Nari Qwen3-TTS and Qwen3-ASR are now generally available with usage-based pricing, $20 welcome credits, and higher production limits.

Research

PersonaPlex 7B: A Full-Duplex Conversation Below 10 Cents an Hour

We serve 80 live PersonaPlex-7B calls on one NVIDIA H100, bringing the estimated cost of a full-duplex voice session to below 10 cents an hour.

Research

Nari Labs Leads Coval’s Voice AI Benchmarks

Nari ranks first in STT latency, second in STT WER, second in TTS latency, and first in TTS WER in Coval’s September 14, 2026 benchmarks.

Put multimodal models into production

Use our pre-built endpoints, bring your workload, or fine-tune a model with us. All served on our specialized inference stack.