Nari Model APIs Are Now Generally Available
Nari Qwen3-TTS and Qwen3-ASR are now generally available with usage-based pricing, $20 welcome credits, and higher production limits.
Ship multimodal models with lower latency, higher efficiency and less infrastructure work.
Model APIs
Optimized multimodal model endpoints built to be fastest in production: starting with speech.
Learn MoreLatencyElevenLabs vs Nari FastTTFA
CostElevenLabs vs Nari StandardUSD / 1M characters
LatencyDeepgram vs Nari FastTTFS
CostDeepgram vs Nari StandardUSD / hour of audio
Based on preliminary independent measurements of publicly available TTS and STT endpoints from US East. Results may vary by region, connection, and workload. Pricing checked Sep 8, 2026 from Artificial Analysis.
Inference
Model and workload-specific runtimes deliver predictable latency and production-scale serving across managed, dedicated, and private infrastructure.
Learn MoreModel-specific optimizations
Kernels, batching, quantization, and routing tuned to your workload.
Proven on our stack
The same technology we used to achieve the Pareto frontier and #1 on Benchmarks.
Deployment flexibility
Managed, dedicated, or private infrastructure.
Training
Fine-tune multimodal models on data that matches your use-case and deploy using our optimized stack. Work with the team behind Dia and Narvatar.
Learn More
Open-source dialogue TTS
Dia
Our dialogue TTS model series with 2M+ downloads and 20K+ GitHub stars.
In-house avatar system
Narvatar
Our best-in-class human-centric avatar model under 5B, available in both streaming and non-streaming variants. Optimized for realtime performance and low costs.
Coming next
Your Model
Expand the frontier by working with us
What's new
Nari Qwen3-TTS and Qwen3-ASR are now generally available with usage-based pricing, $20 welcome credits, and higher production limits.
We serve 80 live PersonaPlex-7B calls on one NVIDIA H100, bringing the estimated cost of a full-duplex voice session to below 10 cents an hour.
Nari ranks first in STT latency, second in STT WER, second in TTS latency, and first in TTS WER in Coval’s September 14, 2026 benchmarks.
Use our pre-built endpoints, bring your workload, or fine-tune a model with us. All served on our specialized inference stack.