Qwen3-TTS and Qwen3-ASR Free Public Beta

Multimodal inferenceat the speed of light

Ship multimodal models with lower latency, higher efficiency and less infrastructure work.

Model APIs

Optimized model APIs

Optimized multimodal model endpoints built to be fastest in production: starting with speech.

Learn More

Realtime TTS

Free Public Beta

LatencyElevenLabs vs Nari FastTTFA

CostElevenLabs vs Nari StandardUSD / 1M characters

Streaming STT

Free Public Beta

LatencyDeepgram vs Nari FastTTFS

CostDeepgram vs Nari StandardUSD / hour of audio

Based on preliminary independent measurements of publicly available TTS and STT endpoints from US East. Results may vary by region, connection, and workload. Free endpoints are best effort and carry no latency or uptime SLO. Free Public Beta is available for a limited time; Early Access pricing is shown. Pricing checked Sep 8, 2026 from Artificial Analysis.

Inference

Dedicated Multimodal Inference for Enterprise

Model and workload-specific runtimes deliver predictable latency and production-scale serving across managed, dedicated, and private infrastructure.

Learn More

Model-specific optimizations

Kernels, batching, quantization, and routing tuned to your workload.

Proven on our stack

The same technology we used to achieve the Pareto frontier and #1 on Benchmarks.

Deployment flexibility

Managed, dedicated, or private infrastructure.

Training

Finetune a model for your use-case

Fine-tune multimodal models on data that matches your use-case and deploy using our optimized stack. Work with the team behind Dia and Narvatar.

Learn More
Two speaker profiles combine into one dialogue audio file

Open-source dialogue TTS

Dia

Our dialogue TTS model series with 2M+ downloads and 20K+ GitHub stars.

In-house avatar system

Narvatar

Our best-in-class human-centric avatar model under 5B, available in both streaming and non-streaming variants. Optimized for realtime performance and low costs.

Coming next

Your Model

Expand the frontier by working with us

What's new

Latest news from Nari Labs

Launch

Expanding the Pareto Frontier of STT with Nari Qwen3-ASR

Nari Qwen3-ASR is now available in Free Public Beta. Fast and Standard endpoints are free for a limited time.

Launch

Introducing Nari Qwen3-TTS

Nari Qwen3-TTS Free Public Beta. Fast and Standard endpoints are free for a limited time.

Research

Nari Qwen3-TTS Ranked #1 in Latency

Preliminary independent measurements put our TTS endpoint first in the benchmark snapshot.

Put multimodal models into production

Use our pre-built endpoints, bring your workload, or fine-tune a model with us. All served on our specialized inference stack.