Launch

Expanding the Pareto Frontier of STT with Nari Qwen3-ASR

Qwen3-ASR is a leading model that has ranked in the top 10 of the Hugging Face Open ASR Leaderboard for over six months, showing competitive performance with closed models.

However, inference support has lagged behind. We built a custom inference engine to preserve accuracy while improving speed and cost efficiency. Here are the results:

  • Ranked Top 5 in accuracy (~4% WER) on Coval Benchmarks
    • Outperforms Gemini Transcribe 3.5, Cartesia Ink 2, and ElevenLabs Scribe v2 RT
  • Fast: Ranked #1 in time-to-final segment on Coval Benchmarks (sub-50 ms)
  • Standard: #1 Lowest Price of All Streaming STT models ($0.06 per hour)

Nari Qwen3-ASR enters public beta today, and everyone can use our Fast and Standard endpoints for free, only for a limited time.

We think that low costs and latency for multimodality will bring so many more applications to life. Ever needed transcription for your product, but didn't bother because it was expensive? Now is the time. Also, we are working hard to lower the prices even further.

Free Public Beta accounts have an organization-level concurrency limit of 2 shared across both ASR endpoints. Daily request limits apply separately to each endpoint. Free endpoints are provided on a best-effort basis. We cannot guarantee SLOs, including uptime or latency.

Endpoint Requests per day
qwen3-asr-fast:free Up to 100
qwen3-asr:free Up to 100

We are rolling out high concurrency support starting with our partners and plan to go GA as soon as possible. If you are interested in early access, please contact founders@narilabs.com.