Use your OpenAI commitments on Baseten open models. Learn more
Transcription

Qwen LogoQwen 3 ASR 1.7B

SOTA Multilingual ASR model from Alibaba. See below for diarization add-on.

Model details

View repository

Qwen3-ASR 1.7B (Streaming) is a real-time speech-to-text service that runs Alibaba's open-weight Qwen3-ASR model behind Baseten's streaming harness over WebSocket. Ideal for voice agents, live captioning, and multilingual customer-support transcription.

Feature highlights

  • Extensive multi-lingual support, including auto-detection across 30+ languages and 20+ Chinese dialects, and easy fine-tuning for even more, achieving batch-level quality at streaming latency.

  • Configurable update cadence for the partial transcriptions delivered

  • Consistent real-time latency under high volume of concurrent audio streams

Where accuracy stands

The Pareto plot below shows how this model compares with leading closed-source ASR solutions across accuracy and latency. Results are measured using Pipecat’s open-source STT Benchmark, which evaluates Semantic WER for transcription accuracy and TTFS (time to final segment) for latency.

Accuracy vs latency — semantic word-error-rate vs median time-to-final-segment (single stream), plotted against closed-source providers.Accuracy vs latency — semantic word-error-rate vs median time-to-final-segment (single stream), plotted against closed-source providers.

Concurrency tradeoffs

The graph below shows how TTFS (time to final segment) and unit economics (cost per audio hour) change as the number of concurrent WebSocket connections per GPU increases. Higher concurrency can significantly reduce cost per audio hour, with a corresponding tradeoff in latency.

Concurrency is configurable and can be tuned to achieve the latency and cost targets that best fit your workload and business requirements. See the concurrency guide for configuration details.

For a balance of latency and throughput, we recommend starting with RTX6000 using concurrency 48, and adjusting based on performance and cost needs based on the chart below if needed.

We've shown that Qwen3-ASR-1.7B can comfortably keep real-time through up to 64 concurrent streams on both NVIDIA RTX 6000 and H100 GPUs, with P95 finalization latency under 0.50 seconds.

Latency vs concurrency — how median latency holds up as concurrent streams scale, by GPU.Latency vs concurrency — how median latency holds up as concurrent streams scale, by GPU.

Adding Diarization to Qwen3-ASR streaming

Our recommended diarization pairing is Nemotron 3 Diarization. To deploy our truss that combines Qwen and Nemotron, click here.

For workload-specific optimization, you can contact an engineer for further assistance.

🔥 Trending models