Alibaba Qwen3-ASR Flash — multilingual batch transcription (25+ languages) with word-level timestamps and automatic language detection, served from Model Studio (Singapore). Handles files up to 12 hours.
25 languages · benchmarked
Qwen3 ASR Flash is a speech-to-text model from Alibaba (Qwen). In our standardized benchmarks it reaches 12.27% word error rate, ranking #5 of 35 models tested. It supports 25 languages and runs at $0.002/min — cheaper than most alternatives.
Its nearest benchmarked alternative is Soniox STT (Async). Best suited for accuracy-critical work and high-volume, cost-sensitive pipelines.
Overall Score
69.1
#5 of 35
Word Error Rate
12.27%
Character Error Rate
7.39%
Match Error Rate
11.23%
Word Info Lost
14.28%
Avg Latency
6.9s
Benchmarks Run
35
last 30 runs
Word error rate
12.27%
↓ 38.2% · 30d
3 runs ago
latest
Avg latency
6.9s
↓ 1.4% · 30d
3 runs ago
latest
overall WER · 35 models
Qwen3 ASR Flash
field · 34 models
lower-left is better
8 categories
WER
lower is better
Legal
3.6%
Technical
4.0%
Finance
6.0%
Medical
7.0%
Conversational
9.7%
General
11.3%
Code-Switching
18.4%
Noisy Environment
29.2%
| Category | WER | CER | MER | WIL | Latency | Benchmarks |
|---|---|---|---|---|---|---|
Code-Switching | 18.41% | 13.78% | 18.02% | 22.84% | 6.8s | 2 |
Conversational | 9.68% | 6.16% | 9.66% | 11.46% | 6.9s | 2 |
Finance | 5.97% | 2.77% | 5.91% | 9.46% | 6.8s | 2 |
General | 11.35% | 7.26% | 11.24% | 14.14% | 6.9s | 19 |
Legal | 3.59% | 1.35% | 3.52% | 4.77% | 6.8s | 2 |
Medical | 6.95% | 2.75% | 6.78% | 9.23% | 6.8s | 2 |
Noisy Environment | 29.17% | 15.63% | 20.99% | 26.43% | 6.8s | 4 |
Technical | 3.99% | 2.35% | 3.99% | 4.92% | 6.9s | 2 |
5 accents
| Accent | WER | Benchmarks |
|---|---|---|
African | 24.62% | 1 |
Australian | 13.85% | 1 |
Indian | 20.63% | 1 |
British | 1.49% | 1 |
American | 1.89% | 1 |
Rate
$0.0021
/min
Per-second billing. Bring your own provider key and pay your provider directly — a 5% routing fee applies (first 100 min/mo free).
Set up BYOKCost estimator
1 hour
$0.13
10 hours
$1.26
100 hours
$12.60
1,000 hours
$126
Billed per second of audio processed.
Model ID
alibaba/qwen3-asr-flash
Authenticate every request with your secret API key as a Bearer token. Issue a key from your dashboard.
Use it with your agent
Paste this into Claude Code, Cursor, Codex or any agent with a shell. It installs the CLI and the OpenTranscription skill, then transcribes with this model.
Quickstart
Upload your audio, create a job with this model, then poll the job or set a webhook_url to be notified when it completes.
Key parameters
file_path
string
required
Storage path returned by the upload step.
model
string
required
The model to run this job on.
models
string[]
Ordered fallback chain: primary first, then backups tried in order if a provider fails.
language
string
ISO 639-1 language code. Omit to auto-detect.
word_timestamps
boolean
Per-word start, end and confidence. Defaults to true; send false to leave them out.
webhook_url
string
HTTPS URL notified with the result when the job completes. Delivered events are signed (X-OT-Signature).
title
string
Display name for this job in the dashboard. Falls back to the file name.
Response
A completed job returns the transcript in the OpenTranscription Unified Schema (OTUS).
Webhooks — skip polling
We POST a signed event to your webhook_url when a job completes or fails; verify the X-OT-Signature (HMAC-SHA256, reject if older than 300 s) and dedupe on the event id, then fetch the full transcript via GET /api/v1/transcriptions/{id}.
Rate-limited per tier — see the X-RateLimit-* response headers.
Full API referenceUptime · 30d
100.00%
Error rate
0.00%
Avg latency
220.6s
Last incident
None in 30d
from the same field
same benchmark, every match
Supported Languages
Supported Formats
Features
Limits
Max file size
2 GB
Max duration
12 h