OpenTranscription
OpenTranscription
RankerModelsPlayground
OpenTranscription
OpenTranscription

One API to every speech-to-text model worth using. Compare them on your audio, route to the best one, pay per second.

Platform status

Product

RankerModelsTranscriptionsPlaygroundBlog

Developers

DocumentationReliabilityAPI VersioningStatus

Legal

Privacy PolicyTerms of ServiceSupport
© 2026 OpenTranscription
Back to Models

Qwen3 ASR Flash vs Soniox STT (Async)

2 models compared on the same benchmark — accuracy, latency, price, language coverage and capabilities. Soniox STT (Async) leads on raw accuracy; Soniox STT (Async) is cheapest.

Open both in playground

Verdict · who wins what

Most accurate

Soniox STT (Async)

8.7% WER

Soniox

Fastest

—

—

Most languages

—

—

Cheapest

Soniox STT (Async)

$0.0017/min

Soniox

Comparing

2 models · same benchmark

1

Batch

Qwen3 ASR Flash

Alibaba (Qwen)

$0.0021/min
#5 · 69.1

Alibaba Qwen3-ASR Flash — multilingual batch transcription (25+ languages) with word-level timestamps and automatic language detection, served from Model Studio (Singapore). Handles files up to 12 hours.

Use This ModelDetails →

2

Batch

Soniox STT (Async)

Soniox

$0.0017/min
#1 · 73.7

Soniox async batch transcription — one multilingual model across 60+ languages with speaker diarization, per-word timestamps, and automatic language identification, at a low flat token-based rate.

Use This ModelDetails →

Overview

30-day benchmark average

Overall score

0–100

69.1

73.7

Accuracy rank

of 35

#5

#1

Word error rate

WER

12.3%

8.7%

Character error rate

CER

7.4%

6.2%

Match error rate

MER

11.2%

8.4%

Word info lost

WIL

14.3%

13.3%

Languages

25

25

Cost vs. accuracy

overall WER vs price

Qwen3 ASR Flash

Soniox STT (Async)

field · 33 models

lower-left is better

Accuracy by category

WER · lower is better

code_switching

18.4%
16.0%

conversational

9.7%
12.6%

finance

6.0%
6.3%

general

11.3%
6.6%

legal

3.6%
3.5%

medical

7.0%
1.6%

noisy

29.2%
21.9%

technical

4.0%
4.9%

Accuracy by accent

WER · English

african

24.6%
12.3%

australian

13.9%
12.3%

indian

20.6%
19.1%

uk

1.5%
7.5%

us

1.9%
11.3%

Speed & latency

p50 → p99

Processing speed

× realtime

0.1×

0.1×

Latency p50

7.0s

6.2s

Latency p90

7.0s

6.2s

Latency p99

7.0s

6.2s

Realtime / streaming

Streaming benchmark averages

Realtime rank

—

—

Realtime score

0–100

—

—

Time to first word

—

—

Final drain

—

—

Real-time factor

—

—

Flicker

—

—

Cadence

—

—

Reliability

Rolling 30 days

Uptime

100.00%

100.00%

Error rate

0.00%

0.00%

Pricing

billed per second

Rate

per minute

$0.0021

$0.0017

Billing granularity

per 1s

per 1s

1 hour of audio

$0.13

$0.10

100 hours

$12.60

$10.08

1,000 hours

$126.00

$100.80

Capabilities

Feature support

Speaker diarization

No

Yes

Word-level timestamps

Yes

Yes

Code-switching

No

No

Auto-detect language

Yes

Yes

Live / streaming

No

No

Custom vocabulary

No

No

Coverage & deployment

Languages · formats · regions

Languages

25

25

Mode

Batch

Batch

Region

Singapore

United States

Audio formats

mp3
wav
flac
m4a
ogg
webm
aac
amr
mp3
wav
flac
m4a
ogg
webm
aac

Max file size

2 GB

No limit

Max duration

12 h

5 h

Compliance

—

—

Frequently asked questions

Is Qwen3 ASR Flash or Soniox STT (Async) more accurate?

Soniox STT (Async) is more accurate, with a 8.7% word error rate versus 12.3% for Qwen3 ASR Flash, on our standardized benchmark.

Which is cheaper, Qwen3 ASR Flash or Soniox STT (Async)?

Soniox STT (Async) is cheaper at $0.0017/min versus $0.0021/min for Qwen3 ASR Flash.

Should you choose Qwen3 ASR Flash or Soniox STT (Async)?

Soniox STT (Async) is the stronger choice overall, leading on accuracy and lower cost.