2 models compared on the same benchmark — accuracy, latency, price, language coverage and capabilities. Soniox STT (Async) leads on raw accuracy; Whisper Large V3 is fastest; Whisper Large V3 covers the most languages; Soniox STT (Async) is cheapest.
Verdict · who wins what
Most accurate
Soniox STT (Async)
8.7% WER
Soniox
Fastest
Whisper Large V3
0.3× realtime
OpenAI
Most languages
Whisper Large V3
98 languages
OpenAI
Cheapest
Soniox STT (Async)
$0.0017/min
Soniox
Comparing
2 models · same benchmark
1
Whisper Large V3
OpenAI
OpenAI's Whisper large-v3 model
2
Soniox STT (Async)
Soniox
Soniox async batch transcription — one multilingual model across 60+ languages with speaker diarization, per-word timestamps, and automatic language identification, at a low flat token-based rate.
30-day benchmark average
overall WER vs price
Soniox STT (Async)
Whisper Large V3
field · 33 models
lower-left is better
WER · lower is better
WER · English
p50 → p99
Streaming benchmark averages
Rolling 30 days
billed per second
Feature support
Languages · formats · regions
general
legal
medical
noisy
technical
uk
us
Latency p99
4.3s
6.2s
$10.08
1,000 hours
$360.00
$100.80
No
Auto-detect language
Yes
Yes
Live / streaming
No
No
Custom vocabulary
No
No
Max file size
25 MB
No limit
Max duration
No limit
5 h
Compliance
—
Is Whisper Large V3 or Soniox STT (Async) more accurate?
Soniox STT (Async) is more accurate, with a 8.7% word error rate versus 17.7% for Whisper Large V3, on our standardized benchmark.
Which is cheaper, Whisper Large V3 or Soniox STT (Async)?
Soniox STT (Async) is cheaper at $0.0017/min versus $0.0060/min for Whisper Large V3.
Should you choose Whisper Large V3 or Soniox STT (Async)?
Choose Whisper Large V3 for speed and language coverage; choose Soniox STT (Async) for accuracy and lower cost.