The ranker
Every model. Same audio. Same metrics. Re-run on every model or pricing change.
10 free transcriptions, up to 2 hours on signup — no card.
34
Models Ranked
1,092
Total Benchmarks
11
Languages Tested
Jul 28, 2026
Last Updated
01 · Leader
Whisper Large V3 Turbo
Groq
81.3
Score
WER
14.31%
Latency
3.4s
Cost
$0.0007/min
02 · Runner-up
GPT-4o Mini Transcribe
OpenAI
81.1
Score
WER
13.58%
Latency
2.0s
Cost
$0.003/min
03 · Third
Voxtral Mini Transcribe
Mistral (Voxtral)
80.7
Score
WER
12.70%
Latency
2.3s
Cost
$0.003/min
| # | Model | Overall Score | WER | CER | Latency | Cost | Benchmarks |
|---|---|---|---|---|---|---|---|
01 | Whisper Large V3 Turbo Groq | 81.3 | 14.31% | 10.63% | 3.4s | $0.0007/min | 33 |
02 | GPT-4o Mini Transcribe OpenAI | 81.1 | 13.58% | 9.32% | 2.0s | $0.003/min | 35 |
03 | Voxtral Mini Transcribe Mistral (Voxtral) | 80.7 | 12.70% | 7.97% | 2.3s | $0.003/min | 30 |
04 | Whisper Large V3 Groq | 79.9 | 14.41% | 11.20% | 3.1s | $0.0019/min | 33 |
05 | Scribe v2 ElevenLabs | 78.2 | 9.15% | 7.19% | 3.1s | $0.004/min | 35 |
06 | Nova-2 Voicemail Deepgram | 76.7 | 17.51% | 12.36% | 954ms | $0.0058/min | 25 |
07 | Nova-2 Deepgram | 76.2 | 16.54% | 11.93% | 1.3s | $0.0058/min | 34 |
08 | Nova-2 Phone Call Deepgram | 75.7 | 16.38% | 12.42% | 1.5s | $0.0058/min | 25 |
09 | Azure Speech Microsoft Azure | 74.4 | 15.34% | 10.67% | 2.0s | $0.006/min | 35 |
10 | GPT-4o Transcribe OpenAI | 74.0 | 14.76% | 11.60% | 2.2s | $0.006/min | 35 |
11 | Soniox STT (Async) Soniox | 73.7 | 8.73% | 6.18% | 6.2s | $0.0017/min | 34 |
12 | Nova-2 Meeting Deepgram | 71.3 | 16.74% | 11.73% | 2.9s | $0.0058/min | 25 |
13 | Nova-3 Deepgram | 71.1 | 20.98% | 11.48% | 1.0s | $0.0077/min | 35 |
14 | Nova-2 Finance Deepgram | 69.6 | 18.92% | 13.09% | 3.1s | $0.0058/min | 25 |
15 | Whisper Large V3 OpenAI | 69.3 | 17.67% | 13.06% | 3.3s | $0.006/min | 35 |
16 | Qwen3 ASR Flash Alibaba (Qwen) | 69.1 | 12.27% | 7.39% | 6.9s | $0.0021/min | 35 |
17 | Whisper 1 (API) OpenAI | 69.0 | 17.63% | 13.03% | 3.4s | $0.006/min | 35 |
18 | Nova-2 Conversational AI Deepgram | 67.6 | 18.23% | 13.27% | 3.9s | $0.0058/min | 25 |
19 | Universal-3.5 Pro AssemblyAI | 67.4 | 17.44% | 10.37% | 5.6s | $0.0035/min | 35 |
20 | Base Deepgram | 65.2 | 25.04% | 18.09% | 772ms | $0.0145/min | 34 |
21 | Enhanced Deepgram | 63.8 | 24.86% | 14.27% | 1.3s | $0.0165/min | 32 |
22 | Speechmatics Standard Speechmatics | 60.9 | 14.69% | 9.80% | 7.3s | $0.005/min | 35 |
23 | Rev AI Reverb Rev AI | 54.9 | 18.18% | 13.46% | 13.7s | $0.003/min | 35 |
24 | Gladia Solaria-3 Gladia | 52.2 | 9.80% | 5.76% | 7.6s | $0.0101/min | 20 |
25 | Amazon Transcribe Amazon Web Services | 51.5 | 13.02% | 9.17% | 12.4s | $0.006/min | 35 |
26 | GPT-4o Transcribe Diarize OpenAI | 50.7 | 14.61% | 9.97% | 15.2s | $0.006/min | 33 |
27 | Speechmatics Enhanced Speechmatics | 47.3 | 12.30% | 7.96% | 12.1s | $0.0083/min | 35 |
28 | Chirp 3 Google Cloud | 45.0 | 9.95% | 5.81% | 11.0s | $0.0107/min | 34 |
29 | Amazon Transcribe Medical Amazon Web Services | 42.5 | 14.92% | 9.67% | 11.4s | $0.075/min | 25 |
30 | Google Telephony Google Cloud | 39.9 | 20.24% | 12.59% | 21.0s | $0.016/min | 32 |
31 | Google Latest (Long) Google Cloud | 38.7 | 22.55% | 13.81% | 13.2s | $0.0107/min | 34 |
32 | Google Default Google Cloud | 32.7 | 34.65% | 25.66% | 13.4s | $0.016/min | 35 |
33 | Google Command & Search Google Cloud | 32.4 | 35.22% | 26.46% | 12.7s | $0.016/min | 35 |
34 | Google Latest (Short) Google Cloud | 21.1 | 57.82% | 52.44% | 11.3s | $0.016/min | 34 |
01
Whisper Large V3 Turbo
81.3
14.31% WER
3.4s
$0.0007/min
02
GPT-4o Mini Transcribe
81.1
13.58% WER
2.0s
$0.003/min
03
Voxtral Mini Transcribe
80.7
12.70% WER
2.3s
$0.003/min
04
Whisper Large V3
79.9
14.41% WER
3.1s
$0.0019/min
05
Scribe v2
78.2
9.15% WER
3.1s
$0.004/min
06
Nova-2 Voicemail
76.7
17.51% WER
954ms
$0.0058/min
07
Nova-2
76.2
16.54% WER
1.3s
$0.0058/min
08
Nova-2 Phone Call
75.7
16.38% WER
1.5s
$0.0058/min
09
Azure Speech
74.4
15.34% WER
2.0s
$0.006/min
10
GPT-4o Transcribe
74.0
14.76% WER
2.2s
$0.006/min
11
Soniox STT (Async)
73.7
8.73% WER
6.2s
$0.0017/min
12
Nova-2 Meeting
71.3
16.74% WER
2.9s
$0.0058/min
13
Nova-3
71.1
20.98% WER
1.0s
$0.0077/min
14
Nova-2 Finance
69.6
18.92% WER
3.1s
$0.0058/min
15
Whisper Large V3
69.3
17.67% WER
3.3s
$0.006/min
16
Qwen3 ASR Flash
69.1
12.27% WER
6.9s
$0.0021/min
17
Whisper 1 (API)
69.0
17.63% WER
3.4s
$0.006/min
18
Nova-2 Conversational AI
67.6
18.23% WER
3.9s
$0.0058/min
19
Universal-3.5 Pro
67.4
17.44% WER
5.6s
$0.0035/min
20
Base
65.2
25.04% WER
772ms
$0.0145/min
21
Enhanced
63.8
24.86% WER
1.3s
$0.0165/min
22
Speechmatics Standard
60.9
14.69% WER
7.3s
$0.005/min
23
Rev AI Reverb
54.9
18.18% WER
13.7s
$0.003/min
24
Gladia Solaria-3
52.2
9.80% WER
7.6s
$0.0101/min
25
Amazon Transcribe
51.5
13.02% WER
12.4s
$0.006/min
26
GPT-4o Transcribe Diarize
50.7
14.61% WER
15.2s
$0.006/min
27
Speechmatics Enhanced
47.3
12.30% WER
12.1s
$0.0083/min
28
Chirp 3
45.0
9.95% WER
11.0s
$0.0107/min
29
Amazon Transcribe Medical
42.5
14.92% WER
11.4s
$0.075/min
30
Google Telephony
39.9
20.24% WER
21.0s
$0.016/min
31
Google Latest (Long)
38.7
22.55% WER
13.2s
$0.0107/min
32
Google Default
32.7
34.65% WER
13.4s
$0.016/min
33
Google Command & Search
32.4
35.22% WER
12.7s
$0.016/min
34
Google Latest (Short)
21.1
57.82% WER
11.3s
$0.016/min
Code-Switching
Scribe v2
ElevenLabs
n=2 · too few to separate
Conversational
Chirp 3
Google Cloud
n=2 · too few to separate
Finance
Universal-3.5 Pro
AssemblyAI
n=2 · too few to separate
General
Amazon Transcribe Medical
Amazon Web Services
within margin · 1.40%–9.96%
Legal
Azure Speech
Microsoft Azure
n=2 · too few to separate
Medical
Soniox STT (Async)
Soniox
n=2 · too few to separate
Noisy Environment
Nova-3
Deepgram
within margin · 6.82%–35.00%
Technical
Nova-3
Deepgram
n=2 · too few to separate
50%
Accuracy
WER + CER vs. reference transcripts
30%
Speed
Median end-to-end latency
20%
Cost efficiency
Price per minute of audio
01
Golden set
Curated test audio with verified human reference transcripts across languages, accents, and noise levels. Some public read-speech corpora (LibriSpeech, FLEURS, Common Voice) may appear in model training data, which can flatter the models trained on them; we counter this by also scoring across domain audio — legal, financial, and conversational — that is unlikely to be in standard training sets.
02
Same audio, vendor-recommended settings
Every model runs the identical test set — no cherry-picked clips. Our rule for calls: each model uses its provider's recommended settings — the configuration real users would run — not a lowest-common-denominator call that penalizes a model for its vendor's best practices (e.g. OpenAI's gpt-4o models get server-side VAD chunking).
03
Event-driven & automated
Benchmarks re-run automatically on every model or pricing change, with no human bias. Every provider gets the same test audio.
04
Scoring
Overall score is a weighted composite: 50% accuracy (WER), 30% speed, 20% cost efficiency. Category leaders show a confidence interval; where clips are too few to separate models, we say so rather than crown a winner.
Corpus sources & licenses
The golden set is built from openly-licensed audio. LibriSpeech and FLEURS are CC BY 4.0 and require attribution; the rest is credited for transparency.
LibriSpeech (test-clean) — CC BY 4.0
Mozilla Common Voice — CC0 1.0
Bangor Miami (TalkBank) — GPLv3
Spoken Wikipedia CS Corpus — CC BY-SA 3.0
U.S. government recordings (SCOTUS, NIH/CDC) — Public domain
Original recordings — OpenTranscription
WER
Word Error Rate
Percentage of words incorrectly transcribed (lower is better)
CER
Character Error Rate
Percentage of characters incorrectly transcribed (lower is better)
MER
Match Error Rate
Ratio of errors to total alignment length (lower is better)
WIL
Word Information Lost
Fraction of word information lost in transcription (lower is better)
One endpoint, every provider. Pin the leader or let us auto-route to the best model under your accuracy and latency budget.