The ranker · benchmark 07 / 2026
Every model. Same audio. Same metrics. Re-run on every model or pricing change.
10 free transcriptions, up to 2 hours on signup — no card.
34
Models Ranked
1,091
Total Benchmarks
11
Languages Tested
Jul 17, 2026
Last Updated
01 · Leader
Ink-Whisper
Cartesia
82.5
score
WER
22.01%
Latency
695ms
Cost
$0.0022/min
02 · Runner-up
Whisper Large V3 Turbo
Groq
81.3
score
WER
14.31%
Latency
3.4s
Cost
$0.0007/min
03 · Third
GPT-4o Mini Transcribe
OpenAI
81.1
score
WER
13.58%
Latency
2.0s
Cost
$0.003/min
01
Ink-Whisper
82.5
22.01% WER
695ms
$0.0022/min
02
Whisper Large V3 Turbo
81.3
14.31% WER
3.4s
$0.0007/min
03
GPT-4o Mini Transcribe
81.1
13.58% WER
2.0s
$0.003/min
04
Voxtral Mini Transcribe
80.7
12.70% WER
2.3s
$0.003/min
05
Whisper Large V3
79.9
14.41% WER
3.1s
$0.0019/min
06
Scribe v2
78.2
9.15% WER
3.1s
$0.004/min
07
Nova-2 Voicemail
76.7
17.51% WER
954ms
$0.0058/min
08
Nova-2
76.2
16.54% WER
1.3s
$0.0058/min
09
Nova-2 Phone Call
75.7
16.38% WER
1.5s
$0.0058/min
10
GPT-4o Transcribe
74.0
14.76% WER
2.2s
$0.006/min
11
Nova-2 Meeting
71.3
16.74% WER
2.9s
$0.0058/min
12
Nova-3
71.1
20.98% WER
1.0s
$0.0077/min
13
Nova-2 Finance
69.6
18.92% WER
3.1s
$0.0058/min
14
Whisper Large V3
69.3
17.67% WER
3.3s
$0.006/min
15
Qwen3 ASR Flash
69.1
12.27% WER
6.9s
$0.0021/min
16
Whisper 1 (API)
69.0
17.63% WER
3.4s
$0.006/min
17
Universal-3 Pro
67.7
17.50% WER
5.5s
$0.0035/min
18
Nova-2 Conversational AI
67.6
18.23% WER
3.9s
$0.0058/min
19
Base
65.2
25.04% WER
772ms
$0.0145/min
20
AssemblyAI Best
61.9
13.31% WER
3.8s
$0.09/min
21
Enhanced
61.7
25.24% WER
1.9s
$0.0165/min
22
Speechmatics Standard
60.9
14.69% WER
7.3s
$0.005/min
23
Rev AI Reverb
54.9
18.18% WER
13.7s
$0.003/min
24
Gladia Solaria-3
52.2
9.80% WER
7.6s
$0.0101/min
25
Amazon Transcribe
51.5
13.02% WER
12.4s
$0.006/min
26
GPT-4o Transcribe Diarize
50.7
14.61% WER
15.2s
$0.006/min
27
Speechmatics Enhanced
49.6
15.60% WER
8.7s
$0.0083/min
28
Chirp 3
45.0
9.95% WER
11.0s
$0.0107/min
29
Amazon Transcribe Medical
42.5
14.92% WER
11.4s
$0.075/min
30
Google Telephony
39.9
20.24% WER
21.0s
$0.016/min
31
Google Latest (Long)
38.7
22.55% WER
13.2s
$0.0107/min
32
Google Default
32.7
34.65% WER
13.4s
$0.016/min
33
Google Command & Search
32.4
35.22% WER
12.7s
$0.016/min
34
Google Latest (Short)
21.1
57.82% WER
11.3s
$0.016/min
WER vs. cost & speed
no single winner
Code-Switching
Scribe v2
ElevenLabs
n=2 · too few to separate
Conversational
Chirp 3
Google Cloud
n=2 · too few to separate
Finance
Universal-3 Pro
AssemblyAI
n=2 · too few to separate
General
Amazon Transcribe Medical
Amazon Web Services
within margin · 1.31%–10.16%
Legal
Universal-3 Pro
AssemblyAI
n=2 · too few to separate
Medical
Whisper Large V3 Turbo
Groq
n=2 · too few to separate
Noisy Environment
Nova-3
Deepgram
within margin · 6.82%–35.05%
Technical
Nova-3
Deepgram
n=2 · too few to separate
weighted composite
50%
Accuracy
WER + CER vs. reference transcripts
30%
Speed
Median end-to-end latency
20%
Cost efficiency
Price per minute of audio
open methodology
01
Golden set
Curated test audio with verified human reference transcripts across languages, accents, and noise levels. Some public read-speech corpora (LibriSpeech, FLEURS, Common Voice) may appear in model training data, which can flatter the models trained on them; we counter this by also scoring across domain audio — legal, financial, and conversational — that is unlikely to be in standard training sets.
02
Same audio, vendor-recommended settings
Every model runs the identical test set — no cherry-picked clips. Our rule for calls: each model uses its provider's recommended settings — the configuration real users would run — not a lowest-common-denominator call that penalizes a model for its vendor's best practices (e.g. OpenAI's gpt-4o models get server-side VAD chunking).
03
Event-driven & automated
Benchmarks re-run automatically on every model or pricing change, with no human bias. Every provider gets the same test audio.
04
Scoring
Overall score is a weighted composite: 50% accuracy (WER), 30% speed, 20% cost efficiency. Category leaders show a confidence interval; where clips are too few to separate models, we say so rather than crown a winner.
Corpus sources & licenses
The golden set is built from openly-licensed audio. LibriSpeech and FLEURS are CC BY 4.0 and require attribution; the rest is credited for transparency.
LibriSpeech (test-clean) — CC BY 4.0
Mozilla Common Voice — CC0 1.0
Bangor Miami (TalkBank) — GPLv3
Spoken Wikipedia CS Corpus — CC BY-SA 3.0
U.S. government recordings (SCOTUS, NIH/CDC) — Public domain
Original recordings — OpenTranscription
how errors are counted
WER
Word Error Rate
Percentage of words incorrectly transcribed (lower is better)
CER
Character Error Rate
Percentage of characters incorrectly transcribed (lower is better)
MER
Match Error Rate
Ratio of errors to total alignment length (lower is better)
WIL
Word Information Lost
Fraction of word information lost in transcription (lower is better)
One endpoint, every provider. Pin the leader or let us auto-route to the best model under your accuracy and latency budget.