OpenTranscription
OpenTranscription
RankerModelsPlayground
OpenTranscription
OpenTranscription

One API to every speech-to-text model worth using. Compare them on your audio, route to the best one, pay per second.

Platform status

Product

RankerModelsTranscriptionsPlaygroundBlog

Developers

DocumentationReliabilityAPI VersioningStatus

Legal

Privacy PolicyTerms of ServiceSupport
© 2026 OpenTranscription

The ranker

Objective transcription benchmarks.

Every model. Same audio. Same metrics. Re-run on every model or pricing change.

10 free transcriptions, up to 2 hours on signup — no card.

Start freeOpen playground

10 free transcriptions, up to 2 hours on signup — no card.

34

Models Ranked

1,092

Total Benchmarks

11

Languages Tested

Jul 28, 2026

Last Updated

01 · Leader

Whisper Large V3 Turbo

Groq

81.3

Score

WER

14.31%

Latency

3.4s

Cost

$0.0007/min

02 · Runner-up

GPT-4o Mini Transcribe

OpenAI

81.1

Score

WER

13.58%

Latency

2.0s

Cost

$0.003/min

03 · Third

Voxtral Mini Transcribe

Mistral (Voxtral)

80.7

Score

WER

12.70%

Latency

2.3s

Cost

$0.003/min

Download CSV
#ModelOverall ScoreWERCERLatencyCostBenchmarks

01

Whisper Large V3 Turbo

Groq

81.3

14.31%

10.63%

3.4s

$0.0007/min

33

02

GPT-4o Mini Transcribe

OpenAI

81.1

13.58%

9.32%

2.0s

$0.003/min

35

03

Voxtral Mini Transcribe

Mistral (Voxtral)

80.7

12.70%

7.97%

2.3s

$0.003/min

30

04

Whisper Large V3

Groq

79.9

14.41%

11.20%

3.1s

$0.0019/min

33

05

Scribe v2

ElevenLabs

78.2

9.15%

7.19%

3.1s

$0.004/min

35

06

Nova-2 Voicemail

Deepgram

76.7

17.51%

12.36%

954ms

$0.0058/min

25

07

Nova-2

Deepgram

76.2

16.54%

11.93%

1.3s

$0.0058/min

34

08

Nova-2 Phone Call

Deepgram

75.7

16.38%

12.42%

1.5s

$0.0058/min

25

09

Azure Speech

Microsoft Azure

74.4

15.34%

10.67%

2.0s

$0.006/min

35

10

GPT-4o Transcribe

OpenAI

74.0

14.76%

11.60%

2.2s

$0.006/min

35

11

Soniox STT (Async)

Soniox

73.7

8.73%

6.18%

6.2s

$0.0017/min

34

12

Nova-2 Meeting

Deepgram

71.3

16.74%

11.73%

2.9s

$0.0058/min

25

13

Nova-3

Deepgram

71.1

20.98%

11.48%

1.0s

$0.0077/min

35

14

Nova-2 Finance

Deepgram

69.6

18.92%

13.09%

3.1s

$0.0058/min

25

15

Whisper Large V3

OpenAI

69.3

17.67%

13.06%

3.3s

$0.006/min

35

16

Qwen3 ASR Flash

Alibaba (Qwen)

69.1

12.27%

7.39%

6.9s

$0.0021/min

35

17

Whisper 1 (API)

OpenAI

69.0

17.63%

13.03%

3.4s

$0.006/min

35

18

Nova-2 Conversational AI

Deepgram

67.6

18.23%

13.27%

3.9s

$0.0058/min

25

19

Universal-3.5 Pro

AssemblyAI

67.4

17.44%

10.37%

5.6s

$0.0035/min

35

20

Base

Deepgram

65.2

25.04%

18.09%

772ms

$0.0145/min

34

21

Enhanced

Deepgram

63.8

24.86%

14.27%

1.3s

$0.0165/min

32

22

Speechmatics Standard

Speechmatics

60.9

14.69%

9.80%

7.3s

$0.005/min

35

23

Rev AI Reverb

Rev AI

54.9

18.18%

13.46%

13.7s

$0.003/min

35

24

Gladia Solaria-3

Gladia

52.2

9.80%

5.76%

7.6s

$0.0101/min

20

25

Amazon Transcribe

Amazon Web Services

51.5

13.02%

9.17%

12.4s

$0.006/min

35

26

GPT-4o Transcribe Diarize

OpenAI

50.7

14.61%

9.97%

15.2s

$0.006/min

33

27

Speechmatics Enhanced

Speechmatics

47.3

12.30%

7.96%

12.1s

$0.0083/min

35

28

Chirp 3

Google Cloud

45.0

9.95%

5.81%

11.0s

$0.0107/min

34

29

Amazon Transcribe Medical

Amazon Web Services

42.5

14.92%

9.67%

11.4s

$0.075/min

25

30

Google Telephony

Google Cloud

39.9

20.24%

12.59%

21.0s

$0.016/min

32

31

Google Latest (Long)

Google Cloud

38.7

22.55%

13.81%

13.2s

$0.0107/min

34

32

Google Default

Google Cloud

32.7

34.65%

25.66%

13.4s

$0.016/min

35

33

Google Command & Search

Google Cloud

32.4

35.22%

26.46%

12.7s

$0.016/min

35

34

Google Latest (Short)

Google Cloud

21.1

57.82%

52.44%

11.3s

$0.016/min

34

01

Whisper Large V3 Turbo

81.3

14.31% WER

3.4s

$0.0007/min

02

GPT-4o Mini Transcribe

81.1

13.58% WER

2.0s

$0.003/min

03

Voxtral Mini Transcribe

80.7

12.70% WER

2.3s

$0.003/min

04

Whisper Large V3

79.9

14.41% WER

3.1s

$0.0019/min

05

Scribe v2

78.2

9.15% WER

3.1s

$0.004/min

06

Nova-2 Voicemail

76.7

17.51% WER

954ms

$0.0058/min

07

Nova-2

76.2

16.54% WER

1.3s

$0.0058/min

08

Nova-2 Phone Call

75.7

16.38% WER

1.5s

$0.0058/min

09

Azure Speech

74.4

15.34% WER

2.0s

$0.006/min

10

GPT-4o Transcribe

74.0

14.76% WER

2.2s

$0.006/min

11

Soniox STT (Async)

73.7

8.73% WER

6.2s

$0.0017/min

12

Nova-2 Meeting

71.3

16.74% WER

2.9s

$0.0058/min

13

Nova-3

71.1

20.98% WER

1.0s

$0.0077/min

14

Nova-2 Finance

69.6

18.92% WER

3.1s

$0.0058/min

15

Whisper Large V3

69.3

17.67% WER

3.3s

$0.006/min

16

Qwen3 ASR Flash

69.1

12.27% WER

6.9s

$0.0021/min

17

Whisper 1 (API)

69.0

17.63% WER

3.4s

$0.006/min

18

Nova-2 Conversational AI

67.6

18.23% WER

3.9s

$0.0058/min

19

Universal-3.5 Pro

67.4

17.44% WER

5.6s

$0.0035/min

20

Base

65.2

25.04% WER

772ms

$0.0145/min

21

Enhanced

63.8

24.86% WER

1.3s

$0.0165/min

22

Speechmatics Standard

60.9

14.69% WER

7.3s

$0.005/min

23

Rev AI Reverb

54.9

18.18% WER

13.7s

$0.003/min

24

Gladia Solaria-3

52.2

9.80% WER

7.6s

$0.0101/min

25

Amazon Transcribe

51.5

13.02% WER

12.4s

$0.006/min

26

GPT-4o Transcribe Diarize

50.7

14.61% WER

15.2s

$0.006/min

27

Speechmatics Enhanced

47.3

12.30% WER

12.1s

$0.0083/min

28

Chirp 3

45.0

9.95% WER

11.0s

$0.0107/min

29

Amazon Transcribe Medical

42.5

14.92% WER

11.4s

$0.075/min

30

Google Telephony

39.9

20.24% WER

21.0s

$0.016/min

31

Google Latest (Long)

38.7

22.55% WER

13.2s

$0.0107/min

32

Google Default

32.7

34.65% WER

13.4s

$0.016/min

33

Google Command & Search

32.4

35.22% WER

12.7s

$0.016/min

34

Google Latest (Short)

21.1

57.82% WER

11.3s

$0.016/min

Accuracy tradeoffs

WER vs. cost & speed

Category leaders

no single winner

Code-Switching

Scribe v2

ElevenLabs

14.64% WER

n=2 · too few to separate

Conversational

Chirp 3

Google Cloud

4.68% WER

n=2 · too few to separate

Finance

Universal-3.5 Pro

AssemblyAI

5.63% WER

n=2 · too few to separate

General

Amazon Transcribe Medical

Amazon Web Services

4.70% WER

within margin · 1.40%–9.96%

Legal

Azure Speech

Microsoft Azure

1.66% WER

n=2 · too few to separate

Medical

Soniox STT (Async)

Soniox

1.62% WER

n=2 · too few to separate

Noisy Environment

Nova-3

Deepgram

21.82% WER

within margin · 6.82%–35.00%

Technical

Nova-3

Deepgram

2.82% WER

n=2 · too few to separate

How the score is built

weighted composite

50%

Accuracy

WER + CER vs. reference transcripts

30%

Speed

Median end-to-end latency

20%

Cost efficiency

Price per minute of audio

How We Benchmark

open methodology

01

Golden set

Curated test audio with verified human reference transcripts across languages, accents, and noise levels. Some public read-speech corpora (LibriSpeech, FLEURS, Common Voice) may appear in model training data, which can flatter the models trained on them; we counter this by also scoring across domain audio — legal, financial, and conversational — that is unlikely to be in standard training sets.

02

Same audio, vendor-recommended settings

Every model runs the identical test set — no cherry-picked clips. Our rule for calls: each model uses its provider's recommended settings — the configuration real users would run — not a lowest-common-denominator call that penalizes a model for its vendor's best practices (e.g. OpenAI's gpt-4o models get server-side VAD chunking).

03

Event-driven & automated

Benchmarks re-run automatically on every model or pricing change, with no human bias. Every provider gets the same test audio.

04

Scoring

Overall score is a weighted composite: 50% accuracy (WER), 30% speed, 20% cost efficiency. Category leaders show a confidence interval; where clips are too few to separate models, we say so rather than crown a winner.

Corpus sources & licenses

The golden set is built from openly-licensed audio. LibriSpeech and FLEURS are CC BY 4.0 and require attribution; the rest is credited for transparency.

LibriSpeech (test-clean) — CC BY 4.0

FLEURS — CC BY 4.0

Mozilla Common Voice — CC0 1.0

Earnings-21 — CC BY-SA 4.0

Bangor Miami (TalkBank) — GPLv3

Spoken Wikipedia CS Corpus — CC BY-SA 3.0

U.S. government recordings (SCOTUS, NIH/CDC) — Public domain

Original recordings — OpenTranscription

Evaluation metrics

how errors are counted

WER

Word Error Rate

(Wrong + Extra + Missed words) ÷ Total words spoken

Percentage of words incorrectly transcribed (lower is better)

CER

Character Error Rate

(Wrong + Extra + Missed characters) ÷ Total characters spoken

Percentage of characters incorrectly transcribed (lower is better)

MER

Match Error Rate

(Wrong + Extra + Missed) ÷ (Correct + Wrong + Extra + Missed)

Ratio of errors to total alignment length (lower is better)

WIL

Word Information Lost

1 − (Correct ÷ Words spoken × Correct ÷ Words predicted)

Fraction of word information lost in transcription (lower is better)

Route to whichever model wins.

One endpoint, every provider. Pin the leader or let us auto-route to the best model under your accuracy and latency budget.

Get started for freeRead the spec ↗