OpenTranscription
OpenTranscription
RankerModelsPlayground
OpenTranscription
OpenTranscription

One API to every speech-to-text model worth using. Compare them on your audio, route to the best one, pay per second.

Platform status

Product

RankerModelsTranscriptionsPlaygroundBlog

Developers

DocumentationReliabilityAPI VersioningStatus

Legal

Privacy PolicyTerms of ServiceSupport
© 2026 OpenTranscription

The ranker

Objective transcription benchmarks.

Every model. Same audio. Same metrics. Re-run on every model or pricing change.

10 free transcriptions, up to 2 hours on signup — no card.

Start freeOpen playground

10 free transcriptions, up to 2 hours on signup — no card.

35

Models Ranked

1,127

Total Benchmarks

11

Languages Tested

Jul 28, 2026

Last Updated

01 · Leader

Ink-Whisper

Cartesia

84.6

Score

WER

17.20%

Latency

773ms

Cost

$0.0022/min

02 · Runner-up

Whisper Large V3 Turbo

Groq

81.3

Score

WER

14.31%

Latency

3.4s

Cost

$0.0007/min

03 · Third

GPT-4o Mini Transcribe

OpenAI

81.1

Score

WER

13.58%

Latency

2.0s

Cost

$0.003/min

Download CSV
#ModelOverall ScoreWERCERLatencyCostBenchmarks

01

Ink-Whisper

Cartesia

84.6

17.20%

11.58%

773ms

$0.0022/min

35

02

Whisper Large V3 Turbo

Groq

81.3

14.31%

10.63%

3.4s

$0.0007/min

33

03

GPT-4o Mini Transcribe

OpenAI

81.1

13.58%

9.32%

2.0s

$0.003/min

35

04

Voxtral Mini Transcribe

Mistral (Voxtral)

80.7

12.70%

7.97%

2.3s

$0.003/min

30

05

Whisper Large V3

Groq

79.9

14.41%

11.20%

3.1s

$0.0019/min

33

06

Scribe v2

ElevenLabs

78.2

9.15%

7.19%

3.1s

$0.004/min

35

07

Nova-2 Voicemail

Deepgram

76.7

17.51%

12.36%

954ms

$0.0058/min

25

08

Nova-2

Deepgram

76.2

16.54%

11.93%

1.3s

$0.0058/min

34

09

Nova-2 Phone Call

Deepgram

75.7

16.38%

12.42%

1.5s

$0.0058/min

25

10

Azure Speech

Microsoft Azure

74.4

15.34%

10.67%

2.0s

$0.006/min

35

11

GPT-4o Transcribe

OpenAI

74.0

14.76%

11.60%

2.2s

$0.006/min

35

12

Soniox STT (Async)

Soniox

73.7

8.73%

6.18%

6.2s

$0.0017/min

34

13

Nova-2 Meeting

Deepgram

71.3

16.74%

11.73%

2.9s

$0.0058/min

25

14

Nova-3

Deepgram

71.1

20.98%

11.48%

1.0s

$0.0077/min

35

15

Nova-2 Finance

Deepgram

69.6

18.92%

13.09%

3.1s

$0.0058/min

25

16

Whisper Large V3

OpenAI

69.3

17.67%

13.06%

3.3s

$0.006/min

35

17

Qwen3 ASR Flash

Alibaba (Qwen)

69.1

12.27%

7.39%

6.9s

$0.0021/min

35

18

Whisper 1 (API)

OpenAI

69.0

17.63%

13.03%

3.4s

$0.006/min

35

19

Nova-2 Conversational AI

Deepgram

67.6

18.23%

13.27%

3.9s

$0.0058/min

25

20

Universal-3.5 Pro

AssemblyAI

67.4

17.44%

10.37%

5.6s

$0.0035/min

35

21

Base

Deepgram

65.2

25.04%

18.09%

772ms

$0.0145/min

34

22

Enhanced

Deepgram

63.8

24.86%

14.27%

1.3s

$0.0165/min

32

23

Speechmatics Standard

Speechmatics

60.9

14.69%

9.80%

7.3s

$0.005/min

35

24

Rev AI Reverb

Rev AI

54.9

18.18%

13.46%

13.7s

$0.003/min

35

25

Gladia Solaria-3

Gladia

52.2

9.80%

5.76%

7.6s

$0.0101/min

20

26

Amazon Transcribe

Amazon Web Services

51.5

13.02%

9.17%

12.4s

$0.006/min

35

27

GPT-4o Transcribe Diarize

OpenAI

50.7

14.61%

9.97%

15.2s

$0.006/min

33

28

Speechmatics Enhanced

Speechmatics

47.3

12.30%

7.96%

12.1s

$0.0083/min

35

29

Chirp 3

Google Cloud

45.0

9.95%

5.81%

11.0s

$0.0107/min

34

30

Amazon Transcribe Medical

Amazon Web Services

42.5

14.92%

9.67%

11.4s

$0.075/min

25

31

Google Telephony

Google Cloud

39.9

20.24%

12.59%

21.0s

$0.016/min

32

32

Google Latest (Long)

Google Cloud

38.7

22.55%

13.81%

13.2s

$0.0107/min

34

33

Google Default

Google Cloud

32.7

34.65%

25.66%

13.4s

$0.016/min

35

34

Google Command & Search

Google Cloud

32.4

35.22%

26.46%

12.7s

$0.016/min

35

35

Google Latest (Short)

Google Cloud

21.1

57.82%

52.44%

11.3s

$0.016/min

34

01

Ink-Whisper

84.6

17.20% WER

773ms

$0.0022/min

02

Whisper Large V3 Turbo

81.3

14.31% WER

3.4s

$0.0007/min

03

GPT-4o Mini Transcribe

81.1

13.58% WER

2.0s

$0.003/min

04

Voxtral Mini Transcribe

80.7

12.70% WER

2.3s

$0.003/min

05

Whisper Large V3

79.9

14.41% WER

3.1s

$0.0019/min

06

Scribe v2

78.2

9.15% WER

3.1s

$0.004/min

07

Nova-2 Voicemail

76.7

17.51% WER

954ms

$0.0058/min

08

Nova-2

76.2

16.54% WER

1.3s

$0.0058/min

09

Nova-2 Phone Call

75.7

16.38% WER

1.5s

$0.0058/min

10

Azure Speech

74.4

15.34% WER

2.0s

$0.006/min

11

GPT-4o Transcribe

74.0

14.76% WER

2.2s

$0.006/min

12

Soniox STT (Async)

73.7

8.73% WER

6.2s

$0.0017/min

13

Nova-2 Meeting

71.3

16.74% WER

2.9s

$0.0058/min

14

Nova-3

71.1

20.98% WER

1.0s

$0.0077/min

15

Nova-2 Finance

69.6

18.92% WER

3.1s

$0.0058/min

16

Whisper Large V3

69.3

17.67% WER

3.3s

$0.006/min

17

Qwen3 ASR Flash

69.1

12.27% WER

6.9s

$0.0021/min

18

Whisper 1 (API)

69.0

17.63% WER

3.4s

$0.006/min

19

Nova-2 Conversational AI

67.6

18.23% WER

3.9s

$0.0058/min

20

Universal-3.5 Pro

67.4

17.44% WER

5.6s

$0.0035/min

21

Base

65.2

25.04% WER

772ms

$0.0145/min

22

Enhanced

63.8

24.86% WER

1.3s

$0.0165/min

23

Speechmatics Standard

60.9

14.69% WER

7.3s

$0.005/min

24

Rev AI Reverb

54.9

18.18% WER

13.7s

$0.003/min

25

Gladia Solaria-3

52.2

9.80% WER

7.6s

$0.0101/min

26

Amazon Transcribe

51.5

13.02% WER

12.4s

$0.006/min

27

GPT-4o Transcribe Diarize

50.7

14.61% WER

15.2s

$0.006/min

28

Speechmatics Enhanced

47.3

12.30% WER

12.1s

$0.0083/min

29

Chirp 3

45.0

9.95% WER

11.0s

$0.0107/min

30

Amazon Transcribe Medical

42.5

14.92% WER

11.4s

$0.075/min

31

Google Telephony

39.9

20.24% WER

21.0s

$0.016/min

32

Google Latest (Long)

38.7

22.55% WER

13.2s

$0.0107/min

33

Google Default

32.7

34.65% WER

13.4s

$0.016/min

34

Google Command & Search

32.4

35.22% WER

12.7s

$0.016/min

35

Google Latest (Short)

21.1

57.82% WER

11.3s

$0.016/min

Accuracy tradeoffs

WER vs. cost & speed

Category leaders

no single winner

Code-Switching

Scribe v2

ElevenLabs

14.64% WER

n=2 · too few to separate

Conversational

Chirp 3

Google Cloud

4.68% WER

n=2 · too few to separate

Finance

Universal-3.5 Pro

AssemblyAI

5.63% WER

n=2 · too few to separate

General

Amazon Transcribe Medical

Amazon Web Services

4.70% WER

within margin · 1.31%–10.09%

Legal

Azure Speech

Microsoft Azure

1.66% WER

n=2 · too few to separate

Medical

Soniox STT (Async)

Soniox

1.62% WER

n=2 · too few to separate

Noisy Environment

Nova-3

Deepgram

21.82% WER

within margin · 6.82%–35.00%

Technical

Nova-3

Deepgram

2.82% WER

n=2 · too few to separate

How the score is built

weighted composite

50%

Accuracy

WER + CER vs. reference transcripts

30%

Speed

Median end-to-end latency

20%

Cost efficiency

Price per minute of audio

How We Benchmark

open methodology

01

Golden set

Curated test audio with verified human reference transcripts across languages, accents, and noise levels. Some public read-speech corpora (LibriSpeech, FLEURS, Common Voice) may appear in model training data, which can flatter the models trained on them; we counter this by also scoring across domain audio — legal, financial, and conversational — that is unlikely to be in standard training sets.

02

Same audio, vendor-recommended settings

Every model runs the identical test set — no cherry-picked clips. Our rule for calls: each model uses its provider's recommended settings — the configuration real users would run — not a lowest-common-denominator call that penalizes a model for its vendor's best practices (e.g. OpenAI's gpt-4o models get server-side VAD chunking).

03

Event-driven & automated

Benchmarks re-run automatically on every model or pricing change, with no human bias. Every provider gets the same test audio.

04

Scoring

Overall score is a weighted composite: 50% accuracy (WER), 30% speed, 20% cost efficiency. Category leaders show a confidence interval; where clips are too few to separate models, we say so rather than crown a winner.

Corpus sources & licenses

The golden set is built from openly-licensed audio. LibriSpeech and FLEURS are CC BY 4.0 and require attribution; the rest is credited for transparency.

LibriSpeech (test-clean) — CC BY 4.0

FLEURS — CC BY 4.0

Mozilla Common Voice — CC0 1.0

Earnings-21 — CC BY-SA 4.0

Bangor Miami (TalkBank) — GPLv3

Spoken Wikipedia CS Corpus — CC BY-SA 3.0

U.S. government recordings (SCOTUS, NIH/CDC) — Public domain

Original recordings — OpenTranscription

Evaluation metrics

how errors are counted

WER

Word Error Rate

(Wrong + Extra + Missed words) ÷ Total words spoken

Percentage of words incorrectly transcribed (lower is better)

CER

Character Error Rate

(Wrong + Extra + Missed characters) ÷ Total characters spoken

Percentage of characters incorrectly transcribed (lower is better)

MER

Match Error Rate

(Wrong + Extra + Missed) ÷ (Correct + Wrong + Extra + Missed)

Ratio of errors to total alignment length (lower is better)

WIL

Word Information Lost

1 − (Correct ÷ Words spoken × Correct ÷ Words predicted)

Fraction of word information lost in transcription (lower is better)

Route to whichever model wins.

One endpoint, every provider. Pin the leader or let us auto-route to the best model under your accuracy and latency budget.

Get started for freeRead the spec ↗