OpenTranscription
OpenTranscription
RankerModelsPlayground
OpenTranscription
OpenTranscription

One API to every speech-to-text model worth using. Compare them on your audio, route to the best one, pay per second.

Platform status

Product

RankerModelsTranscriptionsPlaygroundBlog

Developers

DocumentationReliabilityAPI VersioningStatus

Legal

Privacy PolicyTerms of ServiceSupport
© 2026 OpenTranscription

Catalog · 41 models · 14 providers

Transcription models.

Every speech-to-text model worth using, one rate card.

Price Range

$0.0007

$0.075

Compliance

Capabilities

Best For:

41

Total models

103

Languages

$0.0007

Cheapest /min

8.7%

Best WER

41 models

Base

Deepgram

Batch
United States

Deepgram Base model — high volume, cost-effective

$0.0145/min

WER

25.04%

CER

18.09%

Speed Factor

0.1x

Uptime

100.0%

Billing

per 1s

Retention

No retention

19 languages

(view all)

Auto-detectDiarizationWord Timestamps+1
TranscribeView Benchmarks
Chirp 3

Google Cloud

Batch
United States

Google's latest generative ASR foundation model — 85+ languages

$0.0107/min
Best Conversations
Best Accented

WER

9.95%

CER

5.81%

Speed Factor

0.3x

Uptime

100.0%

Billing

per 15s

Retention

Custom

68 languages

(view all)

GenerativeAuto-detectDiarization+2
TranscribeView Benchmarks
Google Command & Search

Google Cloud

Batch
United States

Optimized for short queries and voice commands

$0.016/min

WER

35.22%

CER

26.46%

Speed Factor

0.1x

Billing

per 15s

Retention

Custom

40 languages

(view all)

Word Timestamps
TranscribeView Benchmarks
Google Default

Google Cloud

Batch
United States

General-purpose model

$0.016/min

WER

34.65%

CER

25.66%

Speed Factor

0.2x

Billing

per 15s

Retention

Custom

40 languages

(view all)

DiarizationWord Timestamps
TranscribeView Benchmarks
Speechmatics Enhanced

Speechmatics

Batch
Europe

Highest accuracy model — 55+ languages, best-in-class

$0.0083/min

WER

12.30%

CER

7.96%

Speed Factor

0.3x

Uptime

100.0%

Billing

per 1s

Retention

No retention

48 languages

(view all)

Async APIAuto-detectDiarization+3
TranscribeView Benchmarks
Enhanced

Deepgram

Batch
United States

Deepgram Enhanced model — high accuracy for uncommon words

$0.0165/min

WER

24.86%

CER

14.27%

Speed Factor

0.1x

Uptime

100.0%

Billing

per 1s

Retention

No retention

15 languages

(view all)

Auto-detectDiarizationWord Timestamps+1
TranscribeView Benchmarks
Flux

Deepgram

Realtime
United States

First conversational ASR model built for voice agents — model-integrated endpointing

$0.0077/min

TTFW (P50)

289 ms

Speed Factor

0.1x

Uptime

92.9%

Billing

per 1s

Retention

No retention

1 languages

(view all)

endpointing
CS
turn detection+1
TranscribeView Benchmarks
GPT-4o Mini Transcribe

OpenAI

Batch
United States

GPT-4o Mini optimized for fast transcription

$0.003/min

WER

13.58%

CER

9.32%

Speed Factor

0.1x

Uptime

100.0%

Billing

per 1s

Retention

No retention

98 languages

(view all)

Auto-detect
TranscribeView Benchmarks
GPT-4o Transcribe

OpenAI

Batch
United States

GPT-4o optimized for transcription with improved WER

$0.006/min

WER

14.76%

CER

11.60%

Speed Factor

0.2x

Uptime

100.0%

Billing

per 1s

Retention

No retention

98 languages

(view all)

Auto-detect
TranscribeView Benchmarks
GPT-4o Transcribe Diarize

OpenAI

Batch
United States

GPT-4o transcription with built-in speaker diarization

$0.006/min

WER

14.61%

CER

9.97%

Speed Factor

0.2x

Uptime

0.0%

Billing

per 1s

Retention

No retention

12 languages

(view all)

Auto-detectDiarization
TranscribeView Benchmarks
Google Latest (Long)

Google Cloud

Batch
United States

Conformer model for long-form audio (minutes to hours)

$0.0107/min

WER

22.55%

CER

13.81%

Speed Factor

0.4x

Uptime

100.0%

Billing

per 15s

Retention

Custom

40 languages

(view all)

Word Timestamps
TranscribeView Benchmarks
Google Latest (Short)

Google Cloud

Batch
United States

Conformer model for short utterances (< 60s)

$0.016/min

WER

57.82%

CER

52.44%

Speed Factor

0.1x

Billing

per 15s

Retention

Custom

40 languages

(view all)

Word Timestamps
TranscribeView Benchmarks
Nova-2

Deepgram

Batch & Realtime
United States

Deepgram's Nova-2 speech recognition

$0.0058/min
Fastest

WER

16.54%

CER

11.93%

Speed Factor

0.1x

Uptime

100.0%

Billing

per 1s

Retention

No retention

33 languages

(view all)

Auto-detectDiarization
CS
+2
TranscribeView Benchmarks
Nova-2 Conversational AI

Deepgram

Batch & Realtime
United States

Optimized for human-to-bot interactions (IVR, voice assistants)

$0.0058/min

WER

18.23%

CER

13.27%

Speed Factor

0.1x

Billing

per 1s

Retention

No retention

1 languages

(view all)

Diarization
CS
Word Timestamps+1
TranscribeView Benchmarks
Nova-2 Finance

Deepgram

Batch & Realtime
Deprecated
United States

Optimized for earnings calls with finance vocabulary

$0.0058/min

WER

18.92%

CER

13.09%

Speed Factor

0.1x

Billing

per 1s

Retention

No retention

1 languages

(view all)

Diarization
CS
Word Timestamps+1
TranscribeView Benchmarks
Nova-2 Meeting

Deepgram

Batch & Realtime
United States

Optimized for conference room audio

$0.0058/min

WER

16.74%

CER

11.73%

Speed Factor

0.1x

Billing

per 1s

Retention

No retention

1 languages

(view all)

Diarization
CS
Word Timestamps+1
TranscribeView Benchmarks
Nova-2 Phone Call

Deepgram

Batch & Realtime
United States

Optimized for low-bandwidth phone call audio

$0.0058/min

WER

16.38%

CER

12.42%

Speed Factor

0.1x

Billing

per 1s

Retention

No retention

1 languages

(view all)

Diarization
CS
Word Timestamps+1
TranscribeView Benchmarks
Nova-2 Voicemail

Deepgram

Batch & Realtime
United States

Optimized for low-bandwidth single speaker voicemail

$0.0058/min

WER

17.51%

CER

12.36%

Speed Factor

0.1x

Billing

per 1s

Retention

No retention

1 languages

(view all)

CS
Word TimestampsCustom Vocabulary
TranscribeView Benchmarks
Nova-3

Deepgram

Batch & Realtime
United States

Deepgram's flagship model — 53% lower WER vs competitors, code-switching support

$0.0077/min
Best Noisy
Best Technical

WER

20.98%

CER

11.48%

Speed Factor

0.1x

Uptime

100.0%

Billing

per 1s

Retention

No retention

47 languages

(view all)

Auto-detectDiarizationSmart Format+3
TranscribeView Benchmarks
Qwen3 ASR Flash

Alibaba (Qwen)

Batch
Singapore

Alibaba Qwen3-ASR Flash — multilingual batch transcription (25+ languages) with word-level timestamps and automatic language detection, served from Model Studio (Singapore). Handles files up to 12 hours.

$0.0021/min

WER

12.27%

CER

7.39%

Speed Factor

0.1x

Uptime

100.0%

Billing

per 1s

Retention

No retention

25 languages

(view all)

Auto-detectWord Timestamps
TranscribeView Benchmarks
Rev AI Reverb

Rev AI

Batch
United States

Rev AI Reverb — async speech-to-text built on Rev's Reverb ASR model (trained on 3M+ hours of human-transcribed audio), with speaker diarization for up to 8 speakers and word-level timestamps. English-core with broad language support via the async API.

$0.003/min

WER

18.18%

CER

13.46%

Speed Factor

0.2x

Uptime

100.0%

Billing

per 1s

Retention

No retention

20 languages

(view all)

DiarizationWord Timestamps
TranscribeView Benchmarks
Scribe v2

ElevenLabs

Batch
United States

State-of-the-art batch STT — 90+ languages, speaker diarization, audio tagging

$0.004/min

WER

9.15%

CER

7.19%

Speed Factor

0.2x

Uptime

100.0%

Billing

per 1s

Retention

30 days

76 languages

(view all)

Auto-detectDiarization
CS
+1
TranscribeView Benchmarks
Scribe v2 Realtime

ElevenLabs

Realtime
United States

Most accurate low-latency STT — <150ms, 90+ languages

$0.0065/min

TTFW (P50)

2113 ms

Speed Factor

0.0x

Billing

per 1s

Retention

30 days

76 languages

(view all)

streamingAuto-detect
CS
+2
TranscribeView Benchmarks
Gladia Solaria-1 (Realtime)

Gladia

Realtime
EU (France)

Gladia Solaria-1 — EU-hosted real-time streaming speech-to-text across 100+ languages with native code-switching and ~103ms partial latency. Word-level timestamps and automatic language identification; speaker diarization is available for batch (Solaria-3), not this streaming model.

$0.0125/min

TTFW (P50)

1372 ms

Speed Factor

0.2x

Billing

per 1s

Retention

No retention

25 languages

(view all)

streamingAuto-detectWord Timestamps
TranscribeView Benchmarks
Gladia Solaria-3

Gladia

Batch
EU (France)

Gladia Solaria-3 — EU-hosted batch speech-to-text tuned for the most accurate business audio across five core European languages (English, French, German, Spanish, Italian), with speaker diarization, custom vocabulary, and word-level timestamps.

$0.0101/min
Best Legal

WER

9.80%

CER

5.76%

Speed Factor

0.2x

Uptime

100.0%

Billing

per 1s

Retention

No retention

5 languages

(view all)

Auto-detectDiarizationWord Timestamps+1
TranscribeView Benchmarks
Azure Speech

Microsoft Azure

Batch & Realtime
United States

Azure default STT model — 140+ languages, diarization, word timestamps

$0.006/min

WER

15.34%

CER

10.67%

Speed Factor

0.3x

Billing

per 60s

Retention

Custom

19 languages

(view all)

DiarizationWord Timestamps
TranscribeView Benchmarks
Speechmatics Standard

Speechmatics

Batch
Europe

Cost-effective model — fast turnaround, 55+ languages

$0.005/min

WER

14.69%

CER

9.80%

Speed Factor

0.1x

Uptime

100.0%

Billing

per 1s

Retention

No retention

48 languages

(view all)

Async APIAuto-detectDiarization+3
TranscribeView Benchmarks
Soniox STT (Async)

Soniox

Batch
United States

Soniox async batch transcription — one multilingual model across 60+ languages with speaker diarization, per-word timestamps, and automatic language identification, at a low flat token-based rate.

$0.0017/min

WER

8.73%

CER

6.18%

Speed Factor

0.1x

Uptime

100.0%

Billing

per 1s

Retention

No retention

25 languages

(view all)

Auto-detectDiarizationWord Timestamps
TranscribeView Benchmarks
Soniox STT Realtime

Soniox

Realtime
United States

Soniox realtime streaming transcription — low-latency multilingual STT across 60+ languages with speaker diarization, per-word timestamps, and automatic language identification over a single WebSocket.

$0.002/min

TTFW (P50)

1459 ms

Speed Factor

0.1x

Billing

per 1s

Retention

No retention

25 languages

(view all)

Auto-detectDiarizationWord Timestamps
TranscribeView Benchmarks
Google Telephony

Google Cloud

Batch
United States

Optimized for telephony audio (8kHz)

$0.016/min

WER

20.24%

CER

12.59%

Speed Factor

0.2x

Uptime

100.0%

Billing

per 15s

Retention

Custom

9 languages

(view all)

Word Timestamps
TranscribeView Benchmarks
Amazon Transcribe

Amazon Web Services

Batch
United States

AWS foundation model-powered ASR — 100+ languages

$0.006/min
Most Accurate

WER

13.02%

CER

9.17%

Speed Factor

0.4x

Uptime

100.0%

Billing

per 1s

Retention

Custom

77 languages

(view all)

Async APIAuto-detectDiarization+2
TranscribeView Benchmarks
Amazon Transcribe Medical

Amazon Web Services

Batch
United States

Medical transcription with HIPAA eligibility

$0.075/min

WER

14.92%

CER

9.67%

Speed Factor

0.4x

Uptime

100.0%

Billing

per 1s

Retention

Custom

1 languages

(view all)

HIPAA CompliantAsync APIDiarization+2
TranscribeView Benchmarks
Universal-3.5 Pro Realtime

AssemblyAI

Realtime
United States

AssemblyAI Universal-3.5 Pro Realtime — the flagship streaming model, with native mid-sentence code-switching across 18 languages and conversation context carried across turns. Word-level timestamps and stable partials that are not rewritten as they arrive; speaker diarization is available on the batch model, not this streaming one.

$0.0075/min

TTFW (P50)

1008 ms

Speed Factor

0.1x

Billing

per 1s

Retention

No retention

18 languages

(view all)

streamingAuto-detect
CS
+1
TranscribeView Benchmarks
Universal-3.5 Pro

AssemblyAI

Batch
United States

AssemblyAI's most powerful speech language model — up to 1000 keyterm phrases

$0.0035/min

WER

17.44%

CER

10.37%

Speed Factor

0.2x

Uptime

100.0%

Billing

per 1s

Retention

No retention

6 languages

(view all)

Auto-detectDiarization
CS
+2
TranscribeView Benchmarks
Universal Streaming

AssemblyAI

Realtime
United States

Purpose-built for real-time voice agents — ~300ms immutable transcripts

$0.0025/min

TTFW (P50)

1269 ms

Speed Factor

0.1x

Billing

per 1s

Retention

No retention

1 languages

(view all)

streamingWord Timestamps
TranscribeView Benchmarks
Universal Streaming Multilingual

AssemblyAI

Realtime
United States

Multilingual streaming STT — English, Spanish, French, German, Italian, Portuguese

$0.0025/min

TTFW (P50)

1496 ms

Speed Factor

0.1x

Billing

per 1s

Retention

No retention

6 languages

(view all)

streaming
CS
Word Timestamps
TranscribeView Benchmarks
Voxtral Mini Transcribe

Mistral (Voxtral)

Batch
France

Mistral Voxtral Mini Transcribe — EU-hosted batch transcription with speaker diarization, segment timestamps, and custom vocabulary (context bias). Strong accuracy at very low cost; handles recordings up to 3 hours.

$0.003/min
Best Value
Budget Pick

WER

12.70%

CER

7.97%

Speed Factor

0.2x

Uptime

100.0%

Billing

per 1s

Retention

No retention

8 languages

(view all)

Auto-detectDiarizationWord Timestamps+1
TranscribeView Benchmarks
Whisper 1 (API)

OpenAI

Batch
United States

OpenAI's Whisper API model

$0.006/min

WER

17.63%

CER

13.03%

Speed Factor

0.3x

Uptime

100.0%

Billing

per 1s

Retention

No retention

98 languages

(view all)

Auto-detectWord Timestamps
TranscribeView Benchmarks
Whisper Large V3

OpenAI

Batch
United States

OpenAI's Whisper large-v3 model

$0.006/min
Best Medical

WER

17.67%

CER

13.06%

Speed Factor

0.3x

Uptime

100.0%

Billing

per 1s

Retention

No retention

98 languages

(view all)

Auto-detectWord Timestamps
TranscribeView Benchmarks
Whisper Large V3

Groq

Batch
United States

OpenAI Whisper large-v3 served on Groq's LPU hardware — top Whisper accuracy at Groq speed and cost, with word-level timestamps and automatic language detection.

$0.0019/min

WER

14.41%

CER

11.20%

Speed Factor

0.0x

Billing

per 1s

Retention

No retention

12 languages

(view all)

Auto-detectWord Timestamps
TranscribeView Benchmarks
Whisper Large V3 Turbo

Groq

Batch
United States

OpenAI Whisper large-v3-turbo served on Groq's LPU hardware — very fast, very low cost batch transcription with word-level timestamps and automatic language detection.

$0.0007/min

WER

14.31%

CER

10.63%

Speed Factor

0.0x

Uptime

100.0%

Billing

per 1s

Retention

No retention

12 languages

(view all)

Auto-detectWord Timestamps
TranscribeView Benchmarks
Benchmarks · 120 head-to-head across the 16 most benchmarked models

Compare any two models.

Every pairing we index, grouped by model. Same audio, same scoring.

Chirp 3

Google Cloud

$0.0107/min · 9.95% WER

vs Voxtral Mini Transcribevs Whisper Large V3 (OpenAI)vs Soniox STT (Async)vs Speechmatics Enhancedvs GPT-4o Mini Transcribevs Whisper Large V3 Turbovs Whisper Large V3 (Groq)vs GPT-4o Transcribe Diarizevs Speechmatics Standard
9 pairs

Nova-3

Deepgram

$0.0077/min · 20.98% WER

vs Chirp 3vs Voxtral Mini Transcribevs Gladia Solaria-3vs Whisper Large V3 (OpenAI)vs Soniox STT (Async)vs Scribe v2vs Speechmatics Enhancedvs GPT-4o Mini Transcribevs Whisper Large V3 Turbovs Whisper Large V3 (Groq)vs GPT-4o Transcribe Diarizevs Speechmatics Standard
12 pairs

Gladia Solaria-3

Gladia

$0.0101/min · 9.80% WER

vs Chirp 3vs Voxtral Mini Transcribevs Whisper Large V3 (OpenAI)vs Soniox STT (Async)vs Speechmatics Enhancedvs GPT-4o Mini Transcribevs Whisper Large V3 Turbovs Whisper Large V3 (Groq)vs GPT-4o Transcribe Diarizevs Speechmatics Standard
10 pairs

Amazon Transcribe

Amazon Web Services

$0.006/min · 13.02% WER

vs Chirp 3vs Voxtral Mini Transcribevs Nova-3vs Gladia Solaria-3vs Nova-2vs Whisper Large V3 (OpenAI)vs Soniox STT (Async)vs Scribe v2vs Speechmatics Enhancedvs GPT-4o Mini Transcribevs Whisper Large V3 Turbovs Whisper Large V3 (Groq)vs GPT-4o Transcribe Diarizevs Speechmatics Standard
14 pairs

Nova-2

Deepgram

$0.0058/min · 16.54% WER

vs Chirp 3vs Voxtral Mini Transcribevs Nova-3vs Gladia Solaria-3vs Whisper Large V3 (OpenAI)vs Soniox STT (Async)vs Scribe v2vs Speechmatics Enhancedvs GPT-4o Mini Transcribevs Whisper Large V3 Turbovs Whisper Large V3 (Groq)vs GPT-4o Transcribe Diarizevs Speechmatics Standard
13 pairs

Scribe v2

ElevenLabs

$0.004/min · 9.15% WER

vs Chirp 3vs Voxtral Mini Transcribevs Gladia Solaria-3vs Whisper Large V3 (OpenAI)vs Soniox STT (Async)vs Speechmatics Enhancedvs GPT-4o Mini Transcribevs Whisper Large V3 Turbovs Whisper Large V3 (Groq)vs GPT-4o Transcribe Diarizevs Speechmatics Standard
11 pairs

Qwen3 ASR Flash

Alibaba (Qwen)

$0.0021/min · 12.27% WER

vs Chirp 3vs Voxtral Mini Transcribevs Nova-3vs Gladia Solaria-3vs Amazon Transcribevs Nova-2vs Whisper Large V3 (OpenAI)vs Soniox STT (Async)vs Scribe v2vs Speechmatics Enhancedvs GPT-4o Mini Transcribevs Whisper Large V3 Turbovs Whisper Large V3 (Groq)vs GPT-4o Transcribe Diarizevs Speechmatics Standard
15 pairs

Voxtral Mini Transcribe

Mistral (Voxtral)

$0.003/min · 12.70% WER

vs Whisper Large V3 (OpenAI)vs Soniox STT (Async)vs Speechmatics Enhancedvs GPT-4o Mini Transcribevs GPT-4o Transcribe Diarizevs Speechmatics Standard
6 pairs

Whisper Large V3 Turbo

Groq

$0.0007/min · 14.31% WER

vs Voxtral Mini Transcribevs Whisper Large V3 (OpenAI)vs Soniox STT (Async)vs Speechmatics Enhancedvs GPT-4o Mini Transcribevs GPT-4o Transcribe Diarizevs Speechmatics Standard
7 pairs

Whisper Large V3 (Groq)

Groq

$0.0019/min · 14.41% WER

vs Voxtral Mini Transcribevs Whisper Large V3 (OpenAI)vs Soniox STT (Async)vs Speechmatics Enhancedvs GPT-4o Mini Transcribevs Whisper Large V3 Turbovs GPT-4o Transcribe Diarizevs Speechmatics Standard
8 pairs

Whisper Large V3 (OpenAI)

OpenAI

$0.006/min · 17.67% WER

vs Soniox STT (Async)vs Speechmatics Enhancedvs Speechmatics Standard
3 pairs

GPT-4o Mini Transcribe

OpenAI

$0.003/min · 13.58% WER

vs Whisper Large V3 (OpenAI)vs Soniox STT (Async)vs Speechmatics Enhancedvs GPT-4o Transcribe Diarizevs Speechmatics Standard
5 pairs

GPT-4o Transcribe Diarize

OpenAI

$0.006/min · 14.61% WER

vs Whisper Large V3 (OpenAI)vs Soniox STT (Async)vs Speechmatics Enhancedvs Speechmatics Standard
4 pairs

Soniox STT (Async)

Soniox

$0.0017/min · 8.73% WER

vs Speechmatics Enhancedvs Speechmatics Standard
2 pairs

Speechmatics Enhanced

Speechmatics

$0.0083/min · 12.30% WER

vs Speechmatics Standard
1 pair
From the blog

Deep dives on the models we index.

Accuracy, latency and pricing, checked against what providers actually ship.

All posts

Amazon Transcribe Medical: what AWS actually ships, and what it won't tell you

What Amazon Transcribe Medical offers in 2026: features, specs, pricing vs Google and Nuance, HIPAA posture, research clues, and where it falls short.

July 3, 2026

Read

Chirp 3: inside Google Cloud's 2025 speech stack, from HD voices to transcription

What Google Cloud Chirp 3 actually is: release timeline, WER and Elo benchmarks, pricing, specs, known issues, and how it compares to Azure and ElevenLabs.

July 3, 2026

Read

Deepgram Base in 2026: what the legacy model still does well

Where Deepgram Base fits in 2026: API behavior, variants, latency, specs, limitations, and when to pick Nova-3 or Flux instead.

July 3, 2026

Read