OpenTranscription
OpenTranscription
RankerModelsPlayground
OpenTranscription
OpenTranscription

One API to every speech-to-text model worth using. Compare them on your audio, route to the best one, pay per second.

Platform status

Product

RankerModelsTranscriptionsPlaygroundBlog

Developers

DocumentationReliabilityAPI VersioningStatus

Legal

Privacy PolicyTerms of ServiceSupport
© 2026 OpenTranscription

Catalog · 43 models · 15 providers

Transcription models.

Every speech-to-text model worth using, one rate card.

Price Range

$0.0007

$0.075

Compliance

Capabilities

Best For:

43

Total models

105

Languages

$0.0007

Cheapest /min

8.7%

Best WER

43 models

Base

Deepgram

Batch
United States

Deepgram Base model — high volume, cost-effective

$0.0145/min

WER

25.04%

CER

18.09%

Speed Factor

0.1x

Uptime

100.0%

Billing

per 1s

Retention

No retention

19 languages

(view all)

Auto-detectDiarizationWord Timestamps+1
TranscribeView Benchmarks
Chirp 3

Google Cloud

Batch
United States

Google's latest generative ASR foundation model — 85+ languages

$0.0107/min
Best Conversations
Best Accented

WER

9.95%

CER

5.81%

Speed Factor

0.3x

Uptime

66.7%

Billing

per 15s

Retention

Custom

68 languages

(view all)

GenerativeAuto-detectDiarization+2
TranscribeView Benchmarks
Google Command & Search

Google Cloud

Batch
United States

Optimized for short queries and voice commands

$0.016/min

WER

35.22%

CER

26.46%

Speed Factor

0.1x

Billing

per 15s

Retention

Custom

40 languages

(view all)

Word Timestamps
TranscribeView Benchmarks
Google Default

Google Cloud

Batch
United States

General-purpose model

$0.016/min

WER

34.65%

CER

25.66%

Speed Factor

0.2x

Billing

per 15s

Retention

Custom

40 languages

(view all)

DiarizationWord Timestamps
TranscribeView Benchmarks
Speechmatics Enhanced

Speechmatics

Batch
Europe

Highest accuracy model — 55+ languages, best-in-class

$0.0083/min

WER

12.30%

CER

7.96%

Speed Factor

0.3x

Uptime

100.0%

Billing

per 1s

Retention

No retention

48 languages

(view all)

Async APIAuto-detectDiarization+3
TranscribeView Benchmarks
Enhanced

Deepgram

Batch
United States

Deepgram Enhanced model — high accuracy for uncommon words

$0.0165/min

WER

24.86%

CER

14.27%

Speed Factor

0.1x

Uptime

100.0%

Billing

per 1s

Retention

No retention

15 languages

(view all)

Auto-detectDiarizationWord Timestamps+1
TranscribeView Benchmarks
Flux

Deepgram

Realtime
United States

First conversational ASR model built for voice agents — model-integrated endpointing

$0.0077/min

TTFW (P50)

289 ms

Speed Factor

0.1x

Uptime

85.7%

Billing

per 1s

Retention

No retention

1 languages

(view all)

endpointing
CS
turn detection+1
TranscribeView Benchmarks
GPT-4o Mini Transcribe

OpenAI

Batch
United States

GPT-4o Mini optimized for fast transcription

$0.003/min

WER

13.58%

CER

9.32%

Speed Factor

0.1x

Uptime

100.0%

Billing

per 1s

Retention

No retention

98 languages

(view all)

Auto-detectWord Timestamps
TranscribeView Benchmarks
GPT-4o Transcribe

OpenAI

Batch
United States

GPT-4o optimized for transcription with improved WER

$0.006/min

WER

14.76%

CER

11.60%

Speed Factor

0.2x

Uptime

100.0%

Billing

per 1s

Retention

No retention

98 languages

(view all)

Auto-detectWord Timestamps
TranscribeView Benchmarks
GPT-4o Transcribe Diarize

OpenAI

Batch
United States

GPT-4o transcription with built-in speaker diarization

$0.006/min

WER

14.61%

CER

9.97%

Speed Factor

0.2x

Uptime

100.0%

Billing

per 1s

Retention

No retention

12 languages

(view all)

Auto-detectDiarization
TranscribeView Benchmarks
Ink-2

Cartesia

Realtime
United States

Cartesia's newest streaming STT — lowest WER of any streaming model, native turn detection, and robust on alphanumerics like phone numbers, emails, and UUIDs. English only for now.

$0.0065/min

TTFW (P50)

1554 ms

Speed Factor

0.1x

Billing

per 1s

Retention

No retention

1 languages

(view all)

streamingendpointingturn detection+1
TranscribeView Benchmarks
Ink-Whisper

Cartesia

Batch & Realtime
United States

Whisper rearchitected for real-time and batch voice AI — fastest TTCT, 99-language coverage

$0.0022/min

WER

17.20%

CER

11.58%

Speed Factor

0.1x

Uptime

100.0%

Billing

per 1s

Retention

No retention

100 languages

(view all)

Word Timestampsdynamic chunking
TranscribeView Benchmarks
Google Latest (Long)

Google Cloud

Batch
United States

Conformer model for long-form audio (minutes to hours)

$0.0107/min

WER

22.55%

CER

13.81%

Speed Factor

0.4x

Uptime

100.0%

Billing

per 15s

Retention

Custom

40 languages

(view all)

Word Timestamps
TranscribeView Benchmarks
Google Latest (Short)

Google Cloud

Batch
United States

Conformer model for short utterances (< 60s)

$0.016/min

WER

57.82%

CER

52.44%

Speed Factor

0.1x

Billing

per 15s

Retention

Custom

40 languages

(view all)

Word Timestamps
TranscribeView Benchmarks
Nova-2

Deepgram

Batch & Realtime
United States

Deepgram's Nova-2 speech recognition

$0.0058/min
Fastest

WER

16.54%

CER

11.93%

Speed Factor

0.1x

Uptime

83.3%

Billing

per 1s

Retention

No retention

33 languages

(view all)

Auto-detectDiarization
CS
+2
TranscribeView Benchmarks
Nova-2 Conversational AI

Deepgram

Batch & Realtime
United States

Optimized for human-to-bot interactions (IVR, voice assistants)

$0.0058/min

WER

18.23%

CER

13.27%

Speed Factor

0.1x

Billing

per 1s

Retention

No retention

1 languages

(view all)

Diarization
CS
Word Timestamps+1
TranscribeView Benchmarks
Nova-2 Finance

Deepgram

Batch & Realtime
Deprecated
United States

Optimized for earnings calls with finance vocabulary

$0.0058/min

WER

18.92%

CER

13.09%

Speed Factor

0.1x

Billing

per 1s

Retention

No retention

1 languages

(view all)

Diarization
CS
Word Timestamps+1
TranscribeView Benchmarks
Nova-2 Meeting

Deepgram

Batch & Realtime
United States

Optimized for conference room audio

$0.0058/min

WER

16.74%

CER

11.73%

Speed Factor

0.1x

Billing

per 1s

Retention

No retention

1 languages

(view all)

Diarization
CS
Word Timestamps+1
TranscribeView Benchmarks
Nova-2 Phone Call

Deepgram

Batch & Realtime
United States

Optimized for low-bandwidth phone call audio

$0.0058/min

WER

16.38%

CER

12.42%

Speed Factor

0.1x

Billing

per 1s

Retention

No retention

1 languages

(view all)

Diarization
CS
Word Timestamps+1
TranscribeView Benchmarks
Nova-2 Voicemail

Deepgram

Batch & Realtime
United States

Optimized for low-bandwidth single speaker voicemail

$0.0058/min

WER

17.51%

CER

12.36%

Speed Factor

0.1x

Billing

per 1s

Retention

No retention

1 languages

(view all)

CS
Word TimestampsCustom Vocabulary
TranscribeView Benchmarks
Nova-3

Deepgram

Batch & Realtime
United States

Deepgram's flagship model — 53% lower WER vs competitors, code-switching support

$0.0077/min
Best Noisy
Best Technical

WER

20.98%

CER

11.48%

Speed Factor

0.1x

Uptime

50.0%

Billing

per 1s

Retention

No retention

47 languages

(view all)

Auto-detectDiarizationSmart Format+3
TranscribeView Benchmarks
Qwen3 ASR Flash

Alibaba (Qwen)

Batch
Singapore

Alibaba Qwen3-ASR Flash — multilingual batch transcription (25+ languages) with word-level timestamps and automatic language detection, served from Model Studio (Singapore). Handles files up to 12 hours.

$0.0021/min

WER

12.27%

CER

7.39%

Speed Factor

0.1x

Uptime

100.0%

Billing

per 1s

Retention

No retention

25 languages

(view all)

Auto-detectWord Timestamps
TranscribeView Benchmarks
Rev AI Reverb

Rev AI

Batch
United States

Rev AI Reverb — async speech-to-text built on Rev's Reverb ASR model (trained on 3M+ hours of human-transcribed audio), with speaker diarization for up to 8 speakers and word-level timestamps. English-core with broad language support via the async API.

$0.003/min

WER

18.18%

CER

13.46%

Speed Factor

0.2x

Uptime

100.0%

Billing

per 1s

Retention

No retention

20 languages

(view all)

DiarizationWord Timestamps
TranscribeView Benchmarks
Scribe v2

ElevenLabs

Batch
United States

State-of-the-art batch STT — 90+ languages, speaker diarization, audio tagging

$0.004/min

WER

9.15%

CER

7.19%

Speed Factor

0.2x

Uptime

100.0%

Billing

per 1s

Retention

30 days

76 languages

(view all)

Auto-detectDiarization
CS
+1
TranscribeView Benchmarks
Scribe v2 Realtime

ElevenLabs

Realtime
United States

Most accurate low-latency STT — <150ms, 90+ languages

$0.0065/min

TTFW (P50)

2113 ms

Speed Factor

0.0x

Billing

per 1s

Retention

30 days

76 languages

(view all)

streamingAuto-detect
CS
+2
TranscribeView Benchmarks
Gladia Solaria-1 (Realtime)

Gladia

Realtime
EU (France)

Gladia Solaria-1 — EU-hosted real-time streaming speech-to-text across 100+ languages with native code-switching and ~103ms partial latency. Word-level timestamps and automatic language identification; speaker diarization is available for batch (Solaria-3), not this streaming model.

$0.0125/min

TTFW (P50)

1372 ms

Speed Factor

0.2x

Billing

per 1s

Retention

No retention

25 languages

(view all)

streamingAuto-detectWord Timestamps
TranscribeView Benchmarks
Gladia Solaria-3

Gladia

Batch
EU (France)

Gladia Solaria-3 — EU-hosted batch speech-to-text tuned for the most accurate business audio across five core European languages (English, French, German, Spanish, Italian), with speaker diarization, custom vocabulary, and word-level timestamps.

$0.0101/min
Best Legal

WER

9.80%

CER

5.76%

Speed Factor

0.2x

Uptime

100.0%

Billing

per 1s

Retention

No retention

5 languages

(view all)

Auto-detectDiarizationWord Timestamps+1
TranscribeView Benchmarks
Azure Speech

Microsoft Azure

Batch & Realtime
United States

Azure default STT model — 140+ languages, diarization, word timestamps

$0.006/min

WER

15.34%

CER

10.67%

Speed Factor

0.3x

Billing

per 60s

Retention

Custom

19 languages

(view all)

DiarizationWord Timestamps
TranscribeView Benchmarks
Speechmatics Standard

Speechmatics

Batch
Europe

Cost-effective model — fast turnaround, 55+ languages

$0.005/min

WER

14.69%

CER

9.80%

Speed Factor

0.1x

Uptime

100.0%

Billing

per 1s

Retention

No retention

48 languages

(view all)

Async APIAuto-detectDiarization+3
TranscribeView Benchmarks
Soniox STT (Async)

Soniox

Batch
United States

Soniox async batch transcription — one multilingual model across 60+ languages with speaker diarization, per-word timestamps, and automatic language identification, at a low flat token-based rate.

$0.0017/min

WER

8.73%

CER

6.18%

Speed Factor

0.1x

Uptime

100.0%

Billing

per 1s

Retention

No retention

25 languages

(view all)

Auto-detectDiarizationWord Timestamps
TranscribeView Benchmarks
Soniox STT Realtime

Soniox

Realtime
United States

Soniox realtime streaming transcription — low-latency multilingual STT across 60+ languages with speaker diarization, per-word timestamps, and automatic language identification over a single WebSocket.

$0.002/min

TTFW (P50)

1459 ms

Speed Factor

0.1x

Billing

per 1s

Retention

No retention

25 languages

(view all)

Auto-detectDiarizationWord Timestamps
TranscribeView Benchmarks
Google Telephony

Google Cloud

Batch
United States

Optimized for telephony audio (8kHz)

$0.016/min

WER

20.24%

CER

12.59%

Speed Factor

0.2x

Uptime

100.0%

Billing

per 15s

Retention

Custom

9 languages

(view all)

Word Timestamps
TranscribeView Benchmarks
Amazon Transcribe

Amazon Web Services

Batch
United States

AWS foundation model-powered ASR — 100+ languages

$0.006/min
Most Accurate

WER

13.02%

CER

9.17%

Speed Factor

0.4x

Uptime

100.0%

Billing

per 1s

Retention

Custom

77 languages

(view all)

Async APIAuto-detectDiarization+2
TranscribeView Benchmarks
Amazon Transcribe Medical

Amazon Web Services

Batch
United States

Medical transcription with HIPAA eligibility

$0.075/min

WER

14.92%

CER

9.67%

Speed Factor

0.4x

Uptime

100.0%

Billing

per 1s

Retention

Custom

1 languages

(view all)

HIPAA CompliantAsync APIDiarization+2
TranscribeView Benchmarks
Universal-3.5 Pro Realtime

AssemblyAI

Realtime
United States

AssemblyAI Universal-3.5 Pro Realtime — the flagship streaming model, with native mid-sentence code-switching across 18 languages and conversation context carried across turns. Word-level timestamps and stable partials that are not rewritten as they arrive; speaker diarization is available on the batch model, not this streaming one.

$0.0075/min

TTFW (P50)

1008 ms

Speed Factor

0.1x

Billing

per 1s

Retention

No retention

18 languages

(view all)

streamingAuto-detect
CS
+1
TranscribeView Benchmarks
Universal-3.5 Pro

AssemblyAI

Batch
United States

AssemblyAI's most powerful speech language model — up to 1000 keyterm phrases

$0.0035/min

WER

17.44%

CER

10.37%

Speed Factor

0.2x

Uptime

88.9%

Billing

per 1s

Retention

No retention

6 languages

(view all)

Auto-detectDiarization
CS
+2
TranscribeView Benchmarks
Universal Streaming

AssemblyAI

Realtime
United States

Purpose-built for real-time voice agents — ~300ms immutable transcripts

$0.0025/min

TTFW (P50)

1269 ms

Speed Factor

0.1x

Billing

per 1s

Retention

No retention

1 languages

(view all)

streamingWord Timestamps
TranscribeView Benchmarks
Universal Streaming Multilingual

AssemblyAI

Realtime
United States

Multilingual streaming STT — English, Spanish, French, German, Italian, Portuguese

$0.0025/min

TTFW (P50)

1496 ms

Speed Factor

0.1x

Billing

per 1s

Retention

No retention

6 languages

(view all)

streaming
CS
Word Timestamps
TranscribeView Benchmarks
Voxtral Mini Transcribe

Mistral (Voxtral)

Batch
France

Mistral Voxtral Mini Transcribe — EU-hosted batch transcription with speaker diarization, segment timestamps, and custom vocabulary (context bias). Strong accuracy at very low cost; handles recordings up to 3 hours.

$0.003/min
Best Value
Budget Pick

WER

12.70%

CER

7.97%

Speed Factor

0.2x

Uptime

100.0%

Billing

per 1s

Retention

No retention

8 languages

(view all)

Auto-detectDiarizationCustom Vocabulary
TranscribeView Benchmarks
Whisper 1 (API)

OpenAI

Batch
United States

OpenAI's Whisper API model

$0.006/min

WER

17.63%

CER

13.03%

Speed Factor

0.3x

Uptime

100.0%

Billing

per 1s

Retention

No retention

98 languages

(view all)

Auto-detectWord Timestamps
TranscribeView Benchmarks
Whisper Large V3

Groq

Batch
United States

OpenAI Whisper large-v3 served on Groq's LPU hardware — top Whisper accuracy at Groq speed and cost, with word-level timestamps and automatic language detection.

$0.0019/min

WER

14.41%

CER

11.20%

Speed Factor

0.0x

Billing

per 1s

Retention

No retention

12 languages

(view all)

Auto-detectWord Timestamps
TranscribeView Benchmarks
Whisper Large V3

OpenAI

Batch
United States

OpenAI's Whisper large-v3 model

$0.006/min
Best Medical

WER

17.67%

CER

13.06%

Speed Factor

0.3x

Uptime

100.0%

Billing

per 1s

Retention

No retention

98 languages

(view all)

Auto-detectWord Timestamps
TranscribeView Benchmarks
Whisper Large V3 Turbo

Groq

Batch
United States

OpenAI Whisper large-v3-turbo served on Groq's LPU hardware — very fast, very low cost batch transcription with word-level timestamps and automatic language detection.

$0.0007/min

WER

14.31%

CER

10.63%

Speed Factor

0.0x

Uptime

100.0%

Billing

per 1s

Retention

No retention

12 languages

(view all)

Auto-detectWord Timestamps
TranscribeView Benchmarks

All model comparisons

Every head-to-head we index, grouped by model.

Chirp 3

vs Voxtral Mini Transcribevs Whisper Large V3 (OpenAI)vs Soniox STT (Async)vs Speechmatics Enhancedvs GPT-4o Mini Transcribevs Whisper Large V3 Turbovs Whisper Large V3 (Groq)vs GPT-4o Transcribe Diarizevs Speechmatics Standard

Nova-3

vs Chirp 3vs Voxtral Mini Transcribevs Gladia Solaria-3vs Whisper Large V3 (OpenAI)vs Soniox STT (Async)vs Scribe v2vs Speechmatics Enhancedvs GPT-4o Mini Transcribevs Whisper Large V3 Turbovs Whisper Large V3 (Groq)vs GPT-4o Transcribe Diarizevs Speechmatics Standard

Gladia Solaria-3

vs Chirp 3vs Voxtral Mini Transcribevs Whisper Large V3 (OpenAI)vs Soniox STT (Async)vs Speechmatics Enhancedvs GPT-4o Mini Transcribevs Whisper Large V3 Turbovs Whisper Large V3 (Groq)vs GPT-4o Transcribe Diarizevs Speechmatics Standard

Amazon Transcribe

vs Chirp 3vs Voxtral Mini Transcribevs Nova-3vs Gladia Solaria-3vs Nova-2vs Whisper Large V3 (OpenAI)vs Soniox STT (Async)vs Scribe v2vs Speechmatics Enhancedvs GPT-4o Mini Transcribevs Whisper Large V3 Turbovs Whisper Large V3 (Groq)vs GPT-4o Transcribe Diarizevs Speechmatics Standard

Nova-2

vs Chirp 3vs Voxtral Mini Transcribevs Nova-3vs Gladia Solaria-3vs Whisper Large V3 (OpenAI)vs Soniox STT (Async)vs Scribe v2vs Speechmatics Enhancedvs GPT-4o Mini Transcribevs Whisper Large V3 Turbovs Whisper Large V3 (Groq)vs GPT-4o Transcribe Diarizevs Speechmatics Standard

Scribe v2

vs Chirp 3vs Voxtral Mini Transcribevs Gladia Solaria-3vs Whisper Large V3 (OpenAI)vs Soniox STT (Async)vs Speechmatics Enhancedvs GPT-4o Mini Transcribevs Whisper Large V3 Turbovs Whisper Large V3 (Groq)vs GPT-4o Transcribe Diarizevs Speechmatics Standard

Qwen3 ASR Flash

vs Chirp 3vs Voxtral Mini Transcribevs Nova-3vs Gladia Solaria-3vs Amazon Transcribevs Nova-2vs Whisper Large V3 (OpenAI)vs Soniox STT (Async)vs Scribe v2vs Speechmatics Enhancedvs GPT-4o Mini Transcribevs Whisper Large V3 Turbovs Whisper Large V3 (Groq)vs GPT-4o Transcribe Diarizevs Speechmatics Standard

Voxtral Mini Transcribe

vs Whisper Large V3 (OpenAI)vs Soniox STT (Async)vs Speechmatics Enhancedvs GPT-4o Mini Transcribevs GPT-4o Transcribe Diarizevs Speechmatics Standard

Whisper Large V3 Turbo

vs Voxtral Mini Transcribevs Whisper Large V3 (OpenAI)vs Soniox STT (Async)vs Speechmatics Enhancedvs GPT-4o Mini Transcribevs GPT-4o Transcribe Diarizevs Speechmatics Standard

Whisper Large V3 (Groq)

vs Voxtral Mini Transcribevs Whisper Large V3 (OpenAI)vs Soniox STT (Async)vs Speechmatics Enhancedvs GPT-4o Mini Transcribevs Whisper Large V3 Turbovs GPT-4o Transcribe Diarizevs Speechmatics Standard

Whisper Large V3 (OpenAI)

vs Soniox STT (Async)vs Speechmatics Enhancedvs Speechmatics Standard

GPT-4o Mini Transcribe

vs Whisper Large V3 (OpenAI)vs Soniox STT (Async)vs Speechmatics Enhancedvs GPT-4o Transcribe Diarizevs Speechmatics Standard

GPT-4o Transcribe Diarize

vs Whisper Large V3 (OpenAI)vs Soniox STT (Async)vs Speechmatics Enhancedvs Speechmatics Standard

Soniox STT (Async)

vs Speechmatics Enhancedvs Speechmatics Standard

Speechmatics Enhanced

vs Speechmatics Standard

From the blog

Deep dives on speech-to-text accuracy, latency and pricing.

All posts
Amazon Transcribe Medical: what AWS actually ships, and what it won't tell you

What Amazon Transcribe Medical offers in 2026: features, specs, pricing vs Google and Nuance, HIPAA posture, research clues, and where it falls short.

Chirp 3: inside Google Cloud's 2025 speech stack, from HD voices to transcription

What Google Cloud Chirp 3 actually is: release timeline, WER and Elo benchmarks, pricing, specs, known issues, and how it compares to Azure and ElevenLabs.

Deepgram Base in 2026: what the legacy model still does well

Where Deepgram Base fits in 2026: API behavior, variants, latency, specs, limitations, and when to pick Nova-3 or Flux instead.