Price Range

$0.0007

$0.09

Compliance

Capabilities

Best For:

Catalog · 40 models · 13 providers

Transcription models.

Every speech-to-text model worth using, one rate card.

40

Total models

106

Languages

$0.0007

Cheapest /min

4.7%

Best WER

40 models

Base

Deepgram

Batch
United States

Deepgram Base model — high volume, cost-effective

$0.0145/min

WER

42.78%

CER

47.68%

Speed Factor

0.1x

Uptime

100.0%

Billing

per 1s

Retention

No retention

19 languages

(view all)

Auto-detectDiarizationWord Timestamps+1
TranscribeView Benchmarks
AssemblyAI Best

AssemblyAI

Batch
United States

AssemblyAI's best transcription model

$0.09/min

WER

20.48%

CER

11.80%

Speed Factor

0.3x

Uptime

100.0%

Billing

per 1s

Retention

No retention

78 languages

(view all)

Auto-detectDiarizationWord Timestamps+1
TranscribeView Benchmarks
Chirp 3

Google Cloud

Batch
United States

Google's latest generative ASR foundation model — 85+ languages

$0.0107/min

WER

9.37%

CER

4.79%

Speed Factor

0.3x

Uptime

100.0%

Billing

per 15s

Retention

Custom

72 languages

(view all)

GenerativeAuto-detectDiarization+2
TranscribeView Benchmarks
Google Command & Search

Google Cloud

Batch
United States

Optimized for short queries and voice commands

$0.016/min

WER

47.19%

CER

21.87%

Speed Factor

0.1x

Billing

per 15s

Retention

Custom

40 languages

(view all)

Word Timestamps
TranscribeView Benchmarks
Google Default

Google Cloud

Batch
United States

General-purpose model

$0.016/min

WER

45.74%

CER

36.56%

Speed Factor

0.2x

Billing

per 15s

Retention

Custom

40 languages

(view all)

DiarizationWord Timestamps
TranscribeView Benchmarks
Enhanced

Deepgram

Batch
United States

Deepgram Enhanced model — high accuracy for uncommon words

$0.0165/min

WER

25.24%

CER

15.53%

Speed Factor

0.1x

Uptime

100.0%

Billing

per 1s

Retention

No retention

15 languages

(view all)

Auto-detectDiarizationWord Timestamps+1
TranscribeView Benchmarks
Speechmatics Enhanced

Speechmatics

Batch
Europe

Highest accuracy model — 55+ languages, best-in-class

$0.0083/min
Best Accented

WER

16.87%

CER

9.29%

Speed Factor

0.3x

Uptime

100.0%

Billing

per 1s

Retention

No retention

48 languages

(view all)

Async APIAuto-detectDiarization+3
TranscribeView Benchmarks
Flux

Deepgram

Realtime
United States

First conversational ASR model built for voice agents — model-integrated endpointing

$0.0077/min

TTFW (P50)

293 ms

Speed Factor

0.1x

Billing

per 1s

Retention

No retention

1 languages

(view all)

endpointing
CS
turn detection+1
TranscribeView Benchmarks
GPT-4o Mini Transcribe

OpenAI

Batch
United States

GPT-4o Mini optimized for fast transcription

$0.003/min

Speed Factor

0.1x

Uptime

100.0%

Billing

per 1s

Retention

No retention

98 languages

(view all)

Auto-detectWord Timestamps
TranscribeView Benchmarks
GPT-4o Transcribe

OpenAI

Batch
United States

GPT-4o optimized for transcription with improved WER

$0.006/min

Speed Factor

0.2x

Uptime

100.0%

Billing

per 1s

Retention

No retention

98 languages

(view all)

Auto-detectWord Timestamps
TranscribeView Benchmarks
GPT-4o Transcribe Diarize

OpenAI

Batch
United States

GPT-4o transcription with built-in speaker diarization

$0.006/min

WER

14.90%

CER

8.59%

Speed Factor

0.2x

Billing

per 1s

Retention

No retention

12 languages

(view all)

Auto-detectDiarization
TranscribeView Benchmarks
Ink-2

Cartesia

Realtime
United States

Cartesia's newest streaming STT — lowest WER of any streaming model, native turn detection, and robust on alphanumerics like phone numbers, emails, and UUIDs. English only for now.

$0.0065/min

TTFW (P50)

1418 ms

Speed Factor

0.1x

Billing

per 1s

Retention

No retention

1 languages

(view all)

streamingendpointingturn detection+1
TranscribeView Benchmarks
Ink-Whisper

Cartesia

Batch & Realtime
United States

Whisper rearchitected for real-time and batch voice AI — fastest TTCT, 99-language coverage

$0.0022/min

WER

23.08%

CER

11.38%

Speed Factor

0.1x

Uptime

100.0%

Billing

per 1s

Retention

No retention

100 languages

(view all)

Word Timestampsdynamic chunking
TranscribeView Benchmarks
Google Latest (Long)

Google Cloud

Batch
United States

Conformer model for long-form audio (minutes to hours)

$0.0107/min

WER

22.54%

CER

10.26%

Speed Factor

0.4x

Billing

per 15s

Retention

Custom

40 languages

(view all)

Word Timestamps
TranscribeView Benchmarks
Google Latest (Short)

Google Cloud

Batch
United States

Conformer model for short utterances (< 60s)

$0.016/min

WER

63.82%

CER

57.52%

Speed Factor

0.1x

Billing

per 15s

Retention

Custom

40 languages

(view all)

Word Timestamps
TranscribeView Benchmarks
Nova-2

Deepgram

Batch & Realtime
United States

Deepgram's Nova-2 speech recognition

$0.0058/min

WER

23.62%

CER

20.01%

Speed Factor

0.1x

Uptime

100.0%

Billing

per 1s

Retention

No retention

33 languages

(view all)

Auto-detectDiarization
CS
+2
TranscribeView Benchmarks
Nova-2 Conversational AI

Deepgram

Batch & Realtime
United States

Optimized for human-to-bot interactions (IVR, voice assistants)

$0.0058/min

WER

19.02%

CER

13.21%

Speed Factor

0.1x

Billing

per 1s

Retention

No retention

1 languages

(view all)

Diarization
CS
Word Timestamps+1
TranscribeView Benchmarks
Nova-2 Finance

Deepgram

Batch & Realtime
United States

Optimized for earnings calls with finance vocabulary

$0.0058/min

WER

18.58%

CER

13.25%

Speed Factor

0.1x

Billing

per 1s

Retention

No retention

1 languages

(view all)

Diarization
CS
Word Timestamps+1
TranscribeView Benchmarks
Nova-2 Meeting

Deepgram

Batch & Realtime
United States

Optimized for conference room audio

$0.0058/min

WER

17.99%

CER

13.41%

Speed Factor

0.1x

Billing

per 1s

Retention

No retention

1 languages

(view all)

Diarization
CS
Word Timestamps+1
TranscribeView Benchmarks
Nova-2 Phone Call

Deepgram

Batch & Realtime
United States

Optimized for low-bandwidth phone call audio

$0.0058/min

WER

16.48%

CER

11.74%

Speed Factor

0.1x

Billing

per 1s

Retention

No retention

1 languages

(view all)

Diarization
CS
Word Timestamps+1
TranscribeView Benchmarks
Nova-2 Voicemail

Deepgram

Batch & Realtime
United States

Optimized for low-bandwidth single speaker voicemail

$0.0058/min

WER

17.33%

CER

12.48%

Speed Factor

0.1x

Billing

per 1s

Retention

No retention

1 languages

(view all)

CS
Word TimestampsCustom Vocabulary
TranscribeView Benchmarks
Nova-3

Deepgram

Batch & Realtime
United States

Deepgram's flagship model — 53% lower WER vs competitors, code-switching support

$0.0077/min

WER

30.75%

CER

35.97%

Speed Factor

0.1x

Uptime

100.0%

Billing

per 1s

Retention

No retention

47 languages

(view all)

Auto-detectDiarizationSmart Format+3
TranscribeView Benchmarks
Qwen3 ASR Flash

Alibaba (Qwen)

Batch
Singapore

Alibaba Qwen3-ASR Flash — multilingual batch transcription (25+ languages) with word-level timestamps and automatic language detection, served from Model Studio (Singapore). Handles files up to 12 hours.

$0.0021/min

Speed Factor

0.1x

Uptime

100.0%

Billing

per 1s

Retention

No retention

25 languages

(view all)

Auto-detectWord Timestamps
TranscribeView Benchmarks
Rev AI Reverb

Rev AI

Batch
United States

Rev AI Reverb — async speech-to-text built on Rev's Reverb ASR model (trained on 3M+ hours of human-transcribed audio), with speaker diarization for up to 8 speakers and word-level timestamps. English-core with broad language support via the async API.

$0.003/min
Best Conversations

WER

18.67%

CER

14.93%

Speed Factor

0.2x

Billing

per 1s

Retention

No retention

20 languages

(view all)

DiarizationWord Timestamps
TranscribeView Benchmarks
Scribe v2

ElevenLabs

Batch
United States

State-of-the-art batch STT — 90+ languages, speaker diarization, audio tagging

$0.004/min

Speed Factor

0.2x

Uptime

100.0%

Billing

per 1s

Retention

30 days

76 languages

(view all)

Auto-detectDiarization
CS
+1
TranscribeView Benchmarks
Scribe v2 Realtime

ElevenLabs

Realtime
United States

Most accurate low-latency STT — <150ms, 90+ languages

$0.0065/min

TTFW (P50)

2084 ms

Speed Factor

0.0x

Billing

per 1s

Retention

30 days

76 languages

(view all)

streamingAuto-detect
CS
+2
TranscribeView Benchmarks
Gladia Solaria-1 (Realtime)

Gladia

Realtime
EU (France)

Gladia Solaria-1 — EU-hosted real-time streaming speech-to-text across 100+ languages with native code-switching and ~103ms partial latency. Word-level timestamps and automatic language identification; speaker diarization is available for batch (Solaria-3), not this streaming model.

$0.0125/min

TTFW (P50)

1332 ms

Speed Factor

0.2x

Billing

per 1s

Retention

No retention

25 languages

(view all)

streamingAuto-detectWord Timestamps
TranscribeView Benchmarks
Gladia Solaria-3

Gladia

Batch
EU (France)

Gladia Solaria-3 — EU-hosted batch speech-to-text tuned for the most accurate business audio across five core European languages (English, French, German, Spanish, Italian), with speaker diarization, custom vocabulary, and word-level timestamps.

$0.0101/min
Best Legal

WER

10.13%

CER

6.26%

Speed Factor

0.2x

Billing

per 1s

Retention

No retention

5 languages

(view all)

Auto-detectDiarizationWord Timestamps+1
TranscribeView Benchmarks
Speechmatics Standard

Speechmatics

Batch
Europe

Cost-effective model — fast turnaround, 55+ languages

$0.005/min

WER

20.05%

CER

12.68%

Speed Factor

0.1x

Uptime

100.0%

Billing

per 1s

Retention

No retention

48 languages

(view all)

Async APIAuto-detectDiarization+3
TranscribeView Benchmarks
Google Telephony

Google Cloud

Batch
United States

Optimized for telephony audio (8kHz)

$0.016/min

WER

22.85%

CER

13.22%

Speed Factor

0.2x

Uptime

100.0%

Billing

per 15s

Retention

Custom

9 languages

(view all)

Word Timestamps
TranscribeView Benchmarks
Amazon Transcribe

Amazon Web Services

Batch
United States

AWS foundation model-powered ASR — 100+ languages

$0.006/min
Most Accurate

WER

4.71%

CER

2.53%

Speed Factor

0.4x

Uptime

100.0%

Billing

per 1s

Retention

Custom

77 languages

(view all)

Async APIAuto-detectDiarization+2
TranscribeView Benchmarks
Amazon Transcribe Medical

Amazon Web Services

Batch
United States

Medical transcription with HIPAA eligibility

$0.075/min
Best Noisy
Best Technical

WER

25.73%

CER

17.32%

Speed Factor

0.4x

Uptime

100.0%

Billing

per 1s

Retention

Custom

1 languages

(view all)

HIPAA CompliantAsync APIDiarization+2
TranscribeView Benchmarks
Universal-3 Pro

AssemblyAI

Batch
United States

AssemblyAI's most powerful speech language model — up to 1000 keyterm phrases

$0.0035/min

WER

17.30%

CER

9.93%

Speed Factor

0.2x

Uptime

100.0%

Billing

per 1s

Retention

No retention

6 languages

(view all)

Auto-detectDiarization
CS
+2
TranscribeView Benchmarks
Universal Streaming

AssemblyAI

Realtime
United States

Purpose-built for real-time voice agents — ~300ms immutable transcripts

$0.0025/min

TTFW (P50)

1277 ms

Speed Factor

0.1x

Billing

per 1s

Retention

No retention

1 languages

(view all)

streamingWord Timestamps
TranscribeView Benchmarks
Universal Streaming Multilingual

AssemblyAI

Realtime
United States

Multilingual streaming STT — English, Spanish, French, German, Italian, Portuguese

$0.0025/min

TTFW (P50)

1469 ms

Speed Factor

0.1x

Billing

per 1s

Retention

No retention

6 languages

(view all)

streaming
CS
Word Timestamps
TranscribeView Benchmarks
Voxtral Mini Transcribe

Mistral (Voxtral)

Batch
France

Mistral Voxtral Mini Transcribe — EU-hosted batch transcription with speaker diarization, segment timestamps, and custom vocabulary (context bias). Strong accuracy at very low cost; handles recordings up to 3 hours.

$0.003/min
Best Value
Fastest
Budget Pick
+1

WER

8.86%

CER

5.30%

Speed Factor

0.2x

Uptime

100.0%

Billing

per 1s

Retention

No retention

8 languages

(view all)

Auto-detectDiarizationCustom Vocabulary
TranscribeView Benchmarks
Whisper 1 (API)

OpenAI

Batch
United States

OpenAI's Whisper API model

$0.006/min

WER

17.63%

CER

13.03%

Speed Factor

0.3x

Billing

per 1s

Retention

No retention

98 languages

(view all)

Auto-detectWord Timestamps
TranscribeView Benchmarks
Whisper Large V3

Groq

Batch
United States

OpenAI Whisper large-v3 served on Groq's LPU hardware — top Whisper accuracy at Groq speed and cost, with word-level timestamps and automatic language detection.

$0.0019/min

Speed Factor

0.0x

Billing

per 1s

Retention

No retention

12 languages

(view all)

Auto-detectWord Timestamps
TranscribeView Benchmarks
Whisper Large V3

OpenAI

Batch
United States

OpenAI's Whisper large-v3 model

$0.006/min

WER

18.81%

CER

15.54%

Speed Factor

0.3x

Uptime

100.0%

Billing

per 1s

Retention

No retention

98 languages

(view all)

Auto-detectWord Timestamps
TranscribeView Benchmarks
Whisper Large V3 Turbo

Groq

Batch
United States

OpenAI Whisper large-v3-turbo served on Groq's LPU hardware — very fast, very low cost batch transcription with word-level timestamps and automatic language detection.

$0.0007/min

Speed Factor

0.0x

Uptime

100.0%

Billing

per 1s

Retention

No retention

12 languages

(view all)

Auto-detectWord Timestamps
TranscribeView Benchmarks