Catalog · 43 models · 15 providers
Every speech-to-text model worth using, one rate card.
10
Total models
49
Languages
$0.0058
Cheapest /min
16.4%
Best WER
10 models
filtered from 43
WER
CER
Speed Factor
0.1x
Uptime
Billing
per 1s
Retention
19 languages
(view all)
WER
CER
Speed Factor
0.1x
Uptime
Billing
per 1s
Retention
15 languages
(view all)
Deepgram
First conversational ASR model built for voice agents — model-integrated endpointing
TTFW (P50)
Speed Factor
0.1x
Uptime
Billing
per 1s
Retention
1 languages
(view all)
WER
CER
Speed Factor
0.1x
Uptime
Billing
per 1s
Retention
33 languages
(view all)
Deepgram
Optimized for human-to-bot interactions (IVR, voice assistants)
WER
CER
Speed Factor
0.1x
Billing
per 1s
Retention
1 languages
(view all)
Deepgram
Optimized for earnings calls with finance vocabulary
WER
CER
Speed Factor
0.1x
Billing
per 1s
Retention
1 languages
(view all)
WER
CER
Speed Factor
0.1x
Billing
per 1s
Retention
1 languages
(view all)
WER
CER
Speed Factor
0.1x
Billing
per 1s
Retention
1 languages
(view all)
Deepgram
Optimized for low-bandwidth single speaker voicemail
WER
CER
Speed Factor
0.1x
Billing
per 1s
Retention
1 languages
(view all)
Deepgram
Deepgram's flagship model — 53% lower WER vs competitors, code-switching support
WER
CER
Speed Factor
0.1x
Uptime
Billing
per 1s
Retention
47 languages
(view all)
Every head-to-head we index, grouped by model.
Deep dives on speech-to-text accuracy, latency and pricing.
What Amazon Transcribe Medical offers in 2026: features, specs, pricing vs Google and Nuance, HIPAA posture, research clues, and where it falls short.
What Google Cloud Chirp 3 actually is: release timeline, WER and Elo benchmarks, pricing, specs, known issues, and how it compares to Azure and ElevenLabs.
Where Deepgram Base fits in 2026: API behavior, variants, latency, specs, limitations, and when to pick Nova-3 or Flux instead.