Catalog · 43 models · 15 providers
Every speech-to-text model worth using, one rate card.
2
Total models
100
Languages
$0.0022
Cheapest /min
17.2%
Best WER
2 models
filtered from 43
Cartesia
Cartesia's newest streaming STT — lowest WER of any streaming model, native turn detection, and robust on alphanumerics like phone numbers, emails, and UUIDs. English only for now.
TTFW (P50)
Speed Factor
0.1x
Billing
per 1s
Retention
1 languages
(view all)
Cartesia
Whisper rearchitected for real-time and batch voice AI — fastest TTCT, 99-language coverage
WER
CER
Speed Factor
0.1x
Uptime
Billing
per 1s
Retention
100 languages
(view all)
Every head-to-head we index, grouped by model.
Deep dives on speech-to-text accuracy, latency and pricing.
What Amazon Transcribe Medical offers in 2026: features, specs, pricing vs Google and Nuance, HIPAA posture, research clues, and where it falls short.
What Google Cloud Chirp 3 actually is: release timeline, WER and Elo benchmarks, pricing, specs, known issues, and how it compares to Azure and ElevenLabs.
Where Deepgram Base fits in 2026: API behavior, variants, latency, specs, limitations, and when to pick Nova-3 or Flux instead.