Catalog · 41 models · 14 providers
Every speech-to-text model worth using, one rate card.
1
Total models
19
Languages
$0.006
Cheapest /min
15.3%
Best WER
1 models
filtered from 41
Microsoft Azure
Azure default STT model — 140+ languages, diarization, word timestamps
WER
CER
Speed Factor
0.3x
Billing
per 60s
Retention
19 languages
(view all)
Every pairing we index, grouped by model. Same audio, same scoring.
Google Cloud
$0.0107/min · 9.95% WER
Deepgram
$0.0077/min · 20.98% WER
Gladia
$0.0101/min · 9.80% WER
Amazon Web Services
$0.006/min · 13.02% WER
Deepgram
$0.0058/min · 16.54% WER
ElevenLabs
$0.004/min · 9.15% WER
Alibaba (Qwen)
$0.0021/min · 12.27% WER
Mistral (Voxtral)
$0.003/min · 12.70% WER
Groq
$0.0007/min · 14.31% WER
Groq
$0.0019/min · 14.41% WER
OpenAI
$0.006/min · 17.67% WER
OpenAI
$0.003/min · 13.58% WER
OpenAI
$0.006/min · 14.61% WER
Soniox
$0.0017/min · 8.73% WER
Accuracy, latency and pricing, checked against what providers actually ship.
What Amazon Transcribe Medical offers in 2026: features, specs, pricing vs Google and Nuance, HIPAA posture, research clues, and where it falls short.
July 3, 2026
Read
What Google Cloud Chirp 3 actually is: release timeline, WER and Elo benchmarks, pricing, specs, known issues, and how it compares to Azure and ElevenLabs.
July 3, 2026
Read
Where Deepgram Base fits in 2026: API behavior, variants, latency, specs, limitations, and when to pick Nova-3 or Flux instead.
July 3, 2026
Read