Catalog · 40 models · 13 providers
Every speech-to-text model worth using, one rate card.
40
Total models
106
Languages
$0.0007
Cheapest /min
4.7%
Best WER
40 models
WER
CER
Speed Factor
0.1x
Uptime
Billing
per 1s
Retention
19 languages
(view all)
WER
CER
Speed Factor
0.3x
Uptime
Billing
per 1s
Retention
78 languages
(view all)
Google Cloud
Google's latest generative ASR foundation model — 85+ languages
WER
CER
Speed Factor
0.3x
Uptime
Billing
per 15s
Retention
72 languages
(view all)
WER
CER
Speed Factor
0.1x
Billing
per 15s
Retention
40 languages
(view all)
WER
CER
Speed Factor
0.2x
Billing
per 15s
Retention
40 languages
(view all)
WER
CER
Speed Factor
0.1x
Uptime
Billing
per 1s
Retention
15 languages
(view all)
WER
CER
Speed Factor
0.3x
Uptime
Billing
per 1s
Retention
48 languages
(view all)
Deepgram
First conversational ASR model built for voice agents — model-integrated endpointing
TTFW (P50)
Speed Factor
0.1x
Billing
per 1s
Retention
1 languages
(view all)
Speed Factor
0.1x
Uptime
Billing
per 1s
Retention
98 languages
(view all)
Speed Factor
0.2x
Uptime
Billing
per 1s
Retention
98 languages
(view all)
OpenAI
GPT-4o transcription with built-in speaker diarization
WER
CER
Speed Factor
0.2x
Billing
per 1s
Retention
12 languages
(view all)
Cartesia
Cartesia's newest streaming STT — lowest WER of any streaming model, native turn detection, and robust on alphanumerics like phone numbers, emails, and UUIDs. English only for now.
TTFW (P50)
Speed Factor
0.1x
Billing
per 1s
Retention
1 languages
(view all)
Cartesia
Whisper rearchitected for real-time and batch voice AI — fastest TTCT, 99-language coverage
WER
CER
Speed Factor
0.1x
Uptime
Billing
per 1s
Retention
100 languages
(view all)
Google Cloud
Conformer model for long-form audio (minutes to hours)
WER
CER
Speed Factor
0.4x
Billing
per 15s
Retention
40 languages
(view all)
WER
CER
Speed Factor
0.1x
Billing
per 15s
Retention
40 languages
(view all)
WER
CER
Speed Factor
0.1x
Uptime
Billing
per 1s
Retention
33 languages
(view all)
Deepgram
Optimized for human-to-bot interactions (IVR, voice assistants)
WER
CER
Speed Factor
0.1x
Billing
per 1s
Retention
1 languages
(view all)
Deepgram
Optimized for earnings calls with finance vocabulary
WER
CER
Speed Factor
0.1x
Billing
per 1s
Retention
1 languages
(view all)
WER
CER
Speed Factor
0.1x
Billing
per 1s
Retention
1 languages
(view all)
WER
CER
Speed Factor
0.1x
Billing
per 1s
Retention
1 languages
(view all)
Deepgram
Optimized for low-bandwidth single speaker voicemail
WER
CER
Speed Factor
0.1x
Billing
per 1s
Retention
1 languages
(view all)
Deepgram
Deepgram's flagship model — 53% lower WER vs competitors, code-switching support
WER
CER
Speed Factor
0.1x
Uptime
Billing
per 1s
Retention
47 languages
(view all)
Alibaba (Qwen)
Alibaba Qwen3-ASR Flash — multilingual batch transcription (25+ languages) with word-level timestamps and automatic language detection, served from Model Studio (Singapore). Handles files up to 12 hours.
Speed Factor
0.1x
Uptime
Billing
per 1s
Retention
25 languages
(view all)
Rev AI
Rev AI Reverb — async speech-to-text built on Rev's Reverb ASR model (trained on 3M+ hours of human-transcribed audio), with speaker diarization for up to 8 speakers and word-level timestamps. English-core with broad language support via the async API.
WER
CER
Speed Factor
0.2x
Billing
per 1s
Retention
20 languages
(view all)
ElevenLabs
State-of-the-art batch STT — 90+ languages, speaker diarization, audio tagging
Speed Factor
0.2x
Uptime
Billing
per 1s
Retention
76 languages
(view all)
ElevenLabs
Most accurate low-latency STT — <150ms, 90+ languages
TTFW (P50)
Speed Factor
0.0x
Billing
per 1s
Retention
76 languages
(view all)
Gladia
Gladia Solaria-1 — EU-hosted real-time streaming speech-to-text across 100+ languages with native code-switching and ~103ms partial latency. Word-level timestamps and automatic language identification; speaker diarization is available for batch (Solaria-3), not this streaming model.
TTFW (P50)
Speed Factor
0.2x
Billing
per 1s
Retention
25 languages
(view all)
Gladia
Gladia Solaria-3 — EU-hosted batch speech-to-text tuned for the most accurate business audio across five core European languages (English, French, German, Spanish, Italian), with speaker diarization, custom vocabulary, and word-level timestamps.
WER
CER
Speed Factor
0.2x
Billing
per 1s
Retention
5 languages
(view all)
WER
CER
Speed Factor
0.1x
Uptime
Billing
per 1s
Retention
48 languages
(view all)
WER
CER
Speed Factor
0.2x
Uptime
Billing
per 15s
Retention
9 languages
(view all)
Amazon Web Services
AWS foundation model-powered ASR — 100+ languages
WER
CER
Speed Factor
0.4x
Uptime
Billing
per 1s
Retention
77 languages
(view all)
Amazon Web Services
Medical transcription with HIPAA eligibility
WER
CER
Speed Factor
0.4x
Uptime
Billing
per 1s
Retention
1 languages
(view all)
AssemblyAI
AssemblyAI's most powerful speech language model — up to 1000 keyterm phrases
WER
CER
Speed Factor
0.2x
Uptime
Billing
per 1s
Retention
6 languages
(view all)
AssemblyAI
Purpose-built for real-time voice agents — ~300ms immutable transcripts
TTFW (P50)
Speed Factor
0.1x
Billing
per 1s
Retention
1 languages
(view all)
AssemblyAI
Multilingual streaming STT — English, Spanish, French, German, Italian, Portuguese
TTFW (P50)
Speed Factor
0.1x
Billing
per 1s
Retention
6 languages
(view all)
Mistral (Voxtral)
Mistral Voxtral Mini Transcribe — EU-hosted batch transcription with speaker diarization, segment timestamps, and custom vocabulary (context bias). Strong accuracy at very low cost; handles recordings up to 3 hours.
WER
CER
Speed Factor
0.2x
Uptime
Billing
per 1s
Retention
8 languages
(view all)
WER
CER
Speed Factor
0.3x
Billing
per 1s
Retention
98 languages
(view all)
Groq
OpenAI Whisper large-v3 served on Groq's LPU hardware — top Whisper accuracy at Groq speed and cost, with word-level timestamps and automatic language detection.
Speed Factor
0.0x
Billing
per 1s
Retention
12 languages
(view all)
WER
CER
Speed Factor
0.3x
Uptime
Billing
per 1s
Retention
98 languages
(view all)
Groq
OpenAI Whisper large-v3-turbo served on Groq's LPU hardware — very fast, very low cost batch transcription with word-level timestamps and automatic language detection.
Speed Factor
0.0x
Uptime
Billing
per 1s
Retention
12 languages
(view all)