Catalog · 43 models · 15 providers
Every speech-to-text model worth using, one rate card.
43
Total models
105
Languages
$0.0007
Cheapest /min
8.7%
Best WER
43 models
WER
CER
Speed Factor
0.1x
Uptime
Billing
per 1s
Retention
19 languages
(view all)
Google Cloud
Google's latest generative ASR foundation model — 85+ languages
WER
CER
Speed Factor
0.3x
Uptime
Billing
per 15s
Retention
68 languages
(view all)
WER
CER
Speed Factor
0.1x
Billing
per 15s
Retention
40 languages
(view all)
WER
CER
Speed Factor
0.2x
Billing
per 15s
Retention
40 languages
(view all)
WER
CER
Speed Factor
0.3x
Uptime
Billing
per 1s
Retention
48 languages
(view all)
WER
CER
Speed Factor
0.1x
Uptime
Billing
per 1s
Retention
15 languages
(view all)
Deepgram
First conversational ASR model built for voice agents — model-integrated endpointing
TTFW (P50)
Speed Factor
0.1x
Uptime
Billing
per 1s
Retention
1 languages
(view all)
WER
CER
Speed Factor
0.1x
Uptime
Billing
per 1s
Retention
98 languages
(view all)
WER
CER
Speed Factor
0.2x
Uptime
Billing
per 1s
Retention
98 languages
(view all)
OpenAI
GPT-4o transcription with built-in speaker diarization
WER
CER
Speed Factor
0.2x
Uptime
Billing
per 1s
Retention
12 languages
(view all)
Cartesia
Cartesia's newest streaming STT — lowest WER of any streaming model, native turn detection, and robust on alphanumerics like phone numbers, emails, and UUIDs. English only for now.
TTFW (P50)
Speed Factor
0.1x
Billing
per 1s
Retention
1 languages
(view all)
Cartesia
Whisper rearchitected for real-time and batch voice AI — fastest TTCT, 99-language coverage
WER
CER
Speed Factor
0.1x
Uptime
Billing
per 1s
Retention
100 languages
(view all)
Google Cloud
Conformer model for long-form audio (minutes to hours)
WER
CER
Speed Factor
0.4x
Uptime
Billing
per 15s
Retention
40 languages
(view all)
WER
CER
Speed Factor
0.1x
Billing
per 15s
Retention
40 languages
(view all)
WER
CER
Speed Factor
0.1x
Uptime
Billing
per 1s
Retention
33 languages
(view all)
Deepgram
Optimized for human-to-bot interactions (IVR, voice assistants)
WER
CER
Speed Factor
0.1x
Billing
per 1s
Retention
1 languages
(view all)
Deepgram
Optimized for earnings calls with finance vocabulary
WER
CER
Speed Factor
0.1x
Billing
per 1s
Retention
1 languages
(view all)
WER
CER
Speed Factor
0.1x
Billing
per 1s
Retention
1 languages
(view all)
WER
CER
Speed Factor
0.1x
Billing
per 1s
Retention
1 languages
(view all)
Deepgram
Optimized for low-bandwidth single speaker voicemail
WER
CER
Speed Factor
0.1x
Billing
per 1s
Retention
1 languages
(view all)
Deepgram
Deepgram's flagship model — 53% lower WER vs competitors, code-switching support
WER
CER
Speed Factor
0.1x
Uptime
Billing
per 1s
Retention
47 languages
(view all)
Alibaba (Qwen)
Alibaba Qwen3-ASR Flash — multilingual batch transcription (25+ languages) with word-level timestamps and automatic language detection, served from Model Studio (Singapore). Handles files up to 12 hours.
WER
CER
Speed Factor
0.1x
Uptime
Billing
per 1s
Retention
25 languages
(view all)
Rev AI
Rev AI Reverb — async speech-to-text built on Rev's Reverb ASR model (trained on 3M+ hours of human-transcribed audio), with speaker diarization for up to 8 speakers and word-level timestamps. English-core with broad language support via the async API.
WER
CER
Speed Factor
0.2x
Uptime
Billing
per 1s
Retention
20 languages
(view all)
ElevenLabs
State-of-the-art batch STT — 90+ languages, speaker diarization, audio tagging
WER
CER
Speed Factor
0.2x
Uptime
Billing
per 1s
Retention
76 languages
(view all)
ElevenLabs
Most accurate low-latency STT — <150ms, 90+ languages
TTFW (P50)
Speed Factor
0.0x
Billing
per 1s
Retention
76 languages
(view all)
Gladia
Gladia Solaria-1 — EU-hosted real-time streaming speech-to-text across 100+ languages with native code-switching and ~103ms partial latency. Word-level timestamps and automatic language identification; speaker diarization is available for batch (Solaria-3), not this streaming model.
TTFW (P50)
Speed Factor
0.2x
Billing
per 1s
Retention
25 languages
(view all)
Gladia
Gladia Solaria-3 — EU-hosted batch speech-to-text tuned for the most accurate business audio across five core European languages (English, French, German, Spanish, Italian), with speaker diarization, custom vocabulary, and word-level timestamps.
WER
CER
Speed Factor
0.2x
Uptime
Billing
per 1s
Retention
5 languages
(view all)
Microsoft Azure
Azure default STT model — 140+ languages, diarization, word timestamps
WER
CER
Speed Factor
0.3x
Billing
per 60s
Retention
19 languages
(view all)
WER
CER
Speed Factor
0.1x
Uptime
Billing
per 1s
Retention
48 languages
(view all)
Soniox
Soniox async batch transcription — one multilingual model across 60+ languages with speaker diarization, per-word timestamps, and automatic language identification, at a low flat token-based rate.
WER
CER
Speed Factor
0.1x
Uptime
Billing
per 1s
Retention
25 languages
(view all)
Soniox
Soniox realtime streaming transcription — low-latency multilingual STT across 60+ languages with speaker diarization, per-word timestamps, and automatic language identification over a single WebSocket.
TTFW (P50)
Speed Factor
0.1x
Billing
per 1s
Retention
25 languages
(view all)
WER
CER
Speed Factor
0.2x
Uptime
Billing
per 15s
Retention
9 languages
(view all)
Amazon Web Services
AWS foundation model-powered ASR — 100+ languages
WER
CER
Speed Factor
0.4x
Uptime
Billing
per 1s
Retention
77 languages
(view all)
Amazon Web Services
Medical transcription with HIPAA eligibility
WER
CER
Speed Factor
0.4x
Uptime
Billing
per 1s
Retention
1 languages
(view all)
AssemblyAI
AssemblyAI Universal-3.5 Pro Realtime — the flagship streaming model, with native mid-sentence code-switching across 18 languages and conversation context carried across turns. Word-level timestamps and stable partials that are not rewritten as they arrive; speaker diarization is available on the batch model, not this streaming one.
TTFW (P50)
Speed Factor
0.1x
Billing
per 1s
Retention
18 languages
(view all)
AssemblyAI
AssemblyAI's most powerful speech language model — up to 1000 keyterm phrases
WER
CER
Speed Factor
0.2x
Uptime
Billing
per 1s
Retention
6 languages
(view all)
AssemblyAI
Purpose-built for real-time voice agents — ~300ms immutable transcripts
TTFW (P50)
Speed Factor
0.1x
Billing
per 1s
Retention
1 languages
(view all)
AssemblyAI
Multilingual streaming STT — English, Spanish, French, German, Italian, Portuguese
TTFW (P50)
Speed Factor
0.1x
Billing
per 1s
Retention
6 languages
(view all)
Mistral (Voxtral)
Mistral Voxtral Mini Transcribe — EU-hosted batch transcription with speaker diarization, segment timestamps, and custom vocabulary (context bias). Strong accuracy at very low cost; handles recordings up to 3 hours.
WER
CER
Speed Factor
0.2x
Uptime
Billing
per 1s
Retention
8 languages
(view all)
WER
CER
Speed Factor
0.3x
Uptime
Billing
per 1s
Retention
98 languages
(view all)
Groq
OpenAI Whisper large-v3 served on Groq's LPU hardware — top Whisper accuracy at Groq speed and cost, with word-level timestamps and automatic language detection.
WER
CER
Speed Factor
0.0x
Billing
per 1s
Retention
12 languages
(view all)
WER
CER
Speed Factor
0.3x
Uptime
Billing
per 1s
Retention
98 languages
(view all)
Groq
OpenAI Whisper large-v3-turbo served on Groq's LPU hardware — very fast, very low cost batch transcription with word-level timestamps and automatic language detection.
WER
CER
Speed Factor
0.0x
Uptime
Billing
per 1s
Retention
12 languages
(view all)
Every head-to-head we index, grouped by model.
Deep dives on speech-to-text accuracy, latency and pricing.
What Amazon Transcribe Medical offers in 2026: features, specs, pricing vs Google and Nuance, HIPAA posture, research clues, and where it falls short.
What Google Cloud Chirp 3 actually is: release timeline, WER and Elo benchmarks, pricing, specs, known issues, and how it compares to Azure and ElevenLabs.
Where Deepgram Base fits in 2026: API behavior, variants, latency, specs, limitations, and when to pick Nova-3 or Flux instead.