One API to every speech-to-text model worth using. Compare side by side, route to the best, pay by the second.
How it works · 3 steps
1
Upload audio
Drop any audio file — MP3, WAV, FLAC, and 10 more formats.
Drop audio file
2
Pick a model — or let us
Choose from 30+ models or use our smart router to auto-select by price, speed, or accuracy.
Whisper Large v3
Nova 3
Chirp 2
3
Get a structured transcript
Receive word-level timestamps, speaker labels, and confidence scores in a unified format.
{
"start": 0.0,
"end": 2.4,
"text": "Welcome to...",
"speaker": "A",
"confidence": 0.97
}Features
The stuff that's table-stakes the day you ship a voice product.
One API for every model. Diarization. 100+ languages. Realtime streams. Bring-your-own keys. Real benchmarks. Six things, all built in.
01 / 06
One API, every model
30+ models from 15 providers through a single unified endpoint. One integration, all the options.
provider : "deepgram" →
response_shape :{ transcript, words[], speakers[], confidence }
endpoint : POST /v1/transcribe // unchanged
01 / 06
One API, every model
30+ models from 15 providers through a single unified endpoint. One integration, all the options.
provider : "deepgram" →
response_shape :{ transcript, words[], speakers[], confidence }
endpoint : POST /v1/transcribe // unchanged
02 / 06
Realtime + batch
WebSocket streaming from mic or file upload. Async batch for pre-recorded audio. Same API.
01:07
02 / 06
Realtime + batch
WebSocket streaming from mic or file upload. Async batch for pre-recorded audio. Same API.
01:07
03 / 06
Speaker diarization
Identify who said what. Available on supported models with a single flag.
0:02
Quick update — the streaming refactor merged last night.
0:05
Nice. Any change to the latency numbers?
0:09
Down about 15%. I'll post the graph.
03 / 06
Speaker diarization
Identify who said what. Available on supported models with a single flag.
0:02
Quick update — the streaming refactor merged last night.
0:05
Nice. Any change to the latency numbers?
0:09
Down about 15%. I'll post the graph.
04 / 06
100+ languages
From English to Yoruba, with auto-detect on verified models. Code-switching for mixed-language audio.
+ 86 more
04 / 06
100+ languages
From English to Yoruba, with auto-detect on verified models. Code-switching for mixed-language audio.
+ 86 more
05 / 06
Bring your own key
Use your existing provider API keys. 100 free minutes per month, 5% platform fee after.
Deepgram
pass-through
AssemblyAI
idle
05 / 06
Bring your own key
Use your existing provider API keys. 100 free minutes per month, 5% platform fee after.
Deepgram
pass-through
AssemblyAI
idle
06 / 06
Real benchmarks
WER scores on golden-set audio across 8 categories. Event-driven, not vibes.
WER
1
2
3
06 / 06
Real benchmarks
WER scores on golden-set audio across 8 categories. Event-driven, not vibes.
WER
1
2
3
One-call integration
The SDKs handle the upload and the wait, in TypeScript or Python. Or call the HTTP API directly.
Deep dives on speech-to-text accuracy, latency and pricing.
What Amazon Transcribe Medical offers in 2026: features, specs, pricing vs Google and Nuance, HIPAA posture, research clues, and where it falls short.
What Google Cloud Chirp 3 actually is: release timeline, WER and Elo benchmarks, pricing, specs, known issues, and how it compares to Azure and ElevenLabs.
Where Deepgram Base fits in 2026: API behavior, variants, latency, specs, limitations, and when to pick Nova-3 or Flux instead.
$0.0007/min
lowest price
43+
models
105
languages
~772ms
lowest latency
Every model worth using
Every model worth using
Deepgram
Google Cloud
Speechmatics
OpenAI
Cartesia
Alibaba (Qwen)
Rev AI
The Ranker
We run automated WER benchmarks against golden-set audio across 8 categories — general, medical, legal, technical, conversational, noisy, finance, and code-switching. Benchmarks trigger on every model or pricing change.
Live ranker
| # | Model | Score | WER | Latency | Cost/min |
|---|---|---|---|---|---|
01 | Ink-Whisper Cartesia | 84.6 | 17.2% | 773ms | $0.0022 |
02 | Whisper Large V3 Turbo Groq | 81.3 | 14.3% | 3.4s | $0.0007 |
03 | GPT-4o Mini Transcribe OpenAI | 81.1 | 13.6% | 2.0s | $0.003 |
04 | Voxtral Mini Transcribe Mistral (Voxtral) | 80.7 | 12.7% | 2.3s | $0.003 |
05 | Whisper Large V3 Groq | 79.9 | 14.4% | 3.1s | $0.0019 |
84.6
WER: 17.2%
Latency: 773ms
Cost/min: $0.0022
WER: 14.3%
Latency: 3.4s
Cost/min: $0.0007
WER: 13.6%
Latency: 2.0s
Cost/min: $0.003
WER: 12.7%
Latency: 2.3s
Cost/min: $0.003
79.9
WER: 14.4%
Latency: 3.1s
Cost/min: $0.0019
+38 more models
View full ranker+38 more models
View full rankerRanked by a composite of accuracy, cost, and speed. See methodology
Pricing · transparent
One rate card across every provider. The price you see is the price you pay. BYOK users pay their provider directly.
Start free
Every new account gets 10 free transcriptions, up to 2 hours of audio — no credit card required.
From
$0.0007
/ minute
Per-second billing. Top up credits anytime and spend them across every provider.
5%
routing fee
Already have provider accounts? Bring your keys, pay your provider directly — we add a 5% routing fee on top.
Custom
For teams with large volumes, custom deployment, or support requirements.
No subscriptions, no minimums.
One API key. No contracts. No minimum spend.