Deepgram's flagship model — 53% lower WER vs competitors, code-switching support
47 languages · benchmarked
Overall Score
71.1
#24 of 28
Word Error Rate
20.98%
Character Error Rate
11.48%
Match Error Rate
20.03%
Word Info Lost
25.79%
Avg Latency
1.0s
Benchmarks Run
35
last 30 runs
Word error rate
15.18%
↓ 31.1% · 30d
4 runs ago
latest
Avg latency
1.0s
↓ 21.8% · 30d
4 runs ago
latest
overall WER · 28 models
Nova-3
field · 27 models
lower-left is better
8 categories
WER
lower is better
Technical
2.8%
Legal
3.8%
Finance
9.5%
Medical
9.5%
Conversational
12.2%
Noisy Environment
21.8%
General
23.4%
Code-Switching
63.0%
p50 → p99 · 4 runs
Latency distribution
lower is better
1.3s
p50
1.8s
p90
1.8s
p95
1.8s
p99
5 accents
Realtime Score
76.6
TTFW (P50)
973 ms
Flicker
9.6%
Cadence
0.8/s
RTF
1.01×
Word Error Rate
18.17%
Character Error Rate
11.05%
Benchmarks Run
34
pay per second
Rate
$0.0077
/min
Per-second billing. Bring your own provider key and pay your provider directly — a 5% routing fee applies (first 100 min/mo free).
Set up BYOKCost estimator
1 hour
$0.46
10 hours
$4.61
100 hours
$46.08
1,000 hours
$460.80
Billed per second of audio processed.
POST /api/v1/transcriptions
Model ID
deepgram/nova-3
Authenticate every request with your secret API key as a Bearer token. Issue a key from your dashboard.
This model also supports realtime streaming over WebSocket.
Quickstart
Upload your audio, create a job with this model, then poll the job or set a webhook_url to be notified when it completes.
Key parameters
file_path
string
required
Storage path returned by the upload step.
model
string
required
The model to run this job on.
language
string
ISO 639-1 language code. Omit to auto-detect.
diarization
boolean
Label speakers (A, B, …) in the output.
webhook_url
string
HTTPS URL notified with the result when the job completes. Delivered events are signed (X-OT-Signature).
Response
A completed job returns the transcript in the OpenTranscription Unified Schema (OTUS).
Webhooks — skip polling
We POST a signed event to your webhook_url when a job completes or fails; verify the X-OT-Signature (HMAC-SHA256, reject if older than 300 s) and dedupe on the event id, then fetch the full transcript via GET /api/v1/transcriptions/{id}.
Rate-limited per tier — see the X-RateLimit-* response headers.
Full API referenceSupported Languages
Supported Formats
Features
Nova-3 is a speech-to-text model from Deepgram. In our standardized benchmarks it reaches 30.75% word error rate, ranking #24 of 28 models tested. It supports 47 languages and runs at $0.008/min — mid-range on price.
Its nearest benchmarked alternative is Amazon Transcribe. Best suited for live captioning and streaming transcription, multi-speaker conversations, and multilingual workloads.