AssemblyAI's most powerful speech language model — up to 1000 keyterm phrases
6 languages · benchmarked
Universal-3.5 Pro is a speech-to-text model from AssemblyAI. In our standardized benchmarks it reaches 17.44% word error rate, ranking #21 of 35 models tested. It supports 6 languages and runs at $0.003/min — cheaper than most alternatives.
Its nearest benchmarked alternative is Soniox STT (Async). Best suited for multi-speaker conversations and high-volume, cost-sensitive pipelines.
Overall Score
67.4
#21 of 35
Word Error Rate
17.44%
Character Error Rate
10.37%
Match Error Rate
16.91%
Word Info Lost
20.23%
Avg Latency
5.6s
Benchmarks Run
35
last 30 runs
Word error rate
9.85%
↓ 43.1% · 30d
4 runs ago
latest
Avg latency
5.8s
↑ 1.4% · 30d
4 runs ago
latest
overall WER · 35 models
Universal-3.5 Pro
field · 34 models
lower-left is better
8 categories
WER
lower is better
Legal
1.8%
Medical
2.6%
Technical
3.8%
Finance
5.6%
Conversational
10.2%
Code-Switching
18.2%
General
20.6%
Noisy Environment
33.9%
| Category | WER | CER | MER | WIL | Latency | Benchmarks |
|---|---|---|---|---|---|---|
Code-Switching | 18.20% | 11.93% | 17.71% | 24.69% | 7.5s | 2 |
Conversational | 10.21% | 6.61% | 10.09% | 12.27% | 7.6s | 2 |
Finance | 5.63% | 2.39% | 5.56% | 9.12% | 9.0s | 2 |
General | 20.57% | 11.23% | 20.57% | 23.60% | 5.2s | 19 |
Legal | 1.79% | 1.18% | 1.77% | 2.10% | 6.6s | 2 |
Medical | 2.59% | 0.93% | 2.59% | 3.70% | 5.2s | 2 |
Noisy Environment | 33.85% | 24.52% | 29.56% | 36.78% | 3.5s | 4 |
Technical | 3.76% | 2.71% | 3.74% | 4.45% | 6.6s | 2 |
p50 → p99 · 4 runs
Latency distribution
lower is better
5.8s
p50
5.8s
p90
5.8s
p95
5.8s
p99
5 accents
| Accent | WER | Benchmarks |
|---|---|---|
African | 26.15% | 1 |
Australian | 3.08% | 1 |
Indian | 11.11% | 1 |
British | 0.00% | 1 |
American | 0.00% | 1 |
Rate
$0.0035
/min
Per-second billing. Bring your own provider key and pay your provider directly — a 5% routing fee applies (first 100 min/mo free).
Set up BYOKCost estimator
1 hour
$0.21
10 hours
$2.09
100 hours
$20.88
1,000 hours
$208.80
Billed per second of audio processed.
Model ID
assemblyai/universal-3-pro
Authenticate every request with your secret API key as a Bearer token. Issue a key from your dashboard.
Use it with your agent
Paste this into Claude Code, Cursor, Codex or any agent with a shell. It installs the CLI and the OpenTranscription skill, then transcribes with this model.
Quickstart
Upload your audio, create a job with this model, then poll the job or set a webhook_url to be notified when it completes.
Key parameters
file_path
string
required
Storage path returned by the upload step.
model
string
required
The model to run this job on.
models
string[]
Ordered fallback chain: primary first, then backups tried in order if a provider fails.
language
string
ISO 639-1 language code. Omit to auto-detect.
diarization
boolean
Label speakers (A, B, …) in the output.
word_timestamps
boolean
Per-word start, end and confidence. Defaults to true; send false to leave them out.
custom_words
string[]
Names, jargon and product terms to bias the model toward. Up to 1000 entries.
vocabulary_list_id
uuid
A vocabulary list saved in your workspace. Merged with custom_words when both are sent.
webhook_url
string
HTTPS URL notified with the result when the job completes. Delivered events are signed (X-OT-Signature).
title
string
Display name for this job in the dashboard. Falls back to the file name.
Response
A completed job returns the transcript in the OpenTranscription Unified Schema (OTUS).
Webhooks — skip polling
We POST a signed event to your webhook_url when a job completes or fails; verify the X-OT-Signature (HMAC-SHA256, reject if older than 300 s) and dedupe on the event id, then fetch the full transcript via GET /api/v1/transcriptions/{id}.
Rate-limited per tier — see the X-RateLimit-* response headers.
Full API referenceSupported Languages
Supported Formats
Features
Limits
Max file size
5 GB
Max duration
10 h