Azure default STT model — 140+ languages, diarization, word timestamps
19 languages · benchmarked
Azure Speech is a speech-to-text model from Microsoft Azure. In our standardized benchmarks it reaches 15.34% word error rate, ranking #16 of 34 models tested. It supports 19 languages and runs at $0.006/min — mid-range on price.
Its nearest benchmarked alternative is Soniox STT (Async). Best suited for live captioning and streaming transcription and multi-speaker conversations.
Overall Score
74.4
#16 of 34
Word Error Rate
15.34%
Character Error Rate
10.67%
Match Error Rate
14.88%
Word Info Lost
18.78%
Avg Latency
2.0s
Benchmarks Run
35
overall WER · 34 models
Azure Speech
field · 33 models
lower-left is better
8 categories
WER
lower is better
Legal
1.7%
Medical
2.8%
Technical
7.0%
Finance
7.1%
Conversational
9.6%
General
10.5%
Noisy Environment
40.0%
Code-Switching
60.1%
| Category | WER | CER | MER | WIL | Latency | Benchmarks |
|---|---|---|---|---|---|---|
Code-Switching | 60.13% | 26.54% | 54.37% | 69.93% | 3.4s | 2 |
Conversational | 9.58% | 6.22% | 9.58% | 11.69% | 2.4s | 2 |
Finance | 7.11% | 3.35% | 7.03% | 11.62% | 2.4s | 2 |
General | 10.55% | 8.84% | 10.32% | 14.17% | 1.8s | 19 |
Legal | 1.66% | 1.09% | 1.66% | 1.98% | 2.2s | 2 |
Medical | 2.75% | 1.44% | 2.75% | 4.17% | 2.2s | 2 |
Noisy Environment | 40.00% | 29.94% | 40.00% | 42.86% | 1.2s | 4 |
Technical | 7.00% | 4.19% | 6.98% | 8.90% | 2.6s | 2 |
5 accents
| Accent | WER | Benchmarks |
|---|---|---|
African | 23.08% | 1 |
Australian | 23.08% | 1 |
Indian | 41.27% | 1 |
British | 14.93% | 1 |
American | 11.32% | 1 |
Realtime Score
72.4
TTFW (P50)
2052 ms
Flicker
13.2%
Cadence
1.7/s
RTF
1.01×
Word Error Rate
15.82%
Character Error Rate
10.48%
Benchmarks Run
35
Rate
$0.006
/min
Per-second billing. Bring your own provider key and pay your provider directly — a 5% routing fee applies (first 100 min/mo free).
Set up BYOKCost estimator
1 hour
$0.36
10 hours
$3.60
100 hours
$36
1,000 hours
$360
Billed per second of audio processed.
Model ID
azure/speech
Authenticate every request with your secret API key as a Bearer token. Issue a key from your dashboard.
This model also supports realtime streaming over WebSocket.
Use it with your agent
Paste this into Claude Code, Cursor, Codex or any agent with a shell. It installs the CLI and the OpenTranscription skill, then transcribes with this model.
Quickstart
Upload your audio, create a job with this model, then poll the job or set a webhook_url to be notified when it completes.
Key parameters
file_path
string
required
Storage path returned by the upload step.
model
string
required
The model to run this job on.
models
string[]
Ordered fallback chain: primary first, then backups tried in order if a provider fails.
language
string
ISO 639-1 language code. Omit to auto-detect.
diarization
boolean
Label speakers (A, B, …) in the output.
word_timestamps
boolean
Per-word start, end and confidence. Defaults to true; send false to leave them out.
webhook_url
string
HTTPS URL notified with the result when the job completes. Delivered events are signed (X-OT-Signature).
title
string
Display name for this job in the dashboard. Falls back to the file name.
Response
A completed job returns the transcript in the OpenTranscription Unified Schema (OTUS).
Webhooks — skip polling
We POST a signed event to your webhook_url when a job completes or fails; verify the X-OT-Signature (HMAC-SHA256, reject if older than 300 s) and dedupe on the event id, then fetch the full transcript via GET /api/v1/transcriptions/{id}.
Rate-limited per tier — see the X-RateLimit-* response headers.
Full API referenceSupported Languages
Supported Formats
Features
Limits
Max file size
1 GB
Max duration
4 h