Soniox async batch transcription — one multilingual model across 60+ languages with speaker diarization, per-word timestamps, and automatic language identification, at a low flat token-based rate.
25 languages · benchmarked
Soniox STT (Async) is a speech-to-text model from Soniox. In our standardized benchmarks it reaches 8.73% word error rate, ranking #1 of 34 models tested. It supports 25 languages and runs at $0.002/min — cheaper than most alternatives.
Its nearest benchmarked alternative is Whisper Large V3 Turbo. Best suited for multi-speaker conversations, accuracy-critical work, and high-volume, cost-sensitive pipelines.
Overall Score
73.7
#1 of 34
Word Error Rate
8.73%
Character Error Rate
6.18%
Match Error Rate
8.39%
Word Info Lost
13.26%
Avg Latency
6.2s
Benchmarks Run
34
overall WER · 34 models
Soniox STT (Async)
field · 33 models
lower-left is better
8 categories
WER
lower is better
Medical
1.6%
Legal
3.5%
Technical
4.9%
Finance
6.3%
General
6.6%
Conversational
12.6%
Code-Switching
16.0%
Noisy Environment
21.9%
| Category | WER | CER | MER | WIL | Latency | Benchmarks |
|---|---|---|---|---|---|---|
Code-Switching | 16.03% | 9.30% | 15.69% | 23.29% | 10.2s | 2 |
Conversational | 12.60% | 6.16% | 12.47% | 16.56% | 8.2s | 2 |
Finance | 6.32% | 2.82% | 6.25% | 10.13% | 5.1s | 2 |
General | 6.62% | 5.30% | 6.47% | 10.89% | 5.4s | 18 |
Legal | 3.52% | 1.90% | 3.49% | 4.55% | 7.6s | 2 |
Medical | 1.62% | 0.83% | 1.62% | 2.25% | 7.6s | 2 |
Noisy Environment | 21.92% | 16.67% | 20.02% | 32.16% | 5.1s | 4 |
Technical | 4.89% | 2.97% | 4.84% | 6.33% | 7.6s | 2 |
5 accents
| Accent | WER | Benchmarks |
|---|---|---|
African | 12.31% | 1 |
Australian | 12.31% | 1 |
Indian | 19.05% | 1 |
British | 7.46% | 1 |
American | 11.32% | 1 |
Rate
$0.0017
/min
Per-second billing. Bring your own provider key and pay your provider directly — a 5% routing fee applies (first 100 min/mo free).
Set up BYOKCost estimator
1 hour
$0.10
10 hours
$1.01
100 hours
$10.08
1,000 hours
$100.80
Billed per second of audio processed.
Model ID
soniox/stt-async
Authenticate every request with your secret API key as a Bearer token. Issue a key from your dashboard.
Use it with your agent
Paste this into Claude Code, Cursor, Codex or any agent with a shell. It installs the CLI and the OpenTranscription skill, then transcribes with this model.
Quickstart
Upload your audio, create a job with this model, then poll the job or set a webhook_url to be notified when it completes.
Key parameters
file_path
string
required
Storage path returned by the upload step.
model
string
required
The model to run this job on.
models
string[]
Ordered fallback chain: primary first, then backups tried in order if a provider fails.
language
string
ISO 639-1 language code. Omit to auto-detect.
diarization
boolean
Label speakers (A, B, …) in the output.
word_timestamps
boolean
Per-word start, end and confidence. Defaults to true; send false to leave them out.
webhook_url
string
HTTPS URL notified with the result when the job completes. Delivered events are signed (X-OT-Signature).
title
string
Display name for this job in the dashboard. Falls back to the file name.
Response
A completed job returns the transcript in the OpenTranscription Unified Schema (OTUS).
Webhooks — skip polling
We POST a signed event to your webhook_url when a job completes or fails; verify the X-OT-Signature (HMAC-SHA256, reject if older than 300 s) and dedupe on the event id, then fetch the full transcript via GET /api/v1/transcriptions/{id}.
Rate-limited per tier — see the X-RateLimit-* response headers.
Full API referenceUptime · 30d
100.00%
Error rate
0.00%
Avg latency
116.8s
Last incident
None in 30d
from the same field
same benchmark, every match
Supported Languages
Supported Formats
Features
Limits
Max file size
No documented limit
Max duration
5 h