Most accurate low-latency STT — <150ms, 90+ languages
76 languages · benchmarked
Scribe v2 Realtime is a speech-to-text model from ElevenLabs. It supports 76 languages and runs at $0.006/min — mid-range on price.
Its nearest benchmarked alternative is Soniox STT (Async). Best suited for live captioning and streaming transcription and multilingual workloads.
Realtime Score
68.9
TTFW (P50)
2113 ms
Flicker
57.9%
Cadence
1.0/s
RTF
1.09×
Word Error Rate
11.60%
Character Error Rate
7.77%
Benchmarks Run
35
Rate
$0.0065
/min
Per-second billing. Bring your own provider key and pay your provider directly — a 5% routing fee applies (first 100 min/mo free).
Set up BYOKCost estimator
1 hour
$0.39
10 hours
$3.89
100 hours
$38.88
1,000 hours
$388.80
Billed per second of audio processed.
Model ID
elevenlabs/scribe-v2-realtime
Authenticate every request with your secret API key as a Bearer token. Issue a key from your dashboard.
This model also supports realtime streaming over WebSocket.
Use it with your agent
Paste this into Claude Code, Cursor, Codex or any agent with a shell. It installs the CLI and the OpenTranscription skill, then transcribes with this model.
Quickstart
Upload your audio, create a job with this model, then poll the job or set a webhook_url to be notified when it completes.
Key parameters
file_path
string
required
Storage path returned by the upload step.
model
string
required
The model to run this job on.
models
string[]
Ordered fallback chain: primary first, then backups tried in order if a provider fails.
language
string
ISO 639-1 language code. Omit to auto-detect.
webhook_url
string
HTTPS URL notified with the result when the job completes. Delivered events are signed (X-OT-Signature).
title
string
Display name for this job in the dashboard. Falls back to the file name.
Response
A completed job returns the transcript in the OpenTranscription Unified Schema (OTUS).
Webhooks — skip polling
We POST a signed event to your webhook_url when a job completes or fails; verify the X-OT-Signature (HMAC-SHA256, reject if older than 300 s) and dedupe on the event id, then fetch the full transcript via GET /api/v1/transcriptions/{id}.
Rate-limited per tier — see the X-RateLimit-* response headers.
Full API referenceSupported Languages
Supported Formats
Features
Limits
Max file size
3 GB
Max duration
10 h