Whisper rearchitected for real-time and batch voice AI — fastest TTCT, 99-language coverage
100 languages · benchmarked
Ink-Whisper is a speech-to-text model from Cartesia. In our standardized benchmarks it reaches 17.20% word error rate, ranking #20 of 35 models tested. It supports 100 languages and runs at $0.002/min — cheaper than most alternatives.
Its nearest benchmarked alternative is Soniox STT (Async). Best suited for live captioning and streaming transcription, multilingual workloads, and high-volume, cost-sensitive pipelines.
Overall Score
84.6
#20 of 35
Word Error Rate
17.20%
Character Error Rate
11.58%
Match Error Rate
15.06%
Word Info Lost
20.23%
Avg Latency
773ms
Benchmarks Run
35
last 30 runs
Word error rate
17.20%
↓ 24.6% · 30d
4 runs ago
latest
Avg latency
773ms
↓ 18.4% · 30d
4 runs ago
latest
overall WER · 35 models
Ink-Whisper
field · 34 models
lower-left is better
8 categories
WER
lower is better
Medical
3.4%
Finance
7.5%
Technical
8.3%
General
9.5%
Conversational
9.8%
Legal
22.7%
Noisy Environment
39.4%
Code-Switching
80.3%
p50 → p99 · 4 runs
Latency distribution
lower is better
748ms
p50
947ms
p90
947ms
p95
947ms
p99
5 accents
Realtime Score
67.1
TTFW (P50)
4992 ms
Flicker
0.0%
Cadence
0.0/s
RTF
1.02×
Word Error Rate
18.19%
Character Error Rate
13.15%
Benchmarks Run
35
Rate
$0.0022
/min
Per-second billing. Bring your own provider key and pay your provider directly — a 5% routing fee applies (first 100 min/mo free).
Set up BYOKCost estimator
1 hour
$0.13
10 hours
$1.33
100 hours
$13.32
1,000 hours
$133.20
Billed per second of audio processed.
Model ID
cartesia/ink-whisper
Authenticate every request with your secret API key as a Bearer token. Issue a key from your dashboard.
This model also supports realtime streaming over WebSocket.
Use it with your agent
Paste this into Claude Code, Cursor, Codex or any agent with a shell. It installs the CLI and the OpenTranscription skill, then transcribes with this model.
Quickstart
Upload your audio, create a job with this model, then poll the job or set a webhook_url to be notified when it completes.
Key parameters
file_path
string
required
Storage path returned by the upload step.
model
string
required
The model to run this job on.
models
string[]
Ordered fallback chain: primary first, then backups tried in order if a provider fails.
language
string
ISO 639-1 language code. Omit to auto-detect.
webhook_url
string
HTTPS URL notified with the result when the job completes. Delivered events are signed (X-OT-Signature).
title
string
Display name for this job in the dashboard. Falls back to the file name.
Response
A completed job returns the transcript in the OpenTranscription Unified Schema (OTUS).
Webhooks — skip polling
We POST a signed event to your webhook_url when a job completes or fails; verify the X-OT-Signature (HMAC-SHA256, reject if older than 300 s) and dedupe on the event id, then fetch the full transcript via GET /api/v1/transcriptions/{id}.
Rate-limited per tier — see the X-RateLimit-* response headers.
Full API referenceSupported Languages
Supported Formats
Features
Limits
Max file size
No documented limit
Max duration
No documented limit