Google Telephony is a speech-to-text model from Google Cloud. In our standardized benchmarks it reaches 20.24% word error rate, ranking #28 of 35 models tested. It supports 9 languages and runs at $0.016/min — a premium option.
Its nearest benchmarked alternative is Soniox STT (Async).
Overall Score
39.9
#28 of 35
Word Error Rate
20.24%
Character Error Rate
12.59%
Match Error Rate
19.65%
Word Info Lost
26.34%
Avg Latency
21.0s
Benchmarks Run
32
last 30 runs
Word error rate
20.90%
↑ 27.1% · 30d
3 runs ago
latest
Avg latency
21.1s
↓ 5.2% · 30d
3 runs ago
latest
overall WER · 35 models
Google Telephony
field · 34 models
lower-left is better
8 categories
WER
lower is better
Legal
2.9%
Medical
3.1%
Technical
5.2%
Conversational
5.8%
Finance
11.1%
General
14.9%
Noisy Environment
53.8%
Code-Switching
68.9%
| Category | WER | CER | MER | WIL | Latency | Benchmarks |
|---|---|---|---|---|---|---|
Code-Switching | 68.89% | 62.50% | 68.26% | 74.67% | 34.9s | 2 |
Conversational | 5.81% | 3.11% | 5.67% | 8.72% | 32.4s | 2 |
Finance | 11.11% | 4.89% | 10.86% | 17.31% | 27.1s | 2 |
General | 14.92% | 8.54% | 14.70% | 22.75% | 17.1s | 16 |
Legal | 2.87% | 1.22% | 2.84% | 4.22% | 27.6s | 2 |
Medical | 3.07% | 1.47% | 3.05% | 4.78% | 24.5s | 2 |
Noisy Environment | 53.78% | 28.67% | 50.45% | 60.59% | 11.5s | 4 |
Technical | 5.21% | 2.61% | 5.14% | 8.50% | 29.6s | 2 |
5 accents
| Accent | WER | Benchmarks |
|---|---|---|
African | 27.69% | 1 |
Australian | 40.00% | 1 |
Indian | 42.86% | 1 |
British | 17.91% | 1 |
American | 26.42% | 1 |
Rate
$0.016
/min
Per-second billing. Bring your own provider key and pay your provider directly — a 5% routing fee applies (first 100 min/mo free).
Set up BYOKCost estimator
1 hour
$0.96
10 hours
$9.61
100 hours
$96.12
1,000 hours
$961.20
Billed per second of audio processed.
Model ID
google/telephony
Authenticate every request with your secret API key as a Bearer token. Issue a key from your dashboard.
Use it with your agent
Paste this into Claude Code, Cursor, Codex or any agent with a shell. It installs the CLI and the OpenTranscription skill, then transcribes with this model.
Quickstart
Upload your audio, create a job with this model, then poll the job or set a webhook_url to be notified when it completes.
Key parameters
file_path
string
required
Storage path returned by the upload step.
model
string
required
The model to run this job on.
models
string[]
Ordered fallback chain: primary first, then backups tried in order if a provider fails.
language
string
ISO 639-1 language code. Omit to auto-detect.
word_timestamps
boolean
Per-word start, end and confidence. Defaults to true; send false to leave them out.
webhook_url
string
HTTPS URL notified with the result when the job completes. Delivered events are signed (X-OT-Signature).
title
string
Display name for this job in the dashboard. Falls back to the file name.
Response
A completed job returns the transcript in the OpenTranscription Unified Schema (OTUS).
Webhooks — skip polling
We POST a signed event to your webhook_url when a job completes or fails; verify the X-OT-Signature (HMAC-SHA256, reject if older than 300 s) and dedupe on the event id, then fetch the full transcript via GET /api/v1/transcriptions/{id}.
Rate-limited per tier — see the X-RateLimit-* response headers.
Full API referenceSupported Languages
Supported Formats
Features
Limits
Max file size
No documented limit
Max duration
8 h