OpenTranscription
OpenTranscription
RankerModelsPlayground
OpenTranscription
OpenTranscription

One API to every speech-to-text model worth using. Compare them on your audio, route to the best one, pay per second.

Platform status

Product

RankerModelsTranscriptionsPlaygroundBlog

Developers

DocumentationReliabilityAPI VersioningStatus

Legal

Privacy PolicyTerms of ServiceSupport
© 2026 OpenTranscription
Back to Models

Soniox STT (Async) vs Speechmatics Enhanced

2 models compared on the same benchmark — accuracy, latency, price, language coverage and capabilities. Soniox STT (Async) leads on raw accuracy; Speechmatics Enhanced is fastest; Speechmatics Enhanced covers the most languages; Soniox STT (Async) is cheapest.

Open both in playground

Verdict · who wins what

Most accurate

Soniox STT (Async)

8.7% WER

Soniox

Fastest

Speechmatics Enhanced

0.3× realtime

Speechmatics

Most languages

Speechmatics Enhanced

48 languages

Speechmatics

Cheapest

Soniox STT (Async)

$0.0017/min

Soniox

Comparing

2 models · same benchmark

1

Batch

Soniox STT (Async)

Soniox

$0.0017/min
#1 · 73.7

Soniox async batch transcription — one multilingual model across 60+ languages with speaker diarization, per-word timestamps, and automatic language identification, at a low flat token-based rate.

Use This ModelDetails →

2

Batch

Speechmatics Enhanced

Speechmatics

$0.0083/min
#6 · 47.3

Highest accuracy model — 55+ languages, best-in-class

Use This ModelDetails →

Overview

30-day benchmark average

Overall score

0–100

73.7

47.3

Accuracy rank

of 35

#1

#6

Word error rate

WER

8.7%

12.3%

Character error rate

CER

6.2%

8.0%

Match error rate

MER

8.4%

11.3%

Word info lost

WIL

13.3%

16.2%

Languages

25

48

Cost vs. accuracy

overall WER vs price

Speechmatics Enhanced

Soniox STT (Async)

field · 33 models

lower-left is better

Accuracy by category

WER · lower is better

code_switching

16.0%
39.2%

conversational

12.6%
6.4%

finance

6.3%

Accuracy by accent

WER · English

african

12.3%
7.7%

australian

12.3%
7.7%

indian

19.1%
9.5%

Speed & latency

p50 → p99

Processing speed

× realtime

0.1×

0.3×

Latency p50

6.2s

8.7s

Latency p90

6.2s

12.1s

Realtime / streaming

Streaming benchmark averages

Realtime rank

—

—

Realtime score

0–100

—

—

Time to first word

—

—

Final drain

—

—

Real-time factor

—

—

Flicker

—

—

Cadence

—

—

Reliability

Rolling 30 days

Uptime

100.00%

100.00%

Error rate

0.00%

0.00%

Pricing

billed per second

Rate

per minute

$0.0017

$0.0083

Billing granularity

per 1s

per 1s

1 hour of audio

$0.10

$0.50

100 hours

Capabilities

Feature support

Speaker diarization

Yes

Yes

Word-level timestamps

Yes

Yes

Code-switching

No

Coverage & deployment

Languages · formats · regions

Languages

25

48

Mode

Batch

Batch

Region

United States

Europe

Audio formats

6.8%

general

6.6%
7.2%

legal

3.5%
2.2%

medical

1.6%
5.7%

noisy

21.9%
40.8%

technical

4.9%
5.3%

uk

7.5%
3.0%

us

11.3%
1.9%

Latency p99

6.2s

12.1s

$10.08

$49.68

1,000 hours

$100.80

$496.80

Yes

Auto-detect language

Yes

Yes

Live / streaming

No

No

Custom vocabulary

No

Yes

mp3
wav
flac
m4a
ogg
webm
aac
mp3
wav
flac
m4a
ogg

Max file size

No limit

1 GB

Max duration

5 h

No limit

Compliance

—

SOC2
GDPR

Frequently asked questions

Is Soniox STT (Async) or Speechmatics Enhanced more accurate?

Soniox STT (Async) is more accurate, with a 8.7% word error rate versus 12.3% for Speechmatics Enhanced, on our standardized benchmark.

Which is cheaper, Soniox STT (Async) or Speechmatics Enhanced?

Soniox STT (Async) is cheaper at $0.0017/min versus $0.0083/min for Speechmatics Enhanced.

Should you choose Soniox STT (Async) or Speechmatics Enhanced?

Choose Soniox STT (Async) for accuracy and lower cost; choose Speechmatics Enhanced for speed and language coverage.