OpenTranscription
OpenTranscription
RankerModelsPlayground
OpenTranscription
OpenTranscription

One API to every speech-to-text model worth using. Compare them on your audio, route to the best one, pay per second.

Platform status

Product

RankerModelsTranscriptionsPlaygroundBlog

Developers

DocumentationReliabilityAPI VersioningStatus

Legal

Privacy PolicyTerms of ServiceSupport
© 2026 OpenTranscription
Back to Models

Nova-3 vs Soniox STT (Async)

2 models compared on the same benchmark — accuracy, latency, price, language coverage and capabilities. Soniox STT (Async) leads on raw accuracy; Soniox STT (Async) is fastest; Nova-3 covers the most languages; Soniox STT (Async) is cheapest.

Open both in playground

Verdict · who wins what

Most accurate

Soniox STT (Async)

8.7% WER

Soniox

Fastest

Soniox STT (Async)

0.1× realtime

Soniox

Most languages

Nova-3

47 languages

Deepgram

Cheapest

Soniox STT (Async)

$0.0017/min

Soniox

Comparing

2 models · same benchmark

1

Batch & Realtime

Nova-3

Deepgram

$0.0077/min
#29 · 71.1

Deepgram's flagship model — 53% lower WER vs competitors, code-switching support

Use This ModelDetails →

2

Batch

Soniox STT (Async)

Soniox

$0.0017/min
#1 · 73.7

Soniox async batch transcription — one multilingual model across 60+ languages with speaker diarization, per-word timestamps, and automatic language identification, at a low flat token-based rate.

Use This ModelDetails →

Overview

30-day benchmark average

Overall score

0–100

71.1

73.7

Accuracy rank

of 35

#29

#1

Word error rate

WER

21.0%

8.7%

Character error rate

CER

11.5%

6.2%

Match error rate

MER

20.0%

8.4%

Word info lost

WIL

25.8%

13.3%

Languages

47

25

Cost vs. accuracy

overall WER vs price

Nova-3

Soniox STT (Async)

field · 33 models

lower-left is better

Accuracy by category

WER · lower is better

code_switching

63.0%
16.0%

conversational

12.2%
12.6%

finance

9.5%
6.3%

general

23.4%
6.6%

legal

3.8%
3.5%

medical

9.5%
1.6%

noisy

21.8%
21.9%

technical

2.8%
4.9%

Accuracy by accent

WER · English

african

26.2%
12.3%

australian

15.4%
12.3%

indian

27.0%
19.1%

uk

6.0%
7.5%

us

11.3%
11.3%

Speed & latency

p50 → p99

Processing speed

× realtime

0.1×

0.1×

Latency p50

1.3s

6.2s

Latency p90

1.8s

6.2s

Latency p99

1.8s

6.2s

Realtime / streaming

Streaming benchmark averages

Realtime rank

of 17

#3

—

Realtime score

0–100

78.2

—

Time to first word

973ms

—

Final drain

181ms

—

Real-time factor

1.01×

—

Flicker

9.6%

—

Cadence

0.8/s

—

Reliability

Rolling 30 days

Uptime

50.00%

100.00%

Error rate

50.00%

0.00%

Pricing

billed per second

Rate

per minute

$0.0077

$0.0017

Billing granularity

per 1s

per 1s

1 hour of audio

$0.46

$0.10

100 hours

$46.08

$10.08

1,000 hours

$460.80

$100.80

Capabilities

Feature support

Speaker diarization

Yes

Yes

Word-level timestamps

Yes

Yes

Code-switching

Yes

No

Auto-detect language

Yes

Yes

Live / streaming

Yes

No

Custom vocabulary

Yes

No

Coverage & deployment

Languages · formats · regions

Languages

47

25

Mode

Batch & Realtime

Batch

Region

United States

United States

Audio formats

mp3
wav
flac
m4a
ogg
webm
mp3
wav
flac
m4a
ogg
webm
aac

Max file size

2 GB

No limit

Max duration

No limit

5 h

Compliance

HIPAA
SOC2
GDPR

—

Frequently asked questions

Is Nova-3 or Soniox STT (Async) more accurate?

Soniox STT (Async) is more accurate, with a 8.7% word error rate versus 21.0% for Nova-3, on our standardized benchmark.

Which is cheaper, Nova-3 or Soniox STT (Async)?

Soniox STT (Async) is cheaper at $0.0017/min versus $0.0077/min for Nova-3.

Should you choose Nova-3 or Soniox STT (Async)?

Choose Nova-3 for language coverage; choose Soniox STT (Async) for accuracy, speed, and lower cost.