OpenTranscription
OpenTranscription
RankerModelsPlayground
OpenTranscription
OpenTranscription

One API to every speech-to-text model worth using. Compare them on your audio, route to the best one, pay per second.

Platform status

Product

RankerModelsTranscriptionsPlaygroundBlog

Developers

DocumentationReliabilityAPI VersioningStatus

Legal

Privacy PolicyTerms of ServiceSupport
© 2026 OpenTranscription
Back to Models

Whisper Large V3 vs Voxtral Mini Transcribe

2 models compared on the same benchmark — accuracy, latency, price, language coverage and capabilities. Voxtral Mini Transcribe leads on raw accuracy; Voxtral Mini Transcribe is fastest; Whisper Large V3 covers the most languages; Whisper Large V3 is cheapest.

Open both in playground

Verdict · who wins what

Most accurate

Voxtral Mini Transcribe

12.7% WER

Mistral (Voxtral)

Fastest

Voxtral Mini Transcribe

0.2× realtime

Mistral (Voxtral)

Most languages

Whisper Large V3

12 languages

Groq

Cheapest

Whisper Large V3

$0.0019/min

Groq

Comparing

2 models · same benchmark

1

Batch

Whisper Large V3

Groq

$0.0019/min
#11 · 79.9

OpenAI Whisper large-v3 served on Groq's LPU hardware — top Whisper accuracy at Groq speed and cost, with word-level timestamps and automatic language detection.

Use This ModelDetails →

2

Batch

Voxtral Mini Transcribe

Mistral (Voxtral)

$0.0030/min
#7 · 80.7

Mistral Voxtral Mini Transcribe — EU-hosted batch transcription with speaker diarization, segment timestamps, and custom vocabulary (context bias). Strong accuracy at very low cost; handles recordings up to 3 hours.

Use This ModelDetails →

Overview

30-day benchmark average

Overall score

0–100

79.9

80.7

Accuracy rank

of 35

#11

#7

Word error rate

WER

14.4%

12.7%

Character error rate

CER

11.2%

8.0%

Match error rate

MER

13.3%

11.5%

Word info lost

WIL

18.2%

15.8%

Languages

12

8

Cost vs. accuracy

overall WER vs price

Voxtral Mini Transcribe

Whisper Large V3

field · 33 models

lower-left is better

Accuracy by category

WER · lower is better

code_switching

68.5%
21.0%

conversational

10.9%
10.0%

finance

6.4%

Accuracy by accent

WER · English

african

26.2%
27.7%

australian

12.3%
6.2%

indian

19.1%
12.7%

Speed & latency

p50 → p99

Processing speed

× realtime

0.0×

0.2×

Latency p50

3.1s

2.3s

Latency p90

3.6s

2.3s

Realtime / streaming

Streaming benchmark averages

Realtime rank

—

—

Realtime score

0–100

—

—

Time to first word

—

—

Final drain

—

—

Real-time factor

—

—

Flicker

—

—

Cadence

—

—

Reliability

Rolling 30 days

Uptime

—

100.00%

Error rate

—

0.00%

Pricing

billed per second

Rate

per minute

$0.0019

$0.0030

Billing granularity

per 1s

per 1s

1 hour of audio

$0.11

$0.18

100 hours

Capabilities

Feature support

Speaker diarization

No

Yes

Word-level timestamps

Yes

No

Code-switching

No

Coverage & deployment

Languages · formats · regions

Languages

12

8

Mode

Batch

Batch

Region

United States

France

Audio formats

6.8%

general

8.1%
6.1%

legal

2.7%
3.1%

medical

2.8%
6.3%

noisy

36.3%
47.4%

technical

5.3%
5.6%

uk

6.0%
3.0%

us

5.7%
0.0%

Latency p99

3.6s

2.3s

$11.16

$18.00

1,000 hours

$111.60

$180.00

No

Auto-detect language

Yes

Yes

Live / streaming

No

No

Custom vocabulary

No

Yes

mp3
wav
flac
m4a
ogg
webm
mp3
wav
flac
m4a
ogg
webm

Max file size

100 MB

No limit

Max duration

No limit

15 min

Compliance

—

GDPR

Frequently asked questions

Is Whisper Large V3 or Voxtral Mini Transcribe more accurate?

Voxtral Mini Transcribe is more accurate, with a 12.7% word error rate versus 14.4% for Whisper Large V3, on our standardized benchmark.

Which is cheaper, Whisper Large V3 or Voxtral Mini Transcribe?

Whisper Large V3 is cheaper at $0.0019/min versus $0.0030/min for Voxtral Mini Transcribe.

Should you choose Whisper Large V3 or Voxtral Mini Transcribe?

Choose Whisper Large V3 for lower cost and language coverage; choose Voxtral Mini Transcribe for accuracy and speed.