OpenTranscription
OpenTranscription
RankerModelsPlayground

Blog

Analysis, model profiles and deep dives on speech-to-text.

OpenTranscription
OpenTranscription

One API to every speech-to-text model worth using. Compare them on your audio, route to the best one, pay per second.

Platform status

Product

RankerModelsTranscriptionsPlaygroundBlog

Developers

DocumentationReliabilityAPI VersioningStatus

Legal

Privacy PolicyTerms of ServiceSupport
© 2026 OpenTranscription

Published July 3, 2026

Google Cloud's latest_short and the batch paradox

Why Google's latest_short model is built for short utterances, not short files, and when running it through batch recognition actually makes sense.

Read post

Published July 3, 2026

Google command_and_search (Google Speech-to-Text): model profile

Reference profile of Google's command_and_search transcription model in Cloud Speech-to-Text, a legacy short-utterance model for voice commands and voice search.

Read post

Published July 3, 2026

Google's command_and_search model: the voice-search engine that quietly became legacy

The history, architecture, and current status of Google's command_and_search speech model, from 2016 Cloud Speech API beta to legacy status behind Chirp.

Read post

Published July 3, 2026

Ink-Whisper: how Cartesia rebuilt Whisper for real-time voice agents

What Cartesia's Ink-Whisper got right on latency, where its accuracy fell behind by 2026, and why it mattered more as a stepping stone than a benchmark.

Read post

Published July 3, 2026

Ink-Whisper: model profile

Reference profile of Ink-Whisper, Cartesia's Whisper-derived streaming speech-to-text model for real-time voice agents, launched June 10, 2025.

Read post

Published July 3, 2026

Microsoft Azure Whisper: model profile

Reference spec sheet for OpenAI's Whisper model as offered on Microsoft Azure: delivery paths, limits, languages, pricing, benchmarks, and release history.

Read post

Published July 3, 2026

OpenAI Whisper large-v3: model profile

Reference profile of OpenAI Whisper large-v3: architecture, training data, release history, deployment options, pricing, limitations, and sources.

Read post

Published July 3, 2026

Scribe v2 Realtime: ElevenLabs makes its play for live speech-to-text

ElevenLabs' Scribe v2 Realtime claims sub-150 ms latency, 93.5% accuracy in 30 languages, and $0.39/hr pricing. What the public record actually supports.

Read post

Published July 3, 2026

Scribe v2 Realtime: model profile

Reference profile of Scribe v2 Realtime, ElevenLabs' streaming speech-to-text model released November 11, 2025: specs, benchmarks, pricing, limits.

Read post

Published July 3, 2026

Universal-3 Pro: what AssemblyAI shipped, and what it still won't say

AssemblyAI's Universal-3 Pro reviewed: promptable transcription, WER benchmarks, pricing, compliance caveats, and what the public record still hides.

Read post

Published July 3, 2026

Whisper large-v3 and the shift from open research to transcription infrastructure

How OpenAI's Whisper large-v3 went from MIT-licensed research artifact to the baseline layer of a managed transcription stack, and what got left unresolved.

Read post

Published July 3, 2026

Whisper on Azure: what Microsoft actually sells, and where it fits now

How Microsoft packages OpenAI's Whisper across Azure OpenAI and Azure Speech: limits, pricing signals, benchmarks, security, and where it fits in 2026.

Read post
Previous

Page 5 of 5