Blog
Analysis, model profiles and deep dives on speech-to-text.
Analysis, model profiles and deep dives on speech-to-text.

Published July 3, 2026
A practitioner's guide to Google Cloud Speech-to-Text latest_long: Conformer roots, pricing, quotas, diarization, and how it compares to V2 and Chirp.

Published July 3, 2026
Why Google's latest_short model is built for short utterances, not short files, and when running it through batch recognition actually makes sense.

Published July 3, 2026
The history, architecture, and current status of Google's command_and_search speech model, from 2016 Cloud Speech API beta to legacy status behind Chirp.

Published July 3, 2026
What Cartesia's Ink-Whisper got right on latency, where its accuracy fell behind by 2026, and why it mattered more as a stepping stone than a benchmark.

Published July 3, 2026
Scribe v2 Realtime by ElevenLabs: specs, pricing, benchmarks, and known limitations behind the sub-150ms, 93.5%-accuracy, 90+ language claims.

Published July 3, 2026
AssemblyAI's Universal-3 Pro reviewed: promptable transcription, WER benchmarks, pricing, compliance caveats, and what the public record still hides.

Published July 3, 2026
How Whisper large-v3 went from OpenAI's MIT-licensed research release to the baseline of a managed transcription stack — specs, data, and what's unresolved.

Published July 3, 2026
How Microsoft packages OpenAI's Whisper across Azure OpenAI and Azure Speech: limits, pricing signals, benchmarks, security, and where it fits in 2026.