Top Open ASR Models of 2026: WER, Languages, Latency, and License Comparison
Cohere
IBM
NVIDIA
Alibaba/Qwen
Mistral
Kyutai
Meta
OpenAI
The open speech recognition landscape has diversified beyond Whisper, with multiple models achieving sub-6% WER on the Open ASR Leaderboard. Key developments include Cohere Transcribe, IBM Granite Speech 4.1, and various models for throughput, streaming, and language coverage. The report emphasizes that license, language support, throughput, and streaming capability are now more critical than raw leaderboard rank.
The open ASR field has moved past the Whisper monoculture. In March 2026, Cohere released Transcribe (2B, Apache 2.0) with 5.42% average WER, followed by IBM's Granite Speech 4.1 2B at 5.33%. ARK-ASR-3B and MOSS-Transcribe-preview-2B later posted lower numbers. However, leaderboard scores are not directly comparable due to different test set compositions; Cohere's score drops from 5.42% to 5.84% when recalculated on the same seven sets as ARK. Private evaluation sets from Appen further reorder the rankings. Cohere Transcribe has strong production support and human preference wins. IBM Granite Speech 4.1 offers more features like keyword biasing and punctuation. Canary-Qwen-2.5B runs in transcription or LLM mode. Qwen3-ASR covers 52 languages. For throughput, Parakeet TDT 0.6B v3 leads at RTFx 3332.74, while Granite Speech 4.2B-NAR is non-autoregressive. For streaming, Voxtral Mini 4B Realtime and Kyutai STT offer low latency. Meta's Omnilingual ASR covers 1600+ languages. Whisper large-v3 remains relevant due to its MIT license and ecosystem. The diffusion-gemma-asr model from Interfaze is architecturally novel but not deployable. MOSS-Transcribe-Diarize integrates speaker labels. License choices split: Apache 2.0 (Cohere, IBM, Qwen), MIT (Whisper), CC-BY-4.0 (Canary, Parakeet, Kyutai). Teams are advised to evaluate based on license, language coverage, streaming vs. batch, their own audio WER, and cost per audio-hour.
- Сокращения
- ASR = Automatic Speech Recognition — Автоматическое распознавание речи
- WER = Word Error Rate — Коэффициент ошибок по словам
- RTFx = Real-Time Factor (x times real-time) — Коэффициент реального времени
- VAD = Voice Activity Detection — Обнаружение речевой активности
- LLM = Large Language Model — Большая языковая модель
- CTC = Connectionist Temporal Classification — Связывающая временная классификация
- CER = Character Error Rate — Коэффициент ошибок по символам
- MIT = Massachusetts Institute of Technology (license) — Лицензия MIT
- CC-BY-4.0 = Creative Commons Attribution 4.0 International — Creative Commons С указанием авторства 4.0
- GPU = Graphics Processing Unit — Графический процессор
Source: MarkTechPost —
original
