ResearchModels 🇺🇸 27.07.2026 10:05

FFASR Leaderboard Launched: Benchmark for Evaluating ASR in Real Acoustic Conditions

Treble TechnologiesTreble Technologies Hugging FaceHugging Face OpenAIOpenAI IBMIBM CohereCohere MetaMeta SpeechBrainSpeechBrain
Hugging Face and Treble Technologies have launched the Far-Field ASR (FFASR) Leaderboard, the first open community-driven benchmark for evaluating ASR models under realistic far-field conditions like reverberation, background noise, and microphone distance. Initial results show far-field WER at low SNR is several times higher than near-field WER, highlighting a significant gap between benchmark performance and real-world deployment.
The FFASR Leaderboard, created by Hugging Face and Treble Technologies, is the first open community-driven benchmark for evaluating ASR models under realistic far-field acoustic conditions. It addresses the persistent gap between benchmark performance and real-world deployment, where models that score well on standard near-field benchmarks often degrade substantially under reverberation, background noise, and microphone distance. The leaderboard evaluates models across nine conditions, with four determining the primary ranking score. Acoustic data is generated using Treble's hybrid simulation engine, combining wave-based and geometrical-acoustics modeling to capture phenomena like diffraction and interference. Fourteen fully furnished rooms are included, ranging from 20 to 470 m³. Alongside WER, the leaderboard reports RTFx (audio seconds per inference second) for every submission, evaluated on an NVIDIA L4 GPU. Initial results show far-field WER at low SNR is several times higher than near-field WER. The leaderboard supports Whisper variants, IBM Granite Speech, Cohere Transcribe, Wav2Vec2, HuBERT, SpeechBrain, and most other ASR architectures via Hugging Face model ID submission. A custom evaluator option allows more complex inference stacks. Future plans include multi-talker scenarios, microphone array evaluation, and echo cancellation.
Сокращения
ASR = Automatic Speech Recognition — автоматическое распознавание речи
WER = Word Error Rate — процент словесных ошибок
SNR = Signal-to-Noise Ratio — отношение сигнал/шум
RTFx = Real-Time Factor (audio seconds per inference second) — фактор реального времени (секунды аудио за секунду инференса)
GPU = Graphics Processing Unit — графический процессор
RIR = Room Impulse Response — импульсная характеристика помещения
CTC = Connectionist Temporal Classification — коннекционистская временная классификация
Source: Hugging Face blog — original
Our earlier posts on this topic ↓
Fresh news