Research 🇷🇺 24.07.2026 11:01

MERA Reason: New Public Leaderboard for Evaluating Reasoning Models on Russian

Т-БанкТ-Банк Alibaba/QwenAlibaba/Qwen
A new public leaderboard called MERA Reason (part of the MERA TEXT platform) has been launched to evaluate reasoning capabilities of AI models specifically on Russian-language mathematical and reasoning tasks. It combines four distinct datasets: Luzitania (hard olympiad-level math), TMath (diverse math problems), MMReD (dense-context multi-hop reasoning), and aims to provide an objective, reproducible benchmark for reasoning models.
The MERA Reason leaderboard has been introduced to address the need for objective measurement of reasoning in large language models, especially for the Russian language. It comprises four datasets: Luzitania (251 hard math problems from high-level olympiads, filtered by difficulty using GPT-OSS-120B), TMath (310 problems from Russian olympiads covering algebra, geometry, combinatorics, etc.), and MMReD (a synthetic benchmark for dense-context reasoning where every token is relevant). The leaderboard allows comparison of reasoning models in a reproducible environment and will be part of a larger update to the MERA platform expected this summer.
Сокращения
GPU = Graphics Processing Unit — графический процессор
NIAH = Needle-In-A-Haystack — иголка в стоге сена
MCQ = Multiple Choice Question — вопрос с множественным выбором
Source: Habr — хаб ИИ — original
Our earlier posts on this topic ↓
Fresh news