Hugging Face published QIMMA, an Arabic LLM leaderboard that validates benchmark quality before evaluating models. Developed by Leen AlQadi, Ahmed Alzubaidi, Mohammed Alyafeai, Maitha Alhammadi, Shaikha Alsuwaidi, Omar saif alkaabi, Basma Boussaha, and Hakim Hacid, the platform consolidates 109 subsets from 14 source benchmarks into over 52,000 samples across seven domains. A multi-stage validation pipeline combining automated assessment and human annotation found systematic quality issues, including incorrect answers and cultural bias, across widely used Arabic benchmarks.
No score is assigned. Sources and their independence are shown in the citation chain below.