February 20, 2024Open Access

ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic

Puntos clave

Los puntos clave no están disponibles para este artículo en este momento.

Resumen

The focus of language model evaluation has transitioned towards reasoning and knowledge-intensive tasks, driven by advancements in pretraining large models. While state-of-the-art models are partially trained on large Arabic texts, evaluating their performance in Arabic remains challenging due to the limited availability of relevant datasets. To bridge this gap, we present ArabicMMLU, the first multi-task language understanding benchmark for Arabic language, sourced from school exams across diverse educational levels in different countries spanning North Africa, the Levant, and the Gulf regions. Our data comprises 40 tasks and 14,575 multiple-choice questions in Modern Standard Arabic (MSA), and is carefully constructed by collaborating with native speakers in the region. Our comprehensive evaluations of 35 models reveal substantial room for improvement, particularly among the best open-source models. Notably, BLOOMZ, mT0, LLama2, and Falcon struggle to achieve a score of 50%, while even the top-performing Arabic-centric model only achieves a score of 62.3%.

Connected Papers

Building similarity graph...

Analyzing shared references across papers

Discussion

Cite this study

Koto et al. (Tue,) studied this question.

www.synapsesocial.com/papers/68e786ffb6db6435876f9c3e — DOI: https://doi.org/10.48550/arxiv.2402.12840

Also consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

Authors

Fajri Koto

Haonan Li

Sara Shatnawi

Actions

References and Citations

Connected Papers

Building similarity graph...

Analyzing shared references across papers

ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic

Puntos clave

Resumen

Citation Network

Connected Papers

Discussion

Cite this study

Also consider

Authors

Actions

References and Citations

Citation Network

Connected Papers

Discussion