Overview
MMLU (Massive Multitask Language Understanding) is a comprehensive benchmark measuring knowledge and reasoning across 57 academic subjects. Created by Dan Hendrycks at UC Berkeley, it became the most widely cited benchmark for comparing frontier language model capabilities.
What’s In It
MMLU contains 15,908 multiple-choice questions spanning 57 subjects from STEM, humanities, social sciences, and professional domains. Questions range from elementary to expert difficulty.
How It’s Used
MMLU scores are the most commonly reported metric for frontier model comparisons. Every major model release reports MMLU performance. The benchmark tests whether models have acquired broad world knowledge and can apply reasoning across domains.
Controversies
MMLU has faced scrutiny for containing erroneous answers. The multiple-choice format limits assessment to recognition rather than generation. Some questions are outdated or rely on specific textbook interpretations.