Home · Datasets · MMLU
DATASET

MMLU

Massive Multitask Language Understanding - 57-subject benchmark measuring broad academic knowledge.

TARGET QUERY mmlu benchmark · ~8K/mo
SIZE
15.9K multiple-choice questions
CREATOR
UC Berkeley
MODALITY
text
LICENSE
MIT
RELEASED
2020-09
OVERVIEW Updated 2026-05-17

Overview

MMLU (Massive Multitask Language Understanding) is a comprehensive benchmark measuring knowledge and reasoning across 57 academic subjects. Created by Dan Hendrycks at UC Berkeley, it became the most widely cited benchmark for comparing frontier language model capabilities.

What’s In It

MMLU contains 15,908 multiple-choice questions spanning 57 subjects from STEM, humanities, social sciences, and professional domains. Questions range from elementary to expert difficulty.

How It’s Used

MMLU scores are the most commonly reported metric for frontier model comparisons. Every major model release reports MMLU performance. The benchmark tests whether models have acquired broad world knowledge and can apply reasoning across domains.

Controversies

MMLU has faced scrutiny for containing erroneous answers. The multiple-choice format limits assessment to recognition rather than generation. Some questions are outdated or rely on specific textbook interpretations.