Overview
GSM8K is OpenAI’s benchmark of 8,500 linguistically diverse grade-school math word problems requiring multi-step reasoning. It exposed a critical weakness in language models and catalyzed research into mathematical reasoning capabilities.
What’s In It
The dataset contains 8,500 math problems requiring 2-8 steps of elementary arithmetic. Problems are written in natural language requiring no knowledge beyond grade-school math but demanding careful multi-step reasoning. Each includes a step-by-step solution.
How It’s Used
GSM8K is the primary benchmark for evaluating mathematical reasoning in language models. It motivated chain-of-thought prompting which showed dramatic improvements. The dataset also benchmarks tool-use approaches where models call calculators.
Controversies
Some researchers argue GSM8K problems are too formulaic and models may learn superficial patterns rather than genuine reasoning. Data contamination is a concern as problems have been widely shared online.