Home · Datasets · The Stack
DATASET

The Stack

The largest open code dataset with opt-out mechanisms covering 358 programming languages from GitHub.

TARGET QUERY the stack dataset · ~3K/mo
SIZE
6.4TB source code
CREATOR
BigCode / Hugging Face
MODALITY
code
LICENSE
Various
RELEASED
2022-11
OVERVIEW Updated 2026-05-17

Overview

The Stack is the largest open code dataset, containing 6.4TB of source code from 358 programming languages collected from permissive-license GitHub repositories. Created by BigCode, it includes a governance framework for developer opt-out.

What’s In It

The Stack v1 contains 6.4TB of deduplicated source code from GitHub repositories with permissive licenses. It covers 358 programming languages identified by file extension. Near-deduplication removes repetitive boilerplate code.

How It’s Used

The Stack provides training data for StarCoder, SantaCoder, and numerous other code models. Its permissive-license filtering makes derived models suitable for commercial deployment.

Controversies

The Am I In The Stack tool lets developers opt out, but the default is opt-in. Debates continue about whether open-source licenses implicitly cover AI training.