Home · Datasets · StarCoder Training Data
DATASET

StarCoder Training Data

A massive code dataset spanning 86 programming languages filtered for quality and licenses.

TARGET QUERY starcoder data · ~2K/mo
SIZE
783GB of code
CREATOR
BigCode
MODALITY
code
LICENSE
Various per license
RELEASED
2023-05
OVERVIEW Updated 2026-05-17

Overview

The StarCoder training data is BigCode’s curated code dataset spanning 86 programming languages, assembled with careful attention to licensing and personal information removal. It powers the StarCoder family of code generation models.

What’s In It

The dataset contains 783GB of code from The Stack, filtered for permissive licenses and processed to remove personal information. It covers 86 programming languages with deduplication applied. Git commit messages and Jupyter notebooks are also included.

How It’s Used

StarCoder data trains the StarCoder model family which achieve strong code generation performance. The permissive licensing focus makes these models suitable for commercial use.

Controversies

Even with license filtering, questions about the legality of training on open-source code persist. Some developers feel that permissive licenses were not intended to cover AI training use cases.