Home · Datasets · Dolma
DATASET

Dolma

A three-trillion token open corpus curated by AI2 for training OLMo and advancing transparent LLM research.

TARGET QUERY dolma dataset · ~2K/mo
SIZE
3T tokens
CREATOR
AI2
MODALITY
text
LICENSE
ODC-BY
RELEASED
2024-01
OVERVIEW Updated 2026-05-17

Overview

Dolma is AI2’s three-trillion token open corpus curated for language model training transparency. It represents the most thoroughly documented and reproducible pre-training dataset, with detailed provenance tracking and a fully open data pipeline.

What’s In It

Dolma contains approximately 3 trillion tokens from Common Crawl, The Pile subsets, Reddit, GitHub code, Semantic Scholar papers, and books. Each document includes metadata about its source, cleaning steps, and quality scores. The full pipeline code is open source.

How It’s Used

Dolma serves as the training corpus for AI2’s OLMo model family, which aims to be the most open and reproducible LLM project. Researchers use it to study pre-training data composition and develop better filtering methods.

Controversies

While Dolma prioritizes transparency, it still inherits copyright concerns from web-crawled content. The project navigates the tension between openness for reproducibility and potential misuse.