Overview
Dolma is AI2’s three-trillion token open corpus curated for language model training transparency. It represents the most thoroughly documented and reproducible pre-training dataset, with detailed provenance tracking and a fully open data pipeline.
What’s In It
Dolma contains approximately 3 trillion tokens from Common Crawl, The Pile subsets, Reddit, GitHub code, Semantic Scholar papers, and books. Each document includes metadata about its source, cleaning steps, and quality scores. The full pipeline code is open source.
How It’s Used
Dolma serves as the training corpus for AI2’s OLMo model family, which aims to be the most open and reproducible LLM project. Researchers use it to study pre-training data composition and develop better filtering methods.
Controversies
While Dolma prioritizes transparency, it still inherits copyright concerns from web-crawled content. The project navigates the tension between openness for reproducibility and potential misuse.