Overview
ROOTS is the 1.6TB multilingual corpus assembled by the BigScience collaborative to train BLOOM, the first large-scale open multilingual language model. It was created through an unprecedented community effort involving hundreds of researchers.
What’s In It
ROOTS spans 46 natural languages and 13 programming languages totaling 1.6TB. Sources include Oscar, Wikipedia, books, scientific papers, and culture-specific web sources for underrepresented languages. Each source was evaluated for quality, consent, and language representation.
How It’s Used
ROOTS trained BLOOM (176B parameters) and its variants. The dataset’s careful documentation established best practices for responsible dataset creation and community-driven data governance.
Controversies
Despite careful curation, achieving representative coverage of all 46 languages proved challenging with high-resource languages dominating. The lengthy governance process delayed the project timeline.