Overview
RefinedWeb is TII’s demonstration that web-only data, when properly filtered and deduplicated, can match the quality of curated multi-source datasets. The 5-trillion-token corpus powers the Falcon model family.
What’s In It
RefinedWeb contains approximately 5 trillion tokens extracted from Common Crawl using intensive filtering including URL filtering, text extraction, language identification, quality heuristics, and deduplication. A 600-billion-token subset is publicly released.
How It’s Used
RefinedWeb trains the Falcon family of models which demonstrated state-of-the-art performance among open models. It proved that careful web curation alone can produce competitive models. Its methodology influenced subsequent data pipeline designs.
Controversies
Despite filtering, web-only data inherits copyright and personal information concerns from Common Crawl. Only a fraction of the full dataset is publicly available. Some argue this may sacrifice certain knowledge domains.