Home · Datasets · RefinedWeb
DATASET

RefinedWeb

Massively filtered web dataset demonstrating that web-only data can match curated corpora quality.

TARGET QUERY refinedweb dataset falcon · ~2K/mo
SIZE
5T tokens from filtered web
CREATOR
TII (Technology Innovation Institute)
MODALITY
text
LICENSE
ODC-BY-1.0
RELEASED
2023-06
OVERVIEW Updated 2026-05-17

Overview

RefinedWeb is TII’s demonstration that web-only data, when properly filtered and deduplicated, can match the quality of curated multi-source datasets. The 5-trillion-token corpus powers the Falcon model family.

What’s In It

RefinedWeb contains approximately 5 trillion tokens extracted from Common Crawl using intensive filtering including URL filtering, text extraction, language identification, quality heuristics, and deduplication. A 600-billion-token subset is publicly released.

How It’s Used

RefinedWeb trains the Falcon family of models which demonstrated state-of-the-art performance among open models. It proved that careful web curation alone can produce competitive models. Its methodology influenced subsequent data pipeline designs.

Controversies

Despite filtering, web-only data inherits copyright and personal information concerns from Common Crawl. Only a fraction of the full dataset is publicly available. Some argue this may sacrifice certain knowledge domains.