Home · Datasets · C4
DATASET

C4

Colossal Clean Crawled Corpus - a filtered and deduplicated version of Common Crawl used to train T5.

TARGET QUERY c4 dataset · ~6K/mo
SIZE
750GB cleaned text
CREATOR
Google
MODALITY
text
LICENSE
ODC-BY
RELEASED
2019-10
OVERVIEW Updated 2026-05-17

Overview

C4, the Colossal Clean Crawled Corpus, is Google’s filtered version of Common Crawl created for training the T5 model. It applies aggressive cleaning heuristics to produce a 750GB English text corpus that became one of the most widely used pre-training datasets in the field.

What’s In It

C4 contains approximately 750GB of clean English text derived from April 2019 Common Crawl. The cleaning pipeline removes pages with offensive content, short pages, pages with boilerplate, and deduplicates at the document level. A multilingual variant mC4 covers 101 languages.

How It’s Used

C4 was the training corpus for T5 and has been used in PaLM, UL2, and Flan-T5. Its well-documented cleaning pipeline became a reference for the field. Many researchers compare their filtering approaches against C4’s heuristics.

Controversies

Research has shown C4’s filtering disproportionately removes text associated with minority dialects. The blocklist approach removes entire documents containing any flagged word, which eliminates legitimate content about marginalized communities.