Overview
C4, the Colossal Clean Crawled Corpus, is Google’s filtered version of Common Crawl created for training the T5 model. It applies aggressive cleaning heuristics to produce a 750GB English text corpus that became one of the most widely used pre-training datasets in the field.
What’s In It
C4 contains approximately 750GB of clean English text derived from April 2019 Common Crawl. The cleaning pipeline removes pages with offensive content, short pages, pages with boilerplate, and deduplicates at the document level. A multilingual variant mC4 covers 101 languages.
How It’s Used
C4 was the training corpus for T5 and has been used in PaLM, UL2, and Flan-T5. Its well-documented cleaning pipeline became a reference for the field. Many researchers compare their filtering approaches against C4’s heuristics.
Controversies
Research has shown C4’s filtering disproportionately removes text associated with minority dialects. The blocklist approach removes entire documents containing any flagged word, which eliminates legitimate content about marginalized communities.