Overview
Conceptual Captions is Google’s automatically harvested dataset of image-description pairs extracted from web alt-text. It pioneered large-scale vision-language pre-training by providing millions of naturally occurring pairs without expensive human annotation.
What’s In It
The original CC3M contains 3.3 million pairs while CC12M expands to 12 million with relaxed filtering. Captions derive from HTML alt-text attributes, cleaned by replacing proper nouns with hypernyms and filtering for quality.
How It’s Used
Conceptual Captions enabled pre-training of vision-language models before larger datasets existed. CLIP, ALIGN, and Florence benefited from this approach for image-text retrieval and visual question answering.
Controversies
Alt-text quality varies enormously on the web with many descriptions being uninformative. The cleaning pipeline removes proper nouns which limits utility for named entity grounding.