Home · Datasets · Conceptual Captions
DATASET

Conceptual Captions

Automatically harvested image-description pairs from the web for vision-language pre-training.

TARGET QUERY conceptual captions dataset · ~2K/mo
SIZE
12M image-caption pairs
CREATOR
Google
MODALITY
multimodal
LICENSE
Custom
RELEASED
2018-09
OVERVIEW Updated 2026-05-17

Overview

Conceptual Captions is Google’s automatically harvested dataset of image-description pairs extracted from web alt-text. It pioneered large-scale vision-language pre-training by providing millions of naturally occurring pairs without expensive human annotation.

What’s In It

The original CC3M contains 3.3 million pairs while CC12M expands to 12 million with relaxed filtering. Captions derive from HTML alt-text attributes, cleaned by replacing proper nouns with hypernyms and filtering for quality.

How It’s Used

Conceptual Captions enabled pre-training of vision-language models before larger datasets existed. CLIP, ALIGN, and Florence benefited from this approach for image-text retrieval and visual question answering.

Controversies

Alt-text quality varies enormously on the web with many descriptions being uninformative. The cleaning pipeline removes proper nouns which limits utility for named entity grounding.