Overview
HowTo100M is a massive instructional video dataset from INRIA containing 136 million clips from 1.2 million YouTube how-to videos. Its narration-aligned clips enable learning text-video embeddings without manual annotation.
What’s In It
The dataset contains 136 million video clips from 1.2 million YouTube instructional videos spanning 23,000 activities. Clips are paired with automatically transcribed narrations. Videos cover cooking, home repair, crafts, fitness, and other instructional content.
How It’s Used
HowTo100M enables self-supervised learning of video-text representations by exploiting natural alignment between narration and visual content. Models like VideoBERT and MIL-NCE use it for pre-training.
Controversies
ASR transcriptions contain significant errors creating noisy alignment. The narration often does not directly describe what is visually happening. YouTube deletions reduce dataset availability over time.