Home · Datasets · HowTo100M
DATASET

HowTo100M

Massive instructional video dataset with narration-aligned clips for learning text-video embeddings.

TARGET QUERY howto100m dataset · ~1K/mo
SIZE
136M video clips from 1.2M videos
CREATOR
INRIA
MODALITY
video
LICENSE
Apache 2.0
RELEASED
2019-06
OVERVIEW Updated 2026-05-17

Overview

HowTo100M is a massive instructional video dataset from INRIA containing 136 million clips from 1.2 million YouTube how-to videos. Its narration-aligned clips enable learning text-video embeddings without manual annotation.

What’s In It

The dataset contains 136 million video clips from 1.2 million YouTube instructional videos spanning 23,000 activities. Clips are paired with automatically transcribed narrations. Videos cover cooking, home repair, crafts, fitness, and other instructional content.

How It’s Used

HowTo100M enables self-supervised learning of video-text representations by exploiting natural alignment between narration and visual content. Models like VideoBERT and MIL-NCE use it for pre-training.

Controversies

ASR transcriptions contain significant errors creating noisy alignment. The narration often does not directly describe what is visually happening. YouTube deletions reduce dataset availability over time.