Home · Datasets · VoxCeleb
DATASET

VoxCeleb

Large-scale speaker identification dataset extracted from YouTube interview videos.

TARGET QUERY voxceleb dataset · ~3K/mo
SIZE
7K+ speakers, 1M+ utterances
CREATOR
University of Oxford
MODALITY
audio
LICENSE
CC-BY-SA 4.0
RELEASED
2017-06
OVERVIEW Updated 2026-05-17

Overview

VoxCeleb is a large-scale speaker recognition dataset from Oxford containing speech from over 7,000 celebrities extracted from YouTube interviews. It provides diverse real-world audio conditions for training speaker identification systems.

What’s In It

VoxCeleb1 has 153,000 utterances from 1,251 speakers. VoxCeleb2 expands to over 1 million utterances from 6,112 speakers. Audio includes background noise, music, reverb, and cross-talk. Speaker identities are verified through face recognition.

How It’s Used

VoxCeleb is the standard benchmark for speaker verification and identification. It trains models like ECAPA-TDNN and TitaNet. The real-world conditions make it valuable for robust speaker recognition research.

Controversies

The dataset was created without speakers’ consent using public YouTube appearances. Celebrity-focused data may not represent the general population’s vocal characteristics.