Overview
VoxCeleb is a large-scale speaker recognition dataset from Oxford containing speech from over 7,000 celebrities extracted from YouTube interviews. It provides diverse real-world audio conditions for training speaker identification systems.
What’s In It
VoxCeleb1 has 153,000 utterances from 1,251 speakers. VoxCeleb2 expands to over 1 million utterances from 6,112 speakers. Audio includes background noise, music, reverb, and cross-talk. Speaker identities are verified through face recognition.
How It’s Used
VoxCeleb is the standard benchmark for speaker verification and identification. It trains models like ECAPA-TDNN and TitaNet. The real-world conditions make it valuable for robust speaker recognition research.
Controversies
The dataset was created without speakers’ consent using public YouTube appearances. Celebrity-focused data may not represent the general population’s vocal characteristics.