Overview
Mozilla Common Voice is the world’s largest open multilingual speech dataset, built through voluntary voice donations. Its crowdsourced approach provides unprecedented language diversity and speaker variety for training inclusive speech recognition systems.
What’s In It
Common Voice contains over 19,000 hours of validated speech across 100+ languages contributed by hundreds of thousands of volunteers. Each clip includes speaker metadata where provided. Monthly releases add new languages and hours.
How It’s Used
Common Voice is essential for building speech recognition in low-resource languages. Whisper, MMS, and XLS-R all incorporate it. It enables research in accent-robust ASR and multilingual speech processing.
Controversies
Contribution quality varies with some recordings having background noise or mispronunciations. Language coverage remains uneven despite recruitment efforts for underrepresented languages.