Comprehensive profiles of the datasets powering AI development. Text, image, code, audio, video, and multimodal training data.
The complete arXiv preprint repository used as training data for scientific language models.
Colossal Clean Crawled Corpus - a filtered and deduplicated version of Common Crawl used to train T5.
The largest open web crawl archive serving as the foundational data source for most modern language models.
Automatically harvested image-description pairs from the web for vision-language pre-training.
A three-trillion token open corpus curated by AI2 for training OLMo and advancing transparent LLM research.
Curated mixture of human preference datasets for Direct Preference Optimization alignment training.
Grade School Math 8K - a benchmark of linguistically diverse math word problems testing multi-step reasoning.
A challenging commonsense reasoning benchmark testing physical and social situation understanding.
Massive instructional video dataset with narration-aligned clips for learning text-video embeddings.
Hand-crafted Python programming challenges used to evaluate functional correctness of code generation models.
The foundational large-scale image classification dataset that catalyzed the deep learning revolution.
Large-scale video understanding benchmark for human action recognition from YouTube clips.
The largest openly available image-text dataset used to train Stable Diffusion and other multimodal models.
A subset of LAION-5B filtered by predicted aesthetic quality scores for training image generation models.
The standard benchmark for automatic speech recognition derived from LibriVox audiobooks.
Massive Multitask Language Understanding - 57-subject benchmark measuring broad academic knowledge.
The largest open multilingual speech dataset built through crowdsourced voice donations.
Common Objects in Context - the standard benchmark for object detection segmentation and image captioning.
Real Google search queries paired with Wikipedia passages to evaluate open-domain question answering.
Human-generated and ranked conversation dataset for training open-source chat assistants.
Open reproduction of the Orca approach with GPT-4 generated explanations for diverse tasks.
An open-source recreation of the WebText dataset used to train GPT-2 built from Reddit-shared URLs.
Large-scale instruction dataset using GPT-4 explanation traces to teach smaller models complex reasoning.
The comprehensive biomedical literature database used for training medical and scientific AI models.
An open reproduction of the LLaMA training dataset enabling transparent large language model research.
Massively filtered web dataset demonstrating that web-only data can match curated corpora quality.
The Responsible Open-science Open-collaboration Text Sources corpus used to train BLOOM.
Community-shared ChatGPT conversations used to train instruction-following open models.
A deduplicated and cleaned version of RedPajama optimized for efficient large-scale language model training.
Stanford Question Answering Dataset - the benchmark that defined modern reading comprehension evaluation.
Self-instruct generated training data from GPT-3.5 that kickstarted the open instruction-tuning movement.
A massive code dataset spanning 86 programming languages filtered for quality and licenses.
A harder successor to GLUE comprising eight challenging language understanding tasks.
A diverse 825GB English text dataset assembled from 22 high-quality sub-datasets for language model training.
The largest open code dataset with opt-out mechanisms covering 358 programming languages from GitHub.
Machine-generated dataset for adversarial and implicit hate speech detection across 13 minority groups.
A large-scale multi-turn dialogue dataset generated through ChatGPT covering diverse topics and tasks.
Large-scale speaker identification dataset extracted from YouTube interview videos.
The curated web text dataset scraped from Reddit-upvoted links used to train GPT-2.
A large-scale commonsense reasoning dataset inspired by the Winograd Schema Challenge.
6:30 AM PT, Monday through Friday. Free.