40 datasets tracked 10 categories
arXiv Papers Scientific

The complete arXiv preprint repository used as training data for scientific language models.

BY arXiv / Cornell University TEXT Various
SIZE 2.3M+ scientific papers
C4 Text

Colossal Clean Crawled Corpus - a filtered and deduplicated version of Common Crawl used to train T5.

BY Google TEXT ODC-BY
SIZE 750GB cleaned text
Common Crawl Text

The largest open web crawl archive serving as the foundational data source for most modern language models.

BY Common Crawl Foundation TEXT CC0 1.0
SIZE Petabytes of web data
Conceptual Captions Multimodal

Automatically harvested image-description pairs from the web for vision-language pre-training.

BY Google MULTIMODAL Custom
SIZE 12M image-caption pairs
Dolma Text

A three-trillion token open corpus curated by AI2 for training OLMo and advancing transparent LLM research.

BY AI2 TEXT ODC-BY
SIZE 3T tokens
DPO-Mix Preference

Curated mixture of human preference datasets for Direct Preference Optimization alignment training.

BY Various TEXT Various
SIZE 60K preference pairs
GSM8K Benchmark

Grade School Math 8K - a benchmark of linguistically diverse math word problems testing multi-step reasoning.

BY OpenAI TEXT MIT
SIZE 8.5K grade-school math problems
HellaSwag Benchmark

A challenging commonsense reasoning benchmark testing physical and social situation understanding.

BY AI2 TEXT MIT
SIZE 70K commonsense completion questions
HowTo100M Video

Massive instructional video dataset with narration-aligned clips for learning text-video embeddings.

BY INRIA VIDEO Apache 2.0
SIZE 136M video clips from 1.2M videos
HumanEval Benchmark

Hand-crafted Python programming challenges used to evaluate functional correctness of code generation models.

BY OpenAI CODE MIT
SIZE 164 hand-written problems
ImageNet Image

The foundational large-scale image classification dataset that catalyzed the deep learning revolution.

BY Princeton University IMAGE Custom (research)
SIZE 14M labeled images
Kinetics Video

Large-scale video understanding benchmark for human action recognition from YouTube clips.

BY DeepMind VIDEO CC-BY-4.0
SIZE 650K video clips across 700 classes
LAION-5B Image

The largest openly available image-text dataset used to train Stable Diffusion and other multimodal models.

BY LAION MULTIMODAL CC-BY-4.0
SIZE 5.85B image-text pairs
LAION-Aesthetics Image

A subset of LAION-5B filtered by predicted aesthetic quality scores for training image generation models.

BY LAION IMAGE CC-BY-4.0
SIZE 600M image-text pairs scored for aesthetics
LibriSpeech Audio

The standard benchmark for automatic speech recognition derived from LibriVox audiobooks.

BY OpenSLR AUDIO CC-BY-4.0
SIZE 1000 hours of read English speech
MMLU Benchmark

Massive Multitask Language Understanding - 57-subject benchmark measuring broad academic knowledge.

BY UC Berkeley TEXT MIT
SIZE 15.9K multiple-choice questions
Mozilla Common Voice Audio

The largest open multilingual speech dataset built through crowdsourced voice donations.

BY Mozilla AUDIO CC-0
SIZE 19K+ hours across 100+ languages
MS COCO Image

Common Objects in Context - the standard benchmark for object detection segmentation and image captioning.

BY Microsoft MULTIMODAL CC-BY-4.0
SIZE 330K images with captions
Natural Questions Benchmark

Real Google search queries paired with Wikipedia passages to evaluate open-domain question answering.

BY Google TEXT CC-BY-SA 3.0
SIZE 307K question-answer pairs
Open Assistant Conversations Instruction

Human-generated and ranked conversation dataset for training open-source chat assistants.

BY LAION TEXT Apache 2.0
SIZE 161K messages across 66K trees
OpenOrca Instruction

Open reproduction of the Orca approach with GPT-4 generated explanations for diverse tasks.

BY Open-Orca TEXT MIT
SIZE 1M+ GPT-4 augmented responses
OpenWebText Text

An open-source recreation of the WebText dataset used to train GPT-2 built from Reddit-shared URLs.

BY Aaron Gokaslan & Vanya Cohen TEXT CC0
SIZE 38GB text
Orca Instruction

Large-scale instruction dataset using GPT-4 explanation traces to teach smaller models complex reasoning.

BY Microsoft TEXT MIT
SIZE 5M instructions with reasoning traces
PubMed Scientific

The comprehensive biomedical literature database used for training medical and scientific AI models.

BY NIH / NLM TEXT Public Domain
SIZE 36M+ biomedical citations
RedPajama Text

An open reproduction of the LLaMA training dataset enabling transparent large language model research.

BY Together AI TEXT Apache 2.0
SIZE 1.2T tokens
RefinedWeb Text

Massively filtered web dataset demonstrating that web-only data can match curated corpora quality.

BY TII (Technology Innovation Institute) TEXT ODC-BY-1.0
SIZE 5T tokens from filtered web
ROOTS Text

The Responsible Open-science Open-collaboration Text Sources corpus used to train BLOOM.

BY BigScience TEXT Various
SIZE 1.6TB spanning 46 languages
ShareGPT Instruction

Community-shared ChatGPT conversations used to train instruction-following open models.

BY Community TEXT Various
SIZE 90K conversations
SlimPajama Text

A deduplicated and cleaned version of RedPajama optimized for efficient large-scale language model training.

BY Cerebras TEXT Apache 2.0
SIZE 627B tokens
SQuAD Benchmark

Stanford Question Answering Dataset - the benchmark that defined modern reading comprehension evaluation.

BY Stanford NLP TEXT CC-BY-SA 4.0
SIZE 150K question-answer pairs
Stanford Alpaca Instruction

Self-instruct generated training data from GPT-3.5 that kickstarted the open instruction-tuning movement.

BY Stanford TEXT CC-BY-NC 4.0
SIZE 52K instructions
StarCoder Training Data Code

A massive code dataset spanning 86 programming languages filtered for quality and licenses.

BY BigCode CODE Various per license
SIZE 783GB of code
SuperGLUE Benchmark

A harder successor to GLUE comprising eight challenging language understanding tasks.

BY NYU / DeepMind / others TEXT Various
SIZE Multi-task benchmark suite
The Pile Text

A diverse 825GB English text dataset assembled from 22 high-quality sub-datasets for language model training.

BY EleutherAI TEXT MIT
SIZE 825GB text
The Stack Code

The largest open code dataset with opt-out mechanisms covering 358 programming languages from GitHub.

BY BigCode / Hugging Face CODE Various
SIZE 6.4TB source code
ToxiGen Benchmark

Machine-generated dataset for adversarial and implicit hate speech detection across 13 minority groups.

BY Microsoft TEXT MIT
SIZE 274K toxic and benign examples
UltraChat Instruction

A large-scale multi-turn dialogue dataset generated through ChatGPT covering diverse topics and tasks.

BY Tsinghua University TEXT CC-BY-NC 4.0
SIZE 1.5M multi-turn dialogues
VoxCeleb Audio

Large-scale speaker identification dataset extracted from YouTube interview videos.

BY University of Oxford AUDIO CC-BY-SA 4.0
SIZE 7K+ speakers, 1M+ utterances
WebText Text

The curated web text dataset scraped from Reddit-upvoted links used to train GPT-2.

BY OpenAI TEXT Proprietary
SIZE 40GB text
WinoGrande Benchmark

A large-scale commonsense reasoning dataset inspired by the Winograd Schema Challenge.

BY AI2 TEXT Apache 2.0
SIZE 44K fill-in-the-blank problems

The AI industry,
just said.

6:30 AM PT, Monday through Friday. Free.

No referral program. No upsell. Unsubscribe is one click.