Home · Datasets · arXiv Papers
DATASET

arXiv Papers

The complete arXiv preprint repository used as training data for scientific language models.

TARGET QUERY arxiv dataset · ~5K/mo
SIZE
2.3M+ scientific papers
CREATOR
arXiv / Cornell University
MODALITY
text
LICENSE
Various
RELEASED
1991-08
OVERVIEW Updated 2026-05-17

Overview

The arXiv Papers corpus represents the complete preprint archive maintained by Cornell University since 1991. With over 2.3 million papers across physics, mathematics, and computer science, it provides the largest open collection of scientific text for training AI models.

What’s In It

The corpus contains 2.3+ million papers in LaTeX source and PDF spanning physics, mathematics, computer science, statistics, and other fields. Papers include full text, citations, equations, and figures with 15,000+ new papers added monthly.

How It’s Used

ArXiv data trains scientific language models like Galactica and SciBERT and contributes to general LLM pre-training. The LaTeX source enables math-aware tokenization. Citation networks support recommendation systems.

Controversies

Papers are posted under various licenses and bulk use for AI training may exceed what authors intended. The corpus skews heavily toward physics and CS. Papers contain errors and varying quality levels.