Overview
The arXiv Papers corpus represents the complete preprint archive maintained by Cornell University since 1991. With over 2.3 million papers across physics, mathematics, and computer science, it provides the largest open collection of scientific text for training AI models.
What’s In It
The corpus contains 2.3+ million papers in LaTeX source and PDF spanning physics, mathematics, computer science, statistics, and other fields. Papers include full text, citations, equations, and figures with 15,000+ new papers added monthly.
How It’s Used
ArXiv data trains scientific language models like Galactica and SciBERT and contributes to general LLM pre-training. The LaTeX source enables math-aware tokenization. Citation networks support recommendation systems.
Controversies
Papers are posted under various licenses and bulk use for AI training may exceed what authors intended. The corpus skews heavily toward physics and CS. Papers contain errors and varying quality levels.