Dataset

The Pile

EleutherAI's 825GB diverse text dataset spanning 22 sources, used to train GPT-Neo/J.

Definition

The Pile aggregates Common Crawl plus high-quality sources — PubMed, GitHub, ArXiv, books — into 825GB of text. It was the open community's answer to GPT-3's training data and powered GPT-Neo, GPT-J and Pythia.

Common use cases

  • Open LLM training
  • Reproducible pre-training
  • Domain studies

Related terms

    The Pile — AI Glossary | Railwail