Dataset
The Pile
EleutherAI's 825GB diverse text dataset spanning 22 sources, used to train GPT-Neo/J.
Definition
The Pile aggregates Common Crawl plus high-quality sources â PubMed, GitHub, ArXiv, books â into 825GB of text. It was the open community's answer to GPT-3's training data and powered GPT-Neo, GPT-J and Pythia.
Common use cases
- Open LLM training
- Reproducible pre-training
- Domain studies