Dataset
RedPajama
Open recreation of the Llama training set — 1.2T tokens of curated web, books and code.
Definition
RedPajama (Together AI, 2023) reproduced the Llama-1 data recipe — Common Crawl, C4, Books, ArXiv, GitHub, Wikipedia, StackExchange — to enable transparent open-weights training. RedPajama-V2 scaled to 30T tokens with quality scores.
Common use cases
- Open Llama-style training
- Filtering experiments
- Reproducibility