Dataset
RedPajama
Open recreation of the Llama training set â 1.2T tokens of curated web, books and code.
Definition
RedPajama (Together AI, 2023) reproduced the Llama-1 data recipe â Common Crawl, C4, Books, ArXiv, GitHub, Wikipedia, StackExchange â to enable transparent open-weights training. RedPajama-V2 scaled to 30T tokens with quality scores.
Common use cases
- Open Llama-style training
- Filtering experiments
- Reproducibility