Dataset

RedPajama

Open recreation of the Llama training set — 1.2T tokens of curated web, books and code.

Definition

RedPajama (Together AI, 2023) reproduced the Llama-1 data recipe — Common Crawl, C4, Books, ArXiv, GitHub, Wikipedia, StackExchange — to enable transparent open-weights training. RedPajama-V2 scaled to 30T tokens with quality scores.

Common use cases

  • Open Llama-style training
  • Filtering experiments
  • Reproducibility

Related terms

    RedPajama — AI Glossary | Railwail