Concept

Scaling Laws

Empirical relations predicting model loss as a function of parameters, data and compute.

Definition

Kaplan et al. (2020) and Chinchilla (2022) showed that pre-training loss follows clean power laws in model size, dataset size and compute. The Chinchilla update — train smaller models on more tokens — reshaped how labs allocate budget across parameters vs data.

Common use cases

  • Training budgets
  • Architecture choices
  • Research planning

Related terms

    Scaling Laws — AI Glossary | Railwail