Concept
Scaling Laws
Empirical relations predicting model loss as a function of parameters, data and compute.
Definition
Kaplan et al. (2020) and Chinchilla (2022) showed that pre-training loss follows clean power laws in model size, dataset size and compute. The Chinchilla update — train smaller models on more tokens — reshaped how labs allocate budget across parameters vs data.
Common use cases
- Training budgets
- Architecture choices
- Research planning