Skip to content

Scaling Laws

At a glance Scaling laws reveal that "loss decreases as a power law with parameter, data, and compute," turning "stacking scale" into a predictable engineering decision. This article covers Kaplan 2020's three power laws, Chinchilla 2022's 1:20 compute-optimal ratio, the emergent abilities debate, inference-time compute as the "fourth dimension," and practical guidance.

Scaling Laws ​

Scaling laws describe the most counter-intuitive empirical pattern of large language models: over a wide range, language model loss decreases as a smooth power law with increasing model parameters, training data, and compute — "scale" isn't magic; it's an engineering variable with a clear curve for prediction. They are the scientific basis for why "large language models" go "large," and the first decision framework for pretraining budget allocation.

One-line summary: scaling laws = answering "how much more compute/parameters/data do I need to push loss down by how much" and how to optimally distribute among the three via a single power law curve. Its scope, prerequisites, and controversies are precisely what matter for understanding it.

1. Kaplan 2020: Three Power Laws ​

OpenAI's Scaling Laws for Neural Language Models (January 2020) systematically measured Transformer language models (architecture see Transformer Architecture Deep Dive) from hundreds of millions to billions of parameters, yielding three core conclusions:

Conclusion 1: performance is primarily determined by model scale N, data volume D, and compute C — the effects of network architecture details (layer/width ratios), optimizers, activation functions, etc., are far smaller than scale itself.

Conclusion 2: loss follows power laws with N, D, and C respectively:

text
Model scale law:  L(N) ≈ (N_c / N)^α_N
Data law:         L(D) ≈ (D_c / D)^α_D
Compute law:      L(C) ≈ (C_c / C)^α_C

Where the power exponents (approximate values, the original paper has piecewise fine-grained estimates):
α_N ≈ 0.076 (for 8B WebText2 subset)
α_D ≈ 0.095
α_C ≈ 0.050

Formal meaning: every 10× scale increase reduces loss by a fixed proportion;
the smaller the exponent, the smaller the return on doubling scale.

Conclusion 3: bigger models are more "sample-efficient." For the same data volume, bigger models have lower loss — in other words, "given the same data, bigger models digest it better." This led to an important judgment at the time: when data is relatively insufficient, adding parameters is more cost-effective than adding data — GPT-3's choice of 175B parameters with ~300B tokens embodies this approach (the formal judgment of "undertraining" wouldn't come until the 2022 Chinchilla paper, see next section).

Kaplan also noted an engineering dividend: within a considerable range, small model training curves can predict big model performance — extrapolating large model loss and performance from small-scale experiments became the standard practice for pretraining budget decisions.

2. Chinchilla 2022: Compute-Optimal Ratio ​

DeepMind's Training Compute-Optimal Large Language Models (March 2022) corrected a critical blind spot: Kaplan measured with fixed data volume, without answering "if given a fixed compute budget, how should parameters and data be distributed?" Chinchilla's answer: parameters and training tokens should grow synchronously at approximately a 1:20 ratio — i.e., ~20 training tokens per parameter on average.

text
For fixed compute budget C, optimal allocation (Chinchilla model family fit):
  N_opt ∝ C^0.5        (parameter count grows as sqrt of budget)
  D_opt ∝ C^0.5        (data volume grows as sqrt of budget)
  → meaning parameters and tokens maintain a fixed ~1:20 ratio

Empirical formula (approximate):  L(N, D) ≈ E + A/N^α + B/D^β
  where α ≈ 0.34, β ≈ 0.28 (fitted from 400+ training runs across scales)

Chinchilla-70B's config: 70B parameters + 1.4 trillion tokens, designed precisely at 1:20.

The impact was enormous: almost all big models before Chinchilla were "over-parameterized and undertrained" — GPT-3 had 175B parameters with 300B tokens (~1:1.7), which by Chinchilla standards used only about 1/12 of the data it should have (should be ~3.5T). The direct consequence: the industry shifted toward "smaller parameters + more data":

ModelParametersTraining tokensParam:tokenRelative to 1:20
GPT-3 (2020)175B~300B~1:1.7Severely undertrained
Chinchilla (2022)70B1.4T1:20Compute-optimal
Llama 2 (2023)70B2T1:28Slightly overtrained
Llama 3 (2024)405B~15T~1:37Clearly overtrained
DeepSeek-V3 (2024)671B (MoE)~14.8T~1:22 (by total params)Close to optimal

"1:20" is an approximation — a starting point, not a final answer

Chinchilla's 1:20 is "training compute-optimal" — i.e., the total cost of one training run is minimized. But deploying a model also requires considering inference cost: bigger parameters → more expensive per inference. So many teams deliberately "overtrain" bigger models (like Llama 3): spend a bit more on training to get "the same capability, a smaller model, cheaper inference." Additionally, MoE, data quality differences all shift the optimal ratio; individual projects should do their own ablation studies.

The Three Curves in One Chart ​

text
Under fixed compute budget C (log-log coordinates, illustrative):
Loss L
 │
 │  ╲  N fixed, add D         (downhill along data axis, diminishing returns)
 │   ╲
 │    ╲___  D fixed, add N
 │       ╲___
 │           ╲___  N and D added synchronously at 1:20 (optimal path, steepest)
  └─────────────────────────▶ Scale (parameters/data/compute)

Intuition: the power law slopes on the three axes differ;
the optimal path = let "currently the scarcest resource" catch up with the others.

A detail that must be emphasized: the power law describes "validation loss," but products care about "downstream capabilities." Loss decreases smoothly and continuously, but capabilities may change in stair-step fashion (see "emergence" next section) — when predicting capability from loss, the same loss difference may correspond to much larger capability differences at bigger scales than at small scales. Therefore, budget decisions should monitor both curves simultaneously: the loss curve (smooth, extrapolatable) and the capability curve (steep, task-dependent).

A Sample Calculation with Power Laws (Illustrative) ​

Assuming L(N) = (N_c/N)^0.076 fitted under a certain architecture and data distribution (Kaplan approximation), going from 1B to 10B parameters (10×), loss drops by approximately 1 − 10^(-0.076) ≈ 16% (relative); from 10B to 100B, another ~16%. Every 10× scale increase yields a fixed percentage drop that gradually diminishes — this is the most practical reading of power laws: first fit your own curve's exponents from small-scale experiments, then extrapolate "how much spending buys how much loss reduction," rather than guessing.

3. Emergent Abilities: How Quantitative Change Becomes Qualitative ​

Alongside the "smooth loss decline" is the phenomenon of sudden capability jumps: some task's abilities suddenly leap from near-random to far-above-random near a certain scale threshold — Wei et al. 2022 called this emergent abilities, including few-shot arithmetic, multi-step reasoning, instruction following, code generation, etc. But the concept quickly sparked debate:

ViewCore claimRepresentative
Emergence is a real capability leapSome tasks can only predictably appear once scale crosses a threshold; small models simply cannot complete themWei et al. 2022
Emergence is a measurement artifactMeasuring "continuously improving capabilities" with non-linear metrics like "accuracy" artificially creates cliffs; switching to linear metrics (like log-loss, token-level evaluation) reveals smooth improvementSchaeffer et al. 2023

The practical significance of this debate isn't "who's right" but methodology: looking at scaling effects through a single task's accuracy gets misled by the metric itself. A more robust framework:

  • The capability axis: eval scores (emergence view, see Evaluation & Benchmarks) and the loss axis (smooth power law) are different things; for "whether something can emerge," examine each task and metric individually.
  • Unpredictability is emergence's practical definition: can you extrapolate a task's performance from small-scale before training? If not, the task has "emergence" for engineers — and that determines budget allocation risk.
  • Practical impact: emergence means "small model validation conclusions" can't be infinitely extrapolated; certain capabilities must be trained at sufficient scale to obtain.

The response to the "artifact" view is equally documented: Wei et al.'s follow-up analysis noted that emergence doesn't hold for all tasks — some tasks indeed rise smoothly with scale, others show clear cliffs, and "when emergence happens and in what form" can't be extrapolated from small-scale pretraining. More interestingly, the same ability measured differently (accuracy vs per-token loss) yields different "emergence shapes," suggesting emergence visibility depends on the metric, but the non-linear growth of the capability itself is real. For engineering, one takeaway: to verify whether a key capability can appear at big scale, small-scale experiments only provide "risk signals," not "guarantees."

A reminder for engineering

Scaling laws give "average expectations"; individual tasks may emerge late or early. For pretraining budgeting, please include both "key task performance at small scale" and "smooth extrapolation of loss curves" in decisions, rather than relying on just one metric.

4. The Fourth Dimension: Inference-Time Compute ​

Starting in 2024, the industry validated another extension axis beyond "training scale" — inference-time compute (test-time scaling): not changing the model, but improving complex task performance by letting the model "think longer" at inference time (generating chain-of-thought, self-reflection, multi-round search).

text
The paradigm of o1 / o1-mini / o3 series (OpenAI, 2024):
  Training phase: teach the model "to plan" via reinforcement learning (RL on long-CoT)
  Inference phase: allow the model to generate thousands to tens of thousands
                   of hidden CoT tokens before giving the final answer —
                   trading inference compute for accuracy
  Observation: for code/math/science tasks, accuracy rises approximately
               smoothly with inference token count — creating "inference-side scaling laws"

New question formulation: given total budget (training + inference),
  how to allocate to training vs. inference optimally? — a new Chinchilla-style question

Inference-time compute represents "scaling laws" extending from the training dimension to the usage dimension, jointly with model distillation and long context to form a "dig one more layer beyond the capability ceiling" approach. Latest frontier developments see Frontier Progress.

An observable example: the o1 series on math and code competition-level problems shows significantly improved accuracy when allowed longer chain-of-thought (more inference tokens) — this is intuitive evidence for "trading inference time for accuracy"; at the same time, it brings longer latency and higher inference cost (one reasoning chain can reach tens of thousands of tokens), so "inference-time compute returns" also obey diminishing marginal returns and must be coordinated with training-side extension. Outside o1, inference-time search (MCTS-like), multi-agent debate, etc., all fall in this category.

5. Practical Implications: Add Data or Parameters First? ​

Scaling laws aren't paper formulas; they directly guide resource allocation:

ScenarioRecommendationBasis
Fixed training budgetDesign model and data at 1:20 compute-optimal ratioChinchilla
Inference cost sensitive (high-concurrency product)Train one "too-large but stronger" model, then distill/choose a smaller one for deploymentInference budget perspective
Want to validate direction quicklySmall model + data subset for ablation, extrapolate big model performance via power lawKaplan's predictability
Data scarce / expensivePrioritize adding data (the real bottleneck for most teams)Undertraining is widespread
Assess whether a single task can meet targetsSmall-scale extrapolation + per-task "emergence threshold" risk assessmentEngineering conclusion of emergence debate

For actual budgeting, follow this sequence: ① Run a scale-staircase experiment with a small model (e.g., 0.1B–1B) + a subset of target data, fitting your team's L(N,D) curve; ② Reverse-derive required N and D from target loss (or key task metrics); ③ Candidate solutions within ~1:20 ratio, then fine-tune based on inference cost preference; ④ Maintain reproducible small-scale ablation baselines throughout as mid-run reference for large-scale training. Scaling laws' value lies precisely in turning "guess-and-stack compute" into "budget decisions backed by curves."

Don't ignore the three prerequisites of scaling laws

① Power laws hold within the range where "architecture and data distribution are fixed" — changing architecture, tokenizer, or data mix shifts the whole curve; ② Loss decline ≠ downstream capability rises synchronously (emergence can lag); ③ It's an empirical fit, not a physical law, and marginal returns will eventually decay; "blindly stacking scale" isn't a free lunch.

6. Hyperparameter Transfer at Scale (μP Primer) ​

Another line tied to "scale affects performance" is hyperparameter transfer. Traditionally, learning rate, initialization scale, etc., need re-tuning as model scale changes; maximal update parametrization (μP, Yang et al.) starts from neural network "tensor program" theory, providing parameterizations that scale correctly with width/depth, so that learning rate, initialization, etc., tuned on a small model can be directly reused on a bigger model.

text
Classic problem: LR=3e-4 tuned on 7B, can it be used on 70B directly?
  (Empirically, LR must decrease with scale — specific curve requires trial)
μP solution: separately scale weight initialization, LR, multiplicative/additive updates by width,
  so that "the impact of each step's update on output" stays at the same magnitude across widths
→ Small-model tuning results transfer to big models ("zero-shot hyperparameter transfer")
Cost: μP uses different LR scaling for embedding and output layers, slightly more complex;
     mainstream frameworks (like certain training stacks) already have built-in support

μP's significance extends scaling laws from performance prediction to training configuration prediction: engineers can cheaply tune the full training config on small scale, then scale up with confidence.

Downstream Capability Rises with Scale — Empirical Evidence ​

Beyond validation loss, the industry also continuously measures "capability-scale" curves: 200+ tasks in BIG-bench show that most tasks' "pass probability" rises steadily with model size, while a minority show a sudden cliff (emergence); MMLU (multiple-choice across 57 disciplines) and GSM8K (elementary math word problems) rose smoothly during the GPT-3 era, providing direct evidence that "loss decline translates to capability improvement." Practical takeaway: for most tasks, the loss curve is a reliable proxy for the capability curve; for the minority "emergent" tasks, capabilities must be measured directly (evaluation methods see Evaluation & Benchmarks).

7. Trade-offs and Boundaries ​

  • Diminishing returns: power laws mean marginal returns from doubling scale gradually shrink; ROI of ultra-large training must be assessed against specific business needs.
  • Data will eventually run out: the total volume of high-quality text data is finite; when curves hit data constraints, emergent returns may slow — which is why synthetic data and multimodal data attract attention.
  • Loss and capability decouple: validation loss isn't a product metric; scaling laws answer "how well trained," not "how strong the product is."
  • Architecture shifts the curve: MoE, better data, novel attention mechanisms can all shift the whole curve left — scaling laws measure "extension within the same tech route," not cross-generational comparison.
  • Power laws don't promise "infinite scaling": data distribution and architecture both have ceilings; extrapolation beyond the curve's fitting interval distorts; write the curve's valid interval into budget docs to avoid treating extrapolation as a guarantee.
  • Beyond scale is "mix": data quality, training duration, context length, inference-time compute can all be "extended" — true optimization is finding the highest marginal return across all dimensions.

Three frequent questions

  • "Only big labs can use this?" No: open-source small models + power law extrapolation for budgeting, inference-time compute for enhancement — small teams can also systematically "scale."
  • "Must 1:20 be strictly enforced?" No: it's an approximation for training compute optimality; inference budget, data quality, MoE all shift it (see MoE: Sparse Expert Models).
  • "No loss decline means no scaling?" Not necessarily: loss is a global average, capabilities are task-specific; some tasks may emerge first.

One-line summary

Scaling laws turn "big" from a slogan into a curve: loss ≈ power law × (parameters, data, compute), optimal ratio ~1:20, emergence is a phenomenon where metric and capability intertwine, inference-time compute is a new dimension. Read it, and you'll calculate before you stack.

Further Reading ​

References ​