Underparameterized · grow data
αD ≈ 0.98–1.11Paper-derived range across the displayed settings. Dataset and architecture assignments are not reconstructed as individual points.
PAPER-DERIVED VALUES (Fig. 1a · top-left)CONTEXT JAMMING / EMPIRICAL LAWS OF AI · PART III
A year after measuring the slopes, Kaplan and four colleagues asked what the slopes were measuring.
Paper framework · §1–2Scaling laws said loss falls as a power law in data and parameters. They did not say why, or why the exponents take the values they do. Bahri, Dyer, Kaplan, Lee and Sharma split the phenomenon into four regimes with two mechanisms. One mechanism gives a universal exponent of one. The other gives exponents set by the intrinsic dimension of the data itself.
Variance-limited
Fluctuations shrink around a smooth limit.
01 · The open question
Kaplan measured the slope. This paper asks what the slope is measuring.
A power law is a relationship in which multiplying the resource changes loss by a predictable factor. On a plot where both axes are logarithmic, it becomes a straight line; for L ∝ X−α, the line’s slope is −α.
The measurement chapter and allocation chapter organize empirical fits. Here the authors ask four different questions:
A common shape does not guarantee a common cause. Start by asking what is scarce.
Intro · p. 1 · questions motivating the paper
02 · Two resources, two limits
Grow one resource while holding the other fixed. The result depends on their hierarchy.
D counts training examples; P counts student features in the solvable linear model. Underparameterized means D ≫ P; overparameterized means P ≫ D. When the growing resource is already abundant, the model approaches a smooth limiting prediction. When that resource is scarce, it still resolves new structure.
§1.1–1.2 · §2 · Fig. 1a
Variance-limited
Fluctuations shrink around a smooth limit.
| Resource grown | Grown resource abundant Variance-limited | Grown resource scarce Resolution-limited |
|---|---|---|
| Grow data D | D ≫ P · αD = 1 Dataset fluctuations shrink around the infinite-data limit. Fig. 1a · top-left | P ≫ D · αD ∝ 1/d Training points resolve finer intrinsic geometry. Fig. 1a · top-right |
| Grow model P / width w | P ≫ D · αW = 1 Width fluctuations shrink around the infinite-width limit. Fig. 1a · bottom-right | D ≫ P · α ∝ 1/d More model capacity resolves finer target structure. Fig. 1a · bottom-left |
Width is not parameter count. In the deep-network scaling described here, w ∝ √P: a width loss-gap exponent of one corresponds to P−1/2. In linear models, variance-limited feature scaling is P−1. The resolution panels in Fig. 1a also measure width; they do not supply αP directly.
§1.1 · §2.1.2 · §2.3.1
03 · The boring exponent
The network’s prediction fluctuates around a limit. Smooth loss turns shrinking variance into a shrinking loss gap.
Concentration means repeated training runs give increasingly similar predictions as a resource grows. Dataset fluctuations have variance of order 1/D; width fluctuations have variance of order 1/w. Theorem 1 transfers that order to expected loss when the centered moments and loss satisfy its conditions.
§2.1 · Theorem 1 · App. B–C
Empirical fit αD ≈ 0.98–1.11 · PAPER-DERIVED VALUES (Fig. 1a · top-left)
Empirical fit αW ≈ 0.98–1.03 · PAPER-DERIVED VALUES (Fig. 1a · bottom-right)
App. C.3 · scope distinction from the language-model fits
04 · Carving the manifold
A manifold is a space described locally by a smaller number of independent coordinates. Its intrinsic dimension is the number of directions the model needs to resolve.
An image has many pixel coordinates, but those coordinates need not vary independently. High input dimension alone says little about intrinsic manifold dimension. Under the paper’s compact-manifold assumptions, a fresh point gets closer to its nearest training point as D grows.
§2.2.1–2.2.2 · Theorems 2–3 · App. D
If both target and student change at bounded rates, matching a training target constrains error nearby. Such a rate bound is called Lipschitz continuity. Theorems 2–3 turn it into upper bounds on test loss, for sampled data or a model that interpolates on a parameter-controlled set of points.
Theorems 2–3
Solid: selected intrinsic d = 4. Dashed: intrinsic d = 1, 2, 4, 8, 16.
05 · From bounds to estimates
An upper bound says how bad the error can be. An estimate says how the error usually behaves. The authors make that extra step explicitly.
Expand loss near the closest training point. If the first surviving local term has order n, raising the spacing to that power gives n/d in the loss exponent. A sufficiently accurate piecewise-linear fit can push the first nonzero term to fourth order. That motivates αD ≈ 4/d.
§2.2.3 · bounds as estimates
When teacher–student input dimension is controlled, 4/αD tracks it. That experiment supports the estimate in a clean setting. For convolutional networks and Wide ResNets on standard datasets, the relationship is less clear. The clean test does not make intrinsic dimension an observed constant for arbitrary real data.
§3.1 · Fig. 1b
06 · A model you can solve
Fix the features. Learn their weights. The network becomes a problem that can be solved exactly.
A feature is a fixed function of an input; a random-feature model samples such functions before training and learns only their linear combination. A teacher creates targets from a feature pool. The student receives a P-dimensional projection of that pool and D examples, then reaches the global optimum of mean squared error—the average squared prediction error—from zero initial weights.
This construction has a connection to wide networks. A Gaussian process describes random functions through Gaussian-distributed values; suitable infinite-width networks admit that description. A neural tangent kernel describes similarities through parameter gradients; under the corresponding wide-network training limit it stays fixed. Those limits motivate the model without making it a description of every finite network.
§2.3 · Eq. 1 · App. E
07 · The spectrum sets the slope
A kernel measures similarity between inputs. Its eigenvalues rank the strength of independent modes the student could learn.
The corresponding feature covariance measures second moments between features. Both carry the same spectral information. The eigenvalue spectrum is the ordered list of mode strengths; spectral decay describes how quickly they fall with rank. A fast-falling spectrum leaves less important structure unresolved.
§2.3.2–2.3.4 · Eqs. 4–6
Empirical fit αK ≈ 0.34–1.25 · PAPER-DERIVED VALUES (Fig. 2b · top)
A Ct kernel has t continuous derivatives. Smoothness limits how heavy its eigenvalue tail can be. On an intrinsic d-dimensional space, the bound involves t/d; if the spectrum saturates that bound, it gives a dimension-dependent exponent.
§2.3.4 · smooth kernels on the d-torus
Hatched: above the normalized bound. A pure power law must decay at least this fast.
Smooth functions on low-dimensional manifolds have spectra that fall fast. Fast-falling spectra mean steep scaling. That is the bridge from geometry to exponent—with the bound-saturation assumption kept in view.
§2.3.4 · saturation assumption
08 · The duality
Project onto random features, or onto random training points. The exact linear loss expressions exchange one projection for the other.
With a shared power-law spectrum, the resolution-limited loss exponent for features at abundant data equals the exponent for data at abundant features: αP = αD = αK. Fig. 2b tests the relationship with pooled MNIST and random ReLU features. Appendix F also compares loss curves using frozen EfficientNet-B5 features under low and tuned regularization.
Double descent means error can rise near interpolation and fall again as a resource grows. Sample-wise double descent is one realization of this duality, rather than a separate general explanation of every scaling curve.
§2.4 · S37–S38 · Fig. 2b · Fig. S6
09 · In the wild
The empirical sweep separates what stays nearly fixed from what moves.
Tests include MNIST, FashionMNIST, CIFAR-10, CIFAR-100 and SVHN; fully connected, convolutional and Wide ResNet architectures; ReLU and Erf activations; mean squared error and cross-entropy; and changes to stochastic-gradient batches. These settings support the taxonomy, with variance-limited exponents near one and resolution-limited exponents that depend on data and model.
§3.1–3.3 · App. A, F · Figs. 1a, S1–S3, S7
Paper-derived range across the displayed settings. Dataset and architecture assignments are not reconstructed as individual points.
PAPER-DERIVED VALUES (Fig. 1a · top-left)Paper-derived range across the displayed settings. Dataset and architecture assignments are not reconstructed as individual points.
PAPER-DERIVED VALUES (Fig. 1a · top-right)Paper-derived range across the displayed settings. Dataset and architecture assignments are not reconstructed as individual points.
PAPER-DERIVED VALUES (Fig. 1a · bottom-left)Paper-derived range across the displayed settings. Dataset and architecture assignments are not reconstructed as individual points.
PAPER-DERIVED VALUES (Fig. 1a · bottom-right)A linear classifier on EfficientNet-B5 embeddings of CIFAR-10 exhibits the same taxonomy.
App. F · Figs. S6–S710 · What moves the exponent
Superclassing merges fine labels into broader categories. In these tests, that changes loss level while leaving the data exponent similar.
Gaussian noise changes the inputs themselves, and the exponent falls as noise grows. The instrument displays the reported superclass band and the verified noise endpoints, without inventing intermediate measurements.
§3.3 · Fig. 3
The authors interpret the contrast as evidence that networks learn input-manifold structure independently of the precise classification task, resembling unsupervised learning. The measured contrast is evidence; that mechanism remains an interpretation.
§4 · interpretation of Fig. 3
§3.3 · App. A.8 · Fig. S5
11 · Scope and failure
Asymptotic means the large-resource limit. A result in that limit needs a real hierarchy before it becomes a useful diagnosis.
§4 · §4.1 · App. C.3, F
When D and P become comparable, the asymptotic predictions lose their premise. The measured curves bend near that region.
§4.1 · Figs. 1a/2aThere is no precise manifold definition for a real dataset. The paper estimates intrinsic dimension through nearest-neighbor distances in a trained final embedding.
§4.1Finite-network kernels evolve during training. The fixed-feature model does not capture that evolution.
§4Teacher-generated targets scale better than real labels in the appendix experiments.
App. F · Fig. S8Non-smooth or unbounded losses can violate the variance-limited exponent.
App. C.312 · Context Jamming extension
Nothing in this section is established by the paper. The paper studies vision datasets, linear and random-feature models, and wide networks. Language models are not its subject.
The comparison asks what its mechanism might suggest elsewhere. It supplies hypotheses to test, with assumptions that can fail.
Context Jamming extension · paper scope: §1–4
Effective dimension of the text or representation distribution an LLM must resolve
Small language-model loss exponents as a possible signature of high effective dimension
A corner where the growing resource is already abundant relative to the other resource
A comparison with the token/parameter symmetry in language-model loss fits
Pretraining as learning input structure beyond a particular downstream label task
A point of contact with the BIG-bench and emergence discussion
Preset source: Kaplan · data loss · Appendix A · Tables 4–5
For a further pointer: Sharma and Kaplan, Scaling laws from the data manifold dimension, JMLR (2022). The Chinchilla chapter provides the allocation comparison; Dyer’s profile provides context for the emergence discussion. Neither establishes the analogies above.
Related work · reference [10], not this paper
13 · Epistemic ledger
Theorem 1’s concentration result. Theorems 2–3 as upper bounds. Exact random-feature losses and projection duality. Empirical examples of all four regimes.
Using bounds as estimates, including 4/d. Networks learning input geometry. A physics-style research program. Emergent abilities as an open question.
The bounded language-model mappings, conditional dimension calculator and falsifiable research questions in Section 12.
14 · Glossary
15 · Conclusion
Two mechanisms sit underneath the straight lines. Concentration gives a universal asymptotic correction. Resolution gives exponents tied to intrinsic geometry and, in fixed-feature models, spectral decay.
The curves in Part I and the allocation in Part II now invite another question: what structure is the limiting resource still resolving? An exponent can hint at a mechanism, provided the hierarchy, model and assumptions match.
§1–4 · synthesis of the framework
The authors pursue a style of theory that combines simple models, contact with realistic systems and experimental checks, drawing on the methodology of physics. Their framework leaves emergent abilities as an open question.
§5 · outlook
16 · Primary record
Bahri, Y., Dyer, E., Kaplan, J., Lee, J., & Sharma, U. Explaining Neural Scaling Laws. arXiv:2102.06701 [cs.LG] · v1: 12 February 2021 · v2: 29 April 2024. Published in PNAS 121(27), e2311878121 (2024). DOI: 10.1073/pnas.2311878121.
Yasaman Bahri · Ethan Dyer · Jared Kaplan · Jaehoon Lee · Utkarsh Sharma
Google DeepMind, Mountain View (Bahri, Dyer, Lee); Physics and Astronomy, Johns Hopkins University (Kaplan, Sharma).
All authors contributed to all aspects of the work. Part of Sharma’s work was completed during a Google internship. Kaplan and Sharma were supported in part by Open Philanthropy.
The paper combines theoretical derivations in fixed-feature models with empirical vision experiments. This explainer reads v2; no underlying training-run data were recovered or invented.
Primary record · attached arXiv v2; journal fields verified against arXiv
| Section | Paper anchor | Figure / appendix | Equations | Fidelity |
|---|---|---|---|---|
| 01 | Intro p.1, refs [1–4] | — | — | Text |
| 02 | §1.1–1.2, §2 | Fig. 1a | — | Conceptual |
| 03 | §2.1, Thm 1 | Fig. 1a TL/BR; App. B, C, C.3 | Thm 1 | Illustrative + paper values |
| 04 | §2.2.1–2.2.2 | — ; App. D | Thms 2–3 | Illustrative + calculated |
| 05 | §2.2.3, §3.1 | Fig. 1b | L ∝ D^(−n/d) | Adapted reconstruction |
| 06 | §2.3 | — ; App. E | Eqs. 1–3, S35–S38 | Diagram |
| 07 | §2.3.2–2.3.4 | Fig. 2b top; App. E.1 | Eqs. 4–6, S46–S47 | Calculated |
| 08 | §2.4 | Fig. 2b bottom; Figs. S6 | S37–S38 | Illustrative |
| 09 | §3.1–3.3; App. A, F | Figs. 1a, S1–S3, S7; Table 1 | — | Paper values |
| 10 | §3.3; App. A.8 | Fig. 3; Fig. S5 | — | Paper values |
| 11 | §4, §4.1; App. C.3, F | Figs. 1a, 2a, S8 | — | Conceptual |
| 12 | §5 (outlook) + Context Jamming | — | — | Extension |
| 13 | §2–5 + editorial synthesis | — | — | Text / epistemic classification |
| 14 | §1–4; App. F | — | — | Definitions |
| 15 | §4–5 + author context | — | — | Text |
| 16 | Full primary record and appendix | — | — | Citation map |
| Claim | Anchor | Class | Caveat |
|---|---|---|---|
| Four regimes: variance- vs resolution-limited, for D and for P | §1, §2, Fig. 1a, Fig. 2a | DIRECT RESULT | Shown on vision datasets and linear/random-feature/pretrained-feature models |
| Variance-limited: L − L∞ ∝ 1/x, x = D or width (√P) in deep nets; D or P in linear models | §1.1, §2.1, Thm 1, App. B–C | DERIVATION | Requires sufficiently smooth loss; App. C.3 gives non-smooth counterexamples |
| Measured variance-limited exponents of about 1 (αD ≈ 0.98–1.11, αW ≈ 0.98–1.03) | Fig. 1a top-left / bottom-right legends | EMPIRICAL FIT | Varies across dataset, architecture (FC, CNN; ReLU/Erf), batch size, loss (MSE/CE) |
| Resolution-limited bounds: L(D) = O(D^(−1/d)), L(P) = O(P^(−1/d)) | §2.2.1–2.2.2, Thms 2–3, App. D | DERIVATION | Lipschitz assumptions; data on a compact d-manifold |
| Bounds as estimates: L ∝ D^(−n/d), n ≥ 2; n ≥ 4 for piecewise-linear fits | §2.2.3 | Hypothesis, not theorem | |
| Teacher-student: 4/αD tracks input dimension d | §3.1, Fig. 1b bottom | EMPIRICAL FIT | Standard datasets (CNN, WRN) relationship "less clear" |
| Resolution-limited measured exponents (αD ≈ 0.26–0.58, αW ≈ 0.34–0.62 on standard datasets) | Fig. 1a top-right / bottom-left legends | EMPIRICAL FIT | Data- and model-dependent by design |
| Linear random-feature teacher-student model; exact L(P), L(D) | §2.3, Eqs. 1–3, App. E | DERIVATION | Models trained to the global optimum with MSE; weights initialized at zero |
| Power-law spectrum λi ∝ i^(−(1+αK)) ⇒ L(D) ∝ D^(−αK), L(P) ∝ P^(−αK) | §2.3.2–2.3.3, Eqs. 4–6, App. E.1 | DERIVATION | Eq. 6 intuition uses the large-αK simplification; general case via replica methods |
| αD ≈ αP ≈ αK empirically (pooled MNIST, random ReLU features) | Fig. 2b | EMPIRICAL FIT | Random features, not trained networks |
| Smooth (C^t) kernels on d-dim manifold: λn ≲ n^(−(1+t/d)) ⇒ αK ∝ 1/d | §2.3.4 (d-torus example) | DERIVATION | For a pure power law, αK ≥ t/d; equality and the dimension estimate require saturation of the bound. |
| Duality between model-size and dataset-size scaling | §2.4, Eqs. S37–S38, Fig. S6 | DERIVATION + EMPIRICAL | Exact for linear models; sample-wise double descent is one realization |
| Pretrained EfficientNet-B5 features show all four regimes | App. F, Figs. S6–S7 | EMPIRICAL FIT | Linear classifier on frozen features (CIFAR-10) |
| Superclassing CIFAR-100 leaves αD ≈ 0.38–0.42 | §3.3, Fig. 3 left | EMPIRICAL FIT | WRN-28-10 |
| Input noise lowers αD (≈0.58 with no noise → ≈0.29 at noise stddev 0.2) | §3.3, Fig. 3 right | EMPIRICAL FIT | Show endpoints; intermediate pairings only if verified against the figure |
| Width matters up to a saturating width; depth effect milder | §3.3, App. A.8, Fig. S5 | EMPIRICAL FIT | WRN-d-k on CIFAR-10 |
| Networks may learn the input manifold "akin to unsupervised learning" | §4 | Speculative | |
| Asymptotic theory; hierarchy breakdown; manifold proxy; feature learning not modeled | §4.1, §4 | LIMITATION | — |
| Teacher targets scale better than real labels | App. F, Fig. S8 | EMPIRICAL FIT | Gap between solvable model and real tasks |
| Outlook: emergent abilities as open question; physics methodology | §5 | — |
Curves are deterministic calculations or clearly labelled conceptual reconstructions. The spectral calculator sums a finite toy spectrum; its integral approximation extends to infinity. Eq. 6 is shown only up to a constant because its printed prefactor conflicts with the integral. The regime table states the actual D/P hierarchy, and the smoothness control keeps the bound’s direction explicit.
Visualization method