CONTEXT JAMMING

Field notes from inside the context window.

CONTEXT JAMMING / EMPIRICAL LAWS OF AI · PART III

WHY THE LINESARE STRAIGHT

A year after measuring the slopes, Kaplan and four colleagues asked what the slopes were measuring.

Paper framework · §1–2

Scaling laws said loss falls as a power law in data and parameters. They did not say why, or why the exponents take the values they do. Bahri, Dyer, Kaplan, Lee and Sharma split the phenomenon into four regimes with two mechanisms. One mechanism gives a universal exponent of one. The other gives exponents set by the intrinsic dimension of the data itself.

CONCEPTUAL RECONSTRUCTION

Which resource is scarce?

Regime map: log data by log featuresAbove P equals D, features outnumber samples. Below it, samples outnumber features. Drag the marker or use the labelled sliders and region buttons.10^110^210^310^410^5P ≫ DD ≫ PLOG DATA D →LOG FEATURES P →→

Variance-limited

Fluctuations shrink around a smooth limit.

PredictionαD = 1
D / P100×
Prediction: αD = 1. D / P: 100×LOSS GAP versus DLogarithmic axes. Illustrative slope −1. 1e+11e-13e+13e-21e+21e-23e+23e-31e+31e-3D · LOG →LOSS GAP · LOG →
§1–2 · Fig. 1a. Boundaries are asymptotic, not sharp lines. P counts features here; deep networks require the width qualification.

01 · The open question

A forecast without a reason.

Direct result / derivation

Kaplan measured the slope. This paper asks what the slope is measuring.

A power law is a relationship in which multiplying the resource changes loss by a predictable factor. On a plot where both axes are logarithmic, it becomes a straight line; for L ∝ X−α, the line’s slope is −α.

The measurement chapter and allocation chapter organize empirical fits. Here the authors ask four different questions:

  1. What sets the exponent’s value?
  2. Is there a taxonomy of distinct regimes?
  3. Are data and parameter exponents related?
  4. Which parts are universal?

A common shape does not guarantee a common cause. Start by asking what is scarce.

Intro · p. 1 · questions motivating the paper

02 · Two resources, two limits

Every scaling law has a bottleneck.

Direct result / derivation

Grow one resource while holding the other fixed. The result depends on their hierarchy.

D counts training examples; P counts student features in the solvable linear model. Underparameterized means D ≫ P; overparameterized means P ≫ D. When the growing resource is already abundant, the model approaches a smooth limiting prediction. When that resource is scarce, it still resolves new structure.

§1.1–1.2 · §2 · Fig. 1a

CONCEPTUAL RECONSTRUCTION

Which resource is scarce?

Regime map: log data by log featuresAbove P equals D, features outnumber samples. Below it, samples outnumber features. Drag the marker or use the labelled sliders and region buttons.10^110^210^310^410^5P ≫ DD ≫ PLOG DATA D →LOG FEATURES P →→

Variance-limited

Fluctuations shrink around a smooth limit.

PredictionαD = 1
D / P100×
Prediction: αD = 1. D / P: 100×LOSS GAP versus DLogarithmic axes. Illustrative slope −1. 1e+11e-13e+13e-21e+21e-23e+23e-31e+31e-3D · LOG →LOSS GAP · LOG →
§1–2 · Fig. 1a. Boundaries are asymptotic, not sharp lines. P counts features here; deep networks require the width qualification.
Direct result / derivation Four regimes · §1–2 · Fig. 1a
Resource grownGrown resource abundant
Variance-limited
Grown resource scarce
Resolution-limited
Grow data DD ≫ P · αD = 1

Dataset fluctuations shrink around the infinite-data limit.

Fig. 1a · top-left
P ≫ D · αD ∝ 1/d

Training points resolve finer intrinsic geometry.

Fig. 1a · top-right
Grow model P / width wP ≫ D · αW = 1

Width fluctuations shrink around the infinite-width limit.

Fig. 1a · bottom-right
D ≫ P · α ∝ 1/d

More model capacity resolves finer target structure.

Fig. 1a · bottom-left
Direct result / derivation

Width is not parameter count. In the deep-network scaling described here, w ∝ √P: a width loss-gap exponent of one corresponds to P−1/2. In linear models, variance-limited feature scaling is P−1. The resolution panels in Fig. 1a also measure width; they do not supply αP directly.

§1.1 · §2.1.2 · §2.3.1

03 · The boring exponent

The universal exponent is the uninteresting one.

Direct result / derivation

The network’s prediction fluctuates around a limit. Smooth loss turns shrinking variance into a shrinking loss gap.

Concentration means repeated training runs give increasingly similar predictions as a resource grows. Dataset fluctuations have variance of order 1/D; width fluctuations have variance of order 1/w. Theorem 1 transfers that order to expected loss when the centered moments and loss satisfy its conditions.

§2.1 · Theorem 1 · App. B–C

ILLUSTRATIVE MODEL

Watch fluctuations concentrate

Fixed normal quantiles · spread σ

Forty illustrative outputs about their limitThe same fixed normal quantiles are rescaled by sigma equal to one over square root of the resource.f∞ · DASHED LIMIT

Smooth loss: solid · non-smooth: dashed

LOSS GAP versus DLogarithmic axes. Selected loss slope −1. 1e+11e-11e+21e-21e+31e-31e+41e-41e+51e-5D · LOG →LOSS GAP · LOG →
Variance σ²1.000e-2
Calculated loss gap1.000e-2
Local log slope-1
Variance σ²: 1.000e-2. Calculated loss gap: 1.000e-2. Local log slope: -1

Empirical fit αD ≈ 0.98–1.11 · PAPER-DERIVED VALUES (Fig. 1a · top-left)

Empirical fit αW ≈ 0.98–1.03 · PAPER-DERIVED VALUES (Fig. 1a · bottom-right)

§2.1 · Theorem 1; App. C.3. The absolute-value example illustrates a failure of smoothness at the limiting target.
Direct result / derivation

App. C.3 · scope distinction from the language-model fits

04 · Carving the manifold

More data is finer resolution.

Direct result / derivation

A manifold is a space described locally by a smaller number of independent coordinates. Its intrinsic dimension is the number of directions the model needs to resolve.

An image has many pixel coordinates, but those coordinates need not vary independently. High input dimension alone says little about intrinsic manifold dimension. Under the paper’s compact-manifold assumptions, a fresh point gets closer to its nearest training point as D grows.

§2.2.1–2.2.2 · Theorems 2–3 · App. D

Direct result / derivation

If both target and student change at bounded rates, matching a training target constrains error nearby. Such a rate bound is called Lipschitz continuity. Theorems 2–3 turn it into upper bounds on test loss, for sampled data or a model that interpolates on a parameter-controlled set of points.

Theorems 2–3

CALCULATED / ILLUSTRATIVE GEOMETRY

A point gets closer in fewer intrinsic dimensions

Intrinsic d = 1 · curve in a plane

One-dimensional manifold embedded in two dimensionsDeterministic training points on a curve. A square test point is linked to its nearest sampled coordinate.

Intrinsic d = 2 · surface patch

Two-dimensional manifold patchThe same number of Halton samples covers a patch. A square test point is linked to its nearest neighbor.
TYPICAL SPACING versus DLogarithmic axes. Intrinsic d = 1; Intrinsic d = 2; Intrinsic d = 4; Intrinsic d = 8; Intrinsic d = 16. 4e+09e-12e+12e-16e+13e-23e+25e-31e+31e-3D · LOG →TYPICAL SPACING · LOG →

Solid: selected intrinsic d = 4. Dashed: intrinsic d = 1, 2, 4, 8, 16.

Typical spacing D^(−1/d)0.42045
Bound slope-0.25
Data to halve spacing×16
Typical spacing D^(−1/d): 0.42045. Bound slope: -0.25. Data to halve spacing: ×16
§2.2 · Theorems 2–3. Constants omitted. Geometry shows intrinsic d = 1 and 2; higher intrinsic dimensions are numerical.

05 · From bounds to estimates

Why four over d.

Author interpretation / hypothesis

An upper bound says how bad the error can be. An estimate says how the error usually behaves. The authors make that extra step explicitly.

Expand loss near the closest training point. If the first surviving local term has order n, raising the spacing to that power gives n/d in the loss exponent. A sufficiently accurate piecewise-linear fit can push the first nonzero term to fourth order. That motivates αD ≈ 4/d.

§2.2.3 · bounds as estimates

ILLUSTRATIVE MODEL

Match the teacher between training points

Analytic teacher and piecewise-linear studentSolid teacher curve is a fixed sum of sinusoids. Dashed student connects exact targets at the selected number of equally spaced knots.x ∈ [0,1] · SOLID TEACHER / DASHED STUDENT
Maximum deviation0.26338
RMS deviation0.16838
Maximum deviation: 0.26338. RMS deviation: 0.16838
Adapted reconstruction of Fig. 1b. The paper uses trained networks in four input dimensions. Here dataset labels select 3, 5 or 9 knots, not actual training runs.
CALCULATED · AUTHOR HYPOTHESIS

Read the estimate in both directions

Dimension estimate: four over exponentCalculated reference line four divided by alpha equals intrinsic dimension. A lighter line shows two divided by alpha under the same four-over-d prediction. The chosen dimension and supplied exponent give square and circular markers.4/αD = d2/αD = d/2INTRINSIC DIMENSION d →4/αD OR 2/αD →
Predicted αD = 4/d0.5
Implied intrinsic d = 4/αD8
Predicted αD = 4/d: 0.5. Implied intrinsic d = 4/αD: 8
Fig. 1b · §3.1. Teacher–student models track 4/αD; for CNN/WRN on standard datasets the authors report a less clear relationship. No measured scatter points are plotted.
Empirical fit

When teacher–student input dimension is controlled, 4/αD tracks it. That experiment supports the estimate in a clean setting. For convolutional networks and Wide ResNets on standard datasets, the relationship is less clear. The clean test does not make intrinsic dimension an observed constant for arbitrary real data.

§3.1 · Fig. 1b

06 · A model you can solve

Shrink the network to something exact.

Direct result / derivation

Fix the features. Learn their weights. The network becomes a problem that can be solved exactly.

A feature is a fixed function of an input; a random-feature model samples such functions before training and learns only their linear combination. A teacher creates targets from a feature pool. The student receives a P-dimensional projection of that pool and D examples, then reaches the global optimum of mean squared error—the average squared prediction error—from zero initial weights.

This construction has a connection to wide networks. A Gaussian process describes random functions through Gaussian-distributed values; suitable infinite-width networks admit that description. A neural tangent kernel describes similarities through parameter gradients; under the corresponding wide-network training limit it stays fixed. Those limits motivate the model without making it a description of every finite network.

§2.3 · Eq. 1 · App. E

CONCEPTUAL DIAGRAM · Eq. 1

Two projections constrain the student

Random-feature teacher–student constructionThe teacher feature pool is projected to P student features. D samples from the intrinsic data manifold provide training targets. Both constrain a linear student trained with mean squared error.TEACHER FEATURE POOL SF = Σ ωM FMP STUDENT FEATURESf = Σ θμ fμINTRINSIC DATA MANIFOLDD SAMPLED POINTSTRAIN LINEAR WEIGHTS θGLOBAL MSE OPTIMUM→→↓
§2.3 · Fixed features; the student learns their weights. Diagram simplifies the model, not its training assumption.
Show the exact losses · Eqs. 2–3

07 · The spectrum sets the slope

The slope was hiding in the eigenvalues.

Direct result / derivation

A kernel measures similarity between inputs. Its eigenvalues rank the strength of independent modes the student could learn.

The corresponding feature covariance measures second moments between features. Both carry the same spectral information. The eigenvalue spectrum is the ordered list of mode strengths; spectral decay describes how quickly they fall with rank. A fast-falling spectrum leaves less important structure unresolved.

§2.3.2–2.3.4 · Eqs. 4–6

CALCULATED

The unresolved tail becomes the loss

Eigenvalues λi · shaded ranks i > D

λi versus EIGENVALUE RANK iLogarithmic axes. Calculated spectrum. The shaded interval marks the selected region.1e+01e+07e+04e-24e+12e-33e+26e-52e+32e-6EIGENVALUE RANK i · LOG →λi · LOG →

Loss: finite sum (solid), integral (dashed)

TAIL SUM versus DLogarithmic axes. Finite tail sum; Infinite integral approximation. 1e+13e-13e+11e-11e+26e-23e+22e-21e+31e-2D · LOG →TAIL SUM · LOG →
Finite tail sum5.622e-2
Infinite integral D^(−αK)/αK5.687e-2
Fitted local slope-0.702
Finite tail sum: 5.622e-2. Infinite integral D^(−αK)/αK: 5.687e-2. Fitted local slope: -0.702

Empirical fit αK ≈ 0.34–1.25 · PAPER-DERIVED VALUES (Fig. 2b · top)

Eqs. 4–6 · App. E.1. The finite sum stops at 100,000 modes. The displayed integral approximation extends to infinity. Eq. 6’s tail intuition uses large αK; replica methods give the general result.
Direct result / derivation

A Ct kernel has t continuous derivatives. Smoothness limits how heavy its eigenvalue tail can be. On an intrinsic d-dimensional space, the bound involves t/d; if the spectrum saturates that bound, it gives a dimension-dependent exponent.

§2.3.4 · smooth kernels on the d-torus

CALCULATED

Smoothness limits how heavy the spectrum can be

NORMALIZED λn BOUND versus EIGENVALUE RANK nLogarithmic axes. Smooth-kernel spectral upper bound. 1e+01e+06e+01e-13e+11e-22e+22e-31e+32e-4EIGENVALUE RANK n · LOG →NORMALIZED λn BOUND · LOG →

Hatched: above the normalized bound. A pure power law must decay at least this fast.

Bound decay exponent 1 + t/d1.25
Minimum αK for a pure power law0.25
Loss slope if saturated-0.25
Bound decay exponent 1 + t/d: 1.25. Minimum αK for a pure power law: 0.25. Loss slope if saturated: -0.25
§2.3.4. Unknown constants are normalized to one. Hatching marks values above that normalized bound; this is an asymptotic order bound. Equality of the power-law exponent needs saturation.
Author interpretation / hypothesis

Smooth functions on low-dimensional manifolds have spectra that fall fast. Fast-falling spectra mean steep scaling. That is the bridge from geometry to exponent—with the bound-saturation assumption kept in view.

§2.3.4 · saturation assumption

08 · The duality

Adding data and adding features are mirror moves.

Direct result / derivation

Project onto random features, or onto random training points. The exact linear loss expressions exchange one projection for the other.

With a shared power-law spectrum, the resolution-limited loss exponent for features at abundant data equals the exponent for data at abundant features: αP = αD = αK. Fig. 2b tests the relationship with pooled MNIST and random ReLU features. Appendix F also compares loss curves using frozen EfficientNet-B5 features under low and tuned regularization.

Double descent means error can rise near interpolation and fall again as a resource grows. Sample-wise double descent is one realization of this duality, rather than a separate general explanation of every scaling curve.

§2.4 · S37–S38 · Fig. 2b · Fig. S6

ILLUSTRATIVE MODEL · LINEAR DUALITY

One spectrum. Mirror resource moves.

Underparameterized: D ≫ P

LOSS TAIL versus FEATURES PLogarithmic axes. Same calculated tail. 1e+13e-13e+11e-11e+25e-23e+22e-21e+31e-2FEATURES P · LOG →LOSS TAIL · LOG →

Overparameterized: P ≫ D

LOSS TAIL versus DATA DLogarithmic axes. Same calculated tail. 1e+13e-13e+11e-11e+25e-23e+22e-21e+31e-2DATA D · LOG →LOSS TAIL · LOG →
Fitted αP0.702
Fitted αD0.702
αK (asymptotic)0.7
αD vs αP within ±0.01Equal
Fitted αP: 0.702. Fitted αD: 0.702. αK (asymptotic): 0.7. αD vs αP within ±0.01: Equal
§2.4 · S37–S38 · Fig. 2b / S6. Exact duality holds in the linear/random-feature construction. These finite tail sums illustrate it; fitted local exponents include truncation effects.

09 · In the wild

Then they checked real networks.

Empirical fit

The empirical sweep separates what stays nearly fixed from what moves.

Tests include MNIST, FashionMNIST, CIFAR-10, CIFAR-100 and SVHN; fully connected, convolutional and Wide ResNet architectures; ReLU and Erf activations; mean squared error and cross-entropy; and changes to stochastic-gradient batches. These settings support the taxonomy, with variance-limited exponents near one and resolution-limited exponents that depend on data and model.

§3.1–3.3 · App. A, F · Figs. 1a, S1–S3, S7

Empirical fit

Underparameterized · grow data

αD ≈ 0.98–1.11

Paper-derived range across the displayed settings. Dataset and architecture assignments are not reconstructed as individual points.

PAPER-DERIVED VALUES (Fig. 1a · top-left)
Empirical fit

Overparameterized · grow data

αD ≈ 0.26–0.58

Paper-derived range across the displayed settings. Dataset and architecture assignments are not reconstructed as individual points.

PAPER-DERIVED VALUES (Fig. 1a · top-right)
Empirical fit

Underparameterized · grow width

αW ≈ 0.34–0.62

Paper-derived range across the displayed settings. Dataset and architecture assignments are not reconstructed as individual points.

PAPER-DERIVED VALUES (Fig. 1a · bottom-left)
Empirical fit

Overparameterized · grow width

αW ≈ 0.98–1.03

Paper-derived range across the displayed settings. Dataset and architecture assignments are not reconstructed as individual points.

PAPER-DERIVED VALUES (Fig. 1a · bottom-right)
Empirical fit

Frozen pretrained features

All four regimes

A linear classifier on EfficientNet-B5 embeddings of CIFAR-10 exhibits the same taxonomy.

App. F · Figs. S6–S7

10 · What moves the exponent

Change the labels: nothing. Change the inputs: everything.

Empirical fit

Superclassing merges fine labels into broader categories. In these tests, that changes loss level while leaving the data exponent similar.

Gaussian noise changes the inputs themselves, and the exponent falls as noise grows. The instrument displays the reported superclass band and the verified noise endpoints, without inventing intermediate measurements.

§3.3 · Fig. 3

PAPER-DERIVED VALUES · Fig. 3 · WRN-28-10

Change labels. Then change inputs.

Reported exponent band under superclassingA horizontal band from alpha 0.38 to 0.42 across 5 to 100 classes. It represents a range, not individual fitted points.αD ≈ 0.38–0.425100CLASS COUNT →
Reported αD band0.38–0.42
LessonLoss level changes; exponent stays similar
Reported αD band: 0.38–0.42. Lesson: Loss level changes; exponent stays similar
The superclass control selects a setting but preserves the reported range. Noise has verified endpoints; interior settings have no assigned measured exponent here.
Author interpretation / hypothesis

The authors interpret the contrast as evidence that networks learn input-manifold structure independently of the precise classification task, resembling unsupervised learning. The measured contrast is evidence; that mechanism remains an interpretation.

§4 · interpretation of Fig. 3

Empirical fit

§3.3 · App. A.8 · Fig. S5

11 · Scope and failure

Where the theory runs out.

Direct result / derivation

Asymptotic means the large-resource limit. A result in that limit needs a real hierarchy before it becomes a useful diagnosis.

§4 · §4.1 · App. C.3, F

ILLUSTRATIVE CROSSOVER · NOT FROM THE PAPER

Lose the hierarchy. Lose the diagnosis.

ILLUSTRATIVE LOSS versus D AT FIXED PLogarithmic axes. Hand-specified illustrative bend. The shaded interval marks the selected region.1e+15e-11e+24e-11e+33e-11e+43e-11e+52e-1D AT FIXED P · LOG →ILLUSTRATIVE LOSS · LOG →
Hierarchy-lost band30–300 examples
Hierarchy-lost band: 30–300 examples
§4.1 · Figs. 1a/2a. The shaded D/P ∈ [0.3, 3] is a conceptual crossover band. The paper reports bending empirically and supplies no crossover formula.
Reported limitation

The hierarchy can vanish

When D and P become comparable, the asymptotic predictions lose their premise. The measured curves bend near that region.

§4.1 · Figs. 1a/2a
Reported limitation

Real manifolds need a proxy

There is no precise manifold definition for a real dataset. The paper estimates intrinsic dimension through nearest-neighbor distances in a trained final embedding.

§4.1
Reported limitation

Features can learn

Finite-network kernels evolve during training. The fixed-feature model does not capture that evolution.

§4
Reported limitation

Targets change the gap

Teacher-generated targets scale better than real labels in the appendix experiments.

App. F · Fig. S8
Reported limitation

Smoothness is load-bearing

Non-smooth or unbounded losses can violate the variance-limited exponent.

App. C.3

12 · Context Jamming extension

From image manifolds to language—structural analogy, not identity.

Context Jamming extension

Nothing in this section is established by the paper. The paper studies vision datasets, linear and random-feature models, and wide networks. Language models are not its subject.

The comparison asks what its mechanism might suggest elsewhere. It supplies hypotheses to test, with assumptions that can fail.

Context Jamming extension · paper scope: §1–4

Context Jamming extension

Intrinsic dimension of the input manifold

Effective dimension of the text or representation distribution an LLM must resolve

Context Jamming extension

Resolution-limited αD ∝ 1/d

Small language-model loss exponents as a possible signature of high effective dimension

Context Jamming extension

Variance-limited α = 1

A corner where the growing resource is already abundant relative to the other resource

Context Jamming extension

Linear duality αD = αP

A comparison with the token/parameter symmetry in language-model loss fits

Context Jamming extension

Superclassing leaves the exponent similar

Pretraining as learning input structure beyond a particular downstream label task

Context Jamming extension

Emergent abilities remain an open question

A point of contact with the BIG-bench and emergence discussion

CONTEXT JAMMING EXTENSION · STRUCTURAL ANALOGY, NOT IDENTITY

What intrinsic dimension would that imply?

Preset source: Kaplan · data loss · Appendix A · Tables 4–5

Input exponent α0.095
Conditional intrinsic d = n/α42.1
Input exponent α: 0.095. Conditional intrinsic d = n/α: 42.1
Calculated n/α. This is a conditional analogy, not an intrinsic dimension measured in any language model. Loss exponents are imported from the sibling pages; compute-allocation exponents are excluded.
Context Jamming extension

Does intrinsic dimension predict LM exponents?

Measurable object
Intrinsic dimension of final-layer token embeddings across corpora.
Experiment
Train small LMs on corpora with different estimated intrinsic dimensions at a fixed recipe; compare αD.
Weakening result
No monotone relation between estimated intrinsic dimension and αD.
Primary confound
Corpus difficulty and entropy can differ independently of dimension.
Context Jamming extension

Is there a variance-limited corner for LMs?

Measurable object
Loss versus data in heavily multi-epoch, tiny-model regimes.
Experiment
Fix a small parameter count, grow data far beyond saturation, and fit the approach to the limit.
Weakening result
No slope near one in a verified asymptotic hierarchy.
Primary confound
Repetition, memorization and training time alter the experiment.
Context Jamming extension

Does input noise move LM exponents more than relabeling?

Measurable object
αD under token corruption versus downstream task changes.
Experiment
Construct a text analogue of the paper’s input/label perturbation experiment.
Weakening result
Task changes move αD as much as input corruption does.
Primary confound
Corruption changes entropy as well as geometry.
Context Jamming extension

For a further pointer: Sharma and Kaplan, Scaling laws from the data manifold dimension, JMLR (2022). The Chinchilla chapter provides the allocation comparison; Dyer’s profile provides context for the emergence discussion. Neither establishes the analogies above.

Related work · reference [10], not this paper

13 · Epistemic ledger

What the paper establishes, proposes, and leaves to us.

Direct result / derivation

Established within the assumptions

Theorem 1’s concentration result. Theorems 2–3 as upper bounds. Exact random-feature losses and projection duality. Empirical examples of all four regimes.

Author interpretation / hypothesis

Proposed and interpreted

Using bounds as estimates, including 4/d. Networks learning input geometry. A physics-style research program. Emergent abilities as an open question.

Context Jamming extension

Added by this explainer

The bounded language-model mappings, conditional dimension calculator and falsifiable research questions in Section 12.

14 · Glossary

Terms, briefly.

Definitions · §1–4 / App. F
Population / test loss
Expected prediction error on fresh examples from the data distribution.
Power law
A relation L ∝ X^(−α): equal resource multipliers give equal loss multipliers.
Scaling exponent
αD, αP and αW describe loss scaling in data, parameters and width; αK describes the spectral tail. The log-log slope is −α.
Variance-limited
Finite-resource fluctuations shrink around a smooth limiting prediction; the loss gap follows their variance.
Resolution-limited
The limiting resource still controls how much of the target’s structure can be resolved.
Over- / underparameterized
More features than training samples (P ≫ D), or many more samples than features (D ≫ P), respectively.
Infinite-width limit
A mathematical limit where each hidden layer becomes arbitrarily wide, under a specified training scaling.
Gaussian process
A distribution over functions for which every finite set of function values has a joint Gaussian distribution.
Neural tangent kernel (NTK)
The similarity induced by network parameter gradients; fixed in the appropriate infinite-width training limit.
Data manifold
A smooth lower-dimensional space supporting the input distribution in the theory; an imperfect proxy for real data.
Intrinsic dimension
The number of independent local coordinates needed to move within that space, distinct from the input’s coordinate count.
Nearest-neighbor distance
Distance from a fresh point to the closest training example in the chosen representation.
Lipschitz
Changing an input by a distance r changes the function value by at most a fixed constant times r.
Teacher–student model
A known teacher generates targets; a student is trained to reproduce them.
Random features
Fixed functions sampled before training; only their linear combination is learned.
Kernel / feature covariance
A similarity function between inputs / a second-moment matrix between features, describing the same feature system.
Eigenvalue spectrum
The ordered strengths of independent modes of a kernel or covariance. Spectral decay is how fast these strengths fall with rank.
Interpolation
Matching all training targets; it does not guarantee accurate predictions between them.
Duality
In this linear model, exchanging feature and sample projections maps the two limiting losses into each other.
Double descent
Error can rise near interpolation and fall again as model or sample size continues to grow.
Superclassing
Merging original labels into broader categories while preserving the inputs.
FC / CNN / WRN-d-k
Fully connected / convolutional networks / Wide ResNet with depth d and width multiplier k. Here architecture d is distinct from intrinsic dimension.
ReLU / Erf
Rectified linear / error-function nonlinearities used in the tested networks.
MSE / CE / SGD
Mean squared error / cross-entropy losses / stochastic gradient descent, an optimizer based on batches of examples.
Ridge regularization
Penalizing squared weight magnitude in linear fitting; related to early stopping in App. F, with different spectral filtering.
Cᵗ / asymptotic / O
t continuous derivatives / behavior in a large-resource limit / an upper-order bound, not an exact equality.
Replica methods
A statistical-physics calculation used in App. E.1 to study general random-feature loss scaling beyond the simple tail argument.

15 · Conclusion

The exponent was a measurement all along.

Direct result / derivation

Two mechanisms sit underneath the straight lines. Concentration gives a universal asymptotic correction. Resolution gives exponents tied to intrinsic geometry and, in fixed-feature models, spectral decay.

The curves in Part I and the allocation in Part II now invite another question: what structure is the limiting resource still resolving? An exponent can hint at a mechanism, provided the hierarchy, model and assumptions match.

§1–4 · synthesis of the framework

Author interpretation / hypothesis

The authors pursue a style of theory that combines simple models, contact with realistic systems and experimental checks, drawing on the methodology of physics. Their framework leaves emergent abilities as an open question.

§5 · outlook

16 · Primary record

Read the claim at its source.

Direct result / derivation

Explaining Neural Scaling Laws

Bahri, Y., Dyer, E., Kaplan, J., Lee, J., & Sharma, U. Explaining Neural Scaling Laws. arXiv:2102.06701 [cs.LG] · v1: 12 February 2021 · v2: 29 April 2024. Published in PNAS 121(27), e2311878121 (2024). DOI: 10.1073/pnas.2311878121.

Yasaman Bahri · Ethan Dyer · Jared Kaplan · Jaehoon Lee · Utkarsh Sharma

Google DeepMind, Mountain View (Bahri, Dyer, Lee); Physics and Astronomy, Johns Hopkins University (Kaplan, Sharma).

All authors contributed to all aspects of the work. Part of Sharma’s work was completed during a Google internship. Kaplan and Sharma were supported in part by Open Philanthropy.

The paper combines theoretical derivations in fixed-feature models with empirical vision experiments. This explainer reads v2; no underlying training-run data were recovered or invented.

Primary record · attached arXiv v2; journal fields verified against arXiv

Argument-to-source map

Primary citation map
SectionPaper anchorFigure / appendixEquationsFidelity
01Intro p.1, refs [1–4]——Text
02§1.1–1.2, §2Fig. 1a—Conceptual
03§2.1, Thm 1Fig. 1a TL/BR; App. B, C, C.3Thm 1Illustrative + paper values
04§2.2.1–2.2.2— ; App. DThms 2–3Illustrative + calculated
05§2.2.3, §3.1Fig. 1bL ∝ D^(−n/d)Adapted reconstruction
06§2.3— ; App. EEqs. 1–3, S35–S38Diagram
07§2.3.2–2.3.4Fig. 2b top; App. E.1Eqs. 4–6, S46–S47Calculated
08§2.4Fig. 2b bottom; Figs. S6S37–S38Illustrative
09§3.1–3.3; App. A, FFigs. 1a, S1–S3, S7; Table 1—Paper values
10§3.3; App. A.8Fig. 3; Fig. S5—Paper values
11§4, §4.1; App. C.3, FFigs. 1a, 2a, S8—Conceptual
12§5 (outlook) + Context Jamming——Extension
13§2–5 + editorial synthesis——Text / epistemic classification
14§1–4; App. F——Definitions
15§4–5 + author context——Text
16Full primary record and appendix——Citation map

Complete claim ledger

Primary record with qualifications
ClaimAnchorClassCaveat
Four regimes: variance- vs resolution-limited, for D and for P§1, §2, Fig. 1a, Fig. 2aDIRECT RESULTShown on vision datasets and linear/random-feature/pretrained-feature models
Variance-limited: L − L∞ ∝ 1/x, x = D or width (√P) in deep nets; D or P in linear models§1.1, §2.1, Thm 1, App. B–CDERIVATIONRequires sufficiently smooth loss; App. C.3 gives non-smooth counterexamples
Measured variance-limited exponents of about 1 (αD ≈ 0.98–1.11, αW ≈ 0.98–1.03)Fig. 1a top-left / bottom-right legendsEMPIRICAL FITVaries across dataset, architecture (FC, CNN; ReLU/Erf), batch size, loss (MSE/CE)
Resolution-limited bounds: L(D) = O(D^(−1/d)), L(P) = O(P^(−1/d))§2.2.1–2.2.2, Thms 2–3, App. DDERIVATIONLipschitz assumptions; data on a compact d-manifold
Bounds as estimates: L ∝ D^(−n/d), n ≥ 2; n ≥ 4 for piecewise-linear fits§2.2.3AUTHOR INTERPRETATIONHypothesis, not theorem
Teacher-student: 4/αD tracks input dimension d§3.1, Fig. 1b bottomEMPIRICAL FITStandard datasets (CNN, WRN) relationship "less clear"
Resolution-limited measured exponents (αD ≈ 0.26–0.58, αW ≈ 0.34–0.62 on standard datasets)Fig. 1a top-right / bottom-left legendsEMPIRICAL FITData- and model-dependent by design
Linear random-feature teacher-student model; exact L(P), L(D)§2.3, Eqs. 1–3, App. EDERIVATIONModels trained to the global optimum with MSE; weights initialized at zero
Power-law spectrum λi ∝ i^(−(1+αK)) ⇒ L(D) ∝ D^(−αK), L(P) ∝ P^(−αK)§2.3.2–2.3.3, Eqs. 4–6, App. E.1DERIVATIONEq. 6 intuition uses the large-αK simplification; general case via replica methods
αD ≈ αP ≈ αK empirically (pooled MNIST, random ReLU features)Fig. 2bEMPIRICAL FITRandom features, not trained networks
Smooth (C^t) kernels on d-dim manifold: λn ≲ n^(−(1+t/d)) ⇒ αK ∝ 1/d§2.3.4 (d-torus example)DERIVATIONFor a pure power law, αK ≥ t/d; equality and the dimension estimate require saturation of the bound.
Duality between model-size and dataset-size scaling§2.4, Eqs. S37–S38, Fig. S6DERIVATION + EMPIRICALExact for linear models; sample-wise double descent is one realization
Pretrained EfficientNet-B5 features show all four regimesApp. F, Figs. S6–S7EMPIRICAL FITLinear classifier on frozen features (CIFAR-10)
Superclassing CIFAR-100 leaves αD ≈ 0.38–0.42§3.3, Fig. 3 leftEMPIRICAL FITWRN-28-10
Input noise lowers αD (≈0.58 with no noise → ≈0.29 at noise stddev 0.2)§3.3, Fig. 3 rightEMPIRICAL FITShow endpoints; intermediate pairings only if verified against the figure
Width matters up to a saturating width; depth effect milder§3.3, App. A.8, Fig. S5EMPIRICAL FITWRN-d-k on CIFAR-10
Networks may learn the input manifold "akin to unsupervised learning"§4AUTHOR INTERPRETATIONSpeculative
Asymptotic theory; hierarchy breakdown; manifold proxy; feature learning not modeled§4.1, §4LIMITATION—
Teacher targets scale better than real labelsApp. F, Fig. S8EMPIRICAL FITGap between solvable model and real tasks
Outlook: emergent abilities as open question; physics methodology§5AUTHOR INTERPRETATION—
Conceptual reconstruction

Curves are deterministic calculations or clearly labelled conceptual reconstructions. The spectral calculator sums a finite toy spectrum; its integral approximation extends to infinity. Eq. 6 is shown only up to a constant because its printed prefactor conflicts with the integral. The regime table states the actual D/P hierarchy, and the smoothness control keeps the bound’s direction explicit.

Visualization method