FounderFiles·N°061·Physics · Scaling · Reasoning
2008 —
Subject·Ethan Dyer, PhD·Theoretical physicist · Scaling-law theorist · Former Google Blueshift, now Anthropic
Ethan Dyer, PhD
Dyer took the physicist’s toolkit of Feynman diagrams, large-N limits and phase diagrams and pointed it at neural networks — then helped build the benchmark and the math model that showed what scale could and couldn’t buy.
He began as a holographer counting black-hole microstates in toy universes. At Google he drew Feynman diagrams for wide networks, found the learning rate at which training “catapults,” co-wrote with Jared Kaplan the paper that explains where scaling laws come from, helped organize the 204-task stress test that made “emergence” a headline, and helped build the language model that learned to do math. Then, quietly, he turned up at Anthropic — the same company as his co-author.
From a boundary term to a boundary theory
His first paper was about boundaries. As a Columbia student in 2008, Dyer co-wrote with Kurt Hinterbichler Boundary terms, variational principles, and higher derivative modified gravity (Phys. Rev. D, January 2009). The paper asks what you must add at the edge of spacetime to make a theory of gravity’s equations well-posed. A companion paper worked out the same question for the DGP “π-Lagrangian” and the galileons. It is still one of his most-cited physics papers, at roughly 270 Google Scholar citations.
He went to MIT for his PhD (2009–2014) and joined the Center for Theoretical Physics. His MIT-era work sits in the AdS/CFT world: monopole operators in three-dimensional conformal field theories with Márk Mezei and Silviu Pufu, and super-Rényi entropies and Wilson loops for N = 4 super-Yang–Mills and their gravity duals with Michael Crossley and Julian Sonner (JHEP, 2014). The title and advisor of his dissertation do not appear in any public thesis record we could find, so this file does not print them.
He then moved to Stanford’s Institute for Theoretical Physics (2014–2017) and went further into three-dimensional gravity and two-dimensional CFT: An Extremal N = 2 Superconformal Field Theory and Universal Bounds on Charged States in 2d CFT and 3d Gravity with Nathan Benjamin, Liam Fitzpatrick and Shamit Kachru; 2D CFT partition functions at late times with Guy Gur-Ari; Spinning Geodesic Witten Diagrams with Daniel Freedman and James Sully; Constraints on Flavored 2d CFT Partition Functions with Fitzpatrick and Yuan Xin, which lists both Stanford and Johns Hopkins; and The most irrational rational theories(2019). He gave a Princeton seminar titled “Small black holes in near extremal gravity,” on counting black-hole microstates in the conjectured duality between pure 3D gravity and “extremal” 2D CFTs.
This is the same subject area as Kaplan’s Aspects of Holography: lower-dimensional boundary theories that encode higher-dimensional gravity. Dyer and Kaplan were not just two physicists who both drifted into machine learning. They were two holographers.
Stanford, Hopkins, and a workshop on physics for machine learning
The crossover can be dated. Dyer sat on the scientific organizing committee of a “Theoretical Physics for Machine Learning” workshop, listed as “Stanford University & Johns Hopkins.” His fellow organizers were Adam Brown, Paul Ginsparg, Guy Gur-Ari and Jaehoon Lee. By 2018 he was at Google. His first ML paper, Gradient Descent Happens in a Tiny Subspace (December 2018, with Gur-Ari and Dan Roberts), showed that during training the gradient quickly settles into a small subspace spanned by the top Hessian eigenvectors. It is a spectral argument of the kind a physicist would make, and it has about 375 citations.
The Johns Hopkins affiliation matters for the Kaplan story. Kaplan joined the Hopkins Department of Physics and Astronomy in 2012–13 and stayed on its faculty after joining OpenAI in 2019. Dyer carried a Hopkins affiliation on papers from about 2017 to 2019, so the two overlapped there. Neither has described that overlap publicly in anything we found, and there is no co-authored physics paper between them. Liam Fitzpatrick, a frequent co-author of both, links the two holography networks.
Why physicists made this crossing at all — and who else did — is the subject of The Data Desert: Physics, AI, and the Diaspora, where Kaplan is the defector and Dyer turns out to be a second one, with a co-author.
Feynman rules for infinitely wide networks
The paper that made Dyer’s name in ML was Asymptotics of Wide Networks from Feynman Diagrams, written with Gur-Ari (arXiv:1909.11304; ICLR 2020). It brings a field theorist’s bookkeeping to neural networks. Treat the width as a large parameter, like N in a large-N gauge theory. Expand correlation functions of the network’s outputs and derivatives in 1/width. Organize the terms with diagrams and read off how each quantity scales before computing anything. It has about 159 citations.
Related papers followed. Asymptotics of Wide Convolutional Neural Networks, with Anders Andreassen, extended the program to CNNs. The large learning rate phase of deep learning: the catapult mechanism (arXiv:2003.02218, March 2020), with Aitor Lewkowycz, Yasaman Bahri, Jascha Sohl-Dickstein and Gur-Ari, identified a regime above the standard “lazy” learning-rate limit. There, loss first rises, curvature collapses, and the network “catapults” into a flatter region and generalizes better. That paper has about 366 citations.
This is Kaplan’s instinct — networks are physical systems, and physical systems have laws — done at the level of dynamics instead of empirical fits. Kaplan measured the curves. Dyer wrote down the perturbation theory.
“How we understand these sharp transitions is a great research question.”
Explaining Neural Scaling Laws
February 12, 2021. Yasaman Bahri, Ethan Dyer, Jared Kaplan, Jaehoon Lee and Utkarsh Sharma posted Explaining Neural Scaling Laws (arXiv:2102.06701). A revised version appeared in PNAS121(27) in 2024. This is the link that holds this file together. It is the only verified paper Dyer and Kaplan share, and it arrived a year after Kaplan’s OpenAI Scaling Laws paper, as its theoretical counterpart.
The claim, from the abstract: variance-limited and resolution-limited scaling behavior for both dataset and model size, for a total of four scaling regimes.
Variance-limitedscaling follows from the existence of a well-behaved infinite-data or infinite-width limit. These are corrections to a smooth limit — the same large-N logic as the Feynman-diagram paper. Resolution-limitedscaling is explained by positing that models are effectively resolving a smooth data manifold; the exponent is set by the manifold’s intrinsic dimension, and higher-dimensional data means slower power laws. In the large-width limit the same exponents come from the spectrum of certain kernels, and the authors present evidence that large-width and large-dataset resolution-limited exponents are related by a duality. Model size and dataset size act as mirror images.
Kaplan’s 2020 paper said that loss follows power laws. Bahri, Dyer, Kaplan, Lee and Sharma said why, and what sets the exponent.
The Kaplan file calls the shape of the power-law claim the most consequential physics result in machine learning history. If so, this paper is its microscopic explanation, and Dyer is one of its five authors — with roughly 720 citations. Play with the curves in the interactive Scaling Laws explainer.
Three months after the preprint, Dyer took the argument to a room of theoretical physicists. His May 2021 talk, “The role of scale in deep neural networks”, is billed with an abstract about understanding how performance improves with scale — and about telling apart the problems scale alone can solve from the ones where new ideas are needed. That second clause is the one worth keeping in view: the theorist of why the curves exist was already asking where they stop.
The benchmark that made “emergence” a headline
In 2020, per Quanta Magazine, Dyer and others at Google Research predicted that large language models would have transformative effects, and asked the research community to contribute examples of difficult and diverse tasks. The result was BIG-bench, Beyond the Imitation Game(arXiv:2206.04615; TMLR 2023). The paper’s abstract counts 204 tasks, contributed by 450 authors across 132 institutions.
The finding that went viral: around 5% of BIG-bench tasks see models achieve sudden score breakthroughs with increasing scale, though that behavior can depend sharply on the metric used to probe performance. The paper’s contributions section says BIG-bench was managed and organized by Guy Gur-Ari, Jascha Sohl-Dickstein, Noah Fiedel and Ethan Dyer, and that BIG-bench Lite was developed by Dyer.
This is where the Dyer file meets the Kaplan file’s “Emergence as Mirage.” The same Quanta piece reports that when Dyer’s team posed the emoji-movie task as multiple choice, the accuracy improvement was less of a sudden jump and more of a gradual increase — the metric-artifact explanation that the 2023–24 “mirage” papers later made formal. BIG-bench’s data supported both the emergence story and its refutation.
Dyer’s own framing was open-ended: “sharp transitions” as a research question, not a proof of magic.
“Despite trying to expect surprises, I’m surprised at the things these models can do.”
Teaching a language model to show its work
June 2022. Solving Quantitative Reasoning Problems with Language Models (arXiv:2206.14858; NeurIPS 2022) introduced Minerva. Dyer is one of fourteen authors, and the Google Research blog post was bylined “Ethan Dyer and Guy Gur-Ari, Research Scientists, Google Research, Blueshift Team.” It describes a model that solves mathematical and scientific questions using step-by-step reasoning, producing numerical calculations and symbolic manipulation without relying on external tools such as a calculator.
Related 2022 papers from the same team: Exploring Length Generalization in Large Language Models (NeurIPS 2022 oral), which documented how badly models extrapolate to longer problems than they were trained on; Block-Recurrent Transformers (about 220 citations); and Effect of scale on catastrophic forgetting (ICLR 2022, about 313 citations). Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models (2023) followed on the same line of work.
Dyer presented this line on November 15, 2022, at a Stanford Applied Physics/Physics Colloquium (video) titled “Lessons from scale for large language models and quantitative reasoning.” The abstract notes that some measures of progress show remarkably robust power-law improvement over many orders of magnitude, while other capabilities remain difficult to extrapolate. That sentence joins ENSL and BIG-bench in a single line.
The quiet move
Blueshift became the reasoning group behind Gemini. Behnam Neyshabur, who co-led it from 2022 to 2024, writes that Blueshift was responsible for the reasoning capabilities of the first version of Gemini, and that it went on to develop a Gemini 1.5 math-specialized model and improve reasoning in Gemini 1.5 Pro and Flash and Gemma 2. Dyer is a credited author on the Gemini 1.0, 1.5 and 2.5 technical reports. Author lists that large, in the thousands, do not show individual roles, so his specific Gemini contribution is not established. As late as the September 2024 Michelangelo long-context evaluation paper, he appears as a contributor listed under Google DeepMind.
Then the affiliation changes. Neyshabur left Google for Anthropic in late 2024 to co-lead a Discovery team aimed at building an AI scientist and engineer; he has since left to co-found Mirendil. The Aspen Center for Physics lists “Ethan Dyer, Anthropic”as an organizer of the January 11–16, 2026 “Theoretical Physics for Artificial Intelligence” meeting, alongside Adam Brown, Dmitry Krotov and Eva Silverstein. Anthropic science posts from 2026 thank “Ethan Dyer” for feedback — “Long-running Claude for scientific computing” (March 23, 2026) and “Paving the way for AI agents in biology” — and his Hugging Face profile lists Anthropic as his organization.
What the record does not give us: his start date, title, or team at Anthropic. That he works on the Discovery or science side is an inference from the acknowledgments and from Neyshabur’s path, not a confirmed fact.
Nor does anything tie him to Gemini 4 Argon. Google’s September 30, 2026 announcement names no individual contributors, and Dyer was listed at Anthropic by then. Argon ships the reasoning line his old team started; the person who helped start it works for the rival.
Two holographers, one company
The paths did not diverge into rival labs, as the obvious frame assumes. They diverged for about five years — Kaplan building a lab, Dyer building reasoning systems inside Google — and then convergedat Anthropic. The side-by-side is below. Dyer’s move also fits a documented 2024–26 flow of Blueshift and Gemini reasoning talent toward Anthropic, including his long-time collaborator Neyshabur; as it applies to Dyer, that pattern is an inference.
The two holographers who co-wrote the theory of why scaling laws exist now work at the same company. That, not rivalry, is the shape of this story.
- 2008–09Columbia — first papers with Kurt Hinterbichler on boundary terms in modified gravity and galileons.
- 2009–14MIT, Center for Theoretical Physics — PhD work in the AdS/CFT world. Thesis title and advisor not publicly verified.
- 2014–17Stanford Institute for Theoretical Physics — postdoc; extremal and 3D gravity, 2D CFT partition functions.
- ~2017–19Johns Hopkins — affiliation on papers, in the same department as Jared Kaplan.
- 2018Google Research — joins the Blueshift team. Gradient Descent Happens in a Tiny Subspace (Dec).
- Sep 2019Asymptotics of Wide Networks from Feynman Diagrams (ICLR 2020).
- Mar 2020The catapult mechanism — a learning-rate “phase” above the lazy limit.
- 2020BIG-bench effort launched at Google Research.
- Feb 12 2021Explaining Neural Scaling Laws posted with Bahri, Kaplan, Lee and Sharma.
- May 17 2021Gives the talk “The role of scale in deep neural networks” for the University of Chicago theoretical physics community (uploaded May 21).
- Jun 2022Minerva; BIG-bench paper posted; length generalization and Block-Recurrent Transformers at NeurIPS.
- Nov 15 2022Presents “Lessons from scale for large language models and quantitative reasoning” at the Stanford Applied Physics/Physics Colloquium (recording uploaded Nov 16).
- Mar 2023Quanta Magazine quotes Dyer on emergence and “sharp transitions.”
- Apr 2024Explaining Neural Scaling Laws published in PNAS.
- Sep 2024Michelangelo long-context evaluation — contributor, listed under Google DeepMind.
- Late 2024Behnam Neyshabur, Blueshift co-lead, leaves Google for Anthropic.
- Jul 2025Credited on the Gemini 2.5 technical report.
- By ~Sep 2025Listed as “Ethan Dyer, Anthropic” among organizers of the Aspen Center for Physics winter 2026 meeting.
- Jan 11–16 2026Aspen — Theoretical Physics for Artificial Intelligence.
- Mar 2026Thanked in Anthropic’s “Long-running Claude for scientific computing.”
- Sep 30 2026Google announces Gemini 4 Argon. The announcement names no individual contributors.
- Nov 2022Lessons from scale for large language models and quantitative reasoningStanford Applied Physics/Physics Colloquium · YouTube · embedded above →
- 2009Boundary terms, variational principles, and higher derivative modified gravityPhys. Rev. D 79 · with Kurt Hinterbichler
- Dec 2018Gradient Descent Happens in a Tiny SubspacearXiv · with Guy Gur-Ari and Dan Roberts
- 2019Asymptotics of Wide Networks from Feynman DiagramsarXiv:1909.11304 · ICLR 2020 · with Gur-Ari →
- Mar 2020The large learning rate phase of deep learning: the catapult mechanismarXiv:2003.02218 · with Lewkowycz, Bahri, Sohl-Dickstein, Gur-Ari →
- Feb 2021Explaining Neural Scaling LawsarXiv:2102.06701 · PNAS 2024 · with Bahri, Kaplan, Lee, Sharma →
- May 2021The role of scale in deep neural networksLeinweber Institute for Theoretical Physics (UChicago) · talk · YouTube →
- Jun 2022Beyond the Imitation Game (BIG-bench)arXiv:2206.04615 · TMLR 2023 →
- Jun 2022Solving Quantitative Reasoning Problems with Language Models (Minerva)arXiv:2206.14858 · NeurIPS 2022 →
- Mar 2023The Unpredictable Abilities Emerging From Large AI ModelsQuanta Magazine · Stephen Ornes →
- 2026The Data Desert: Physics, AI, and the DiasporaContext Jamming · dispatch →
Education.Columbia University (undergraduate physics research, 2008–09). MIT, PhD in physics, Center for Theoretical Physics (2009–2014; dissertation details not yet verified). Postdoc at the Stanford Institute for Theoretical Physics (2014–2017).
Affiliations.Johns Hopkins Physics & Astronomy (affiliation around 2017–19). Google Research, Blueshift Team (2018 to about 2024), later listed under Google DeepMind. Anthropic (by about 2025; role and start date unconfirmed).
Collaborators worth naming. Guy Gur-Ari (co-founder of Augment), Aitor Lewkowycz, Yasaman Bahri, Jaehoon Lee, Jared Kaplan, Utkarsh Sharma, Jascha Sohl-Dickstein, Behnam Neyshabur, Vinay Ramasesh, Liam Fitzpatrick, and Adam Brown, with whom he co-organized both physics-for-ML meetings.
Range. Beyond physics and ML, a genomics paper (Tanigawa, Dyer, Bejerano, PLoS Computational Biology, 2022) and a glass-physics paper, Linking dynamical heterogeneity to static amorphous order.
Public voice. Thin on the record: two Quanta quotes (2023), the Stanford colloquium (Nov 2022, embedded at the top of this page), a 2021 University of Chicago theoretical-physics talk, the Minerva blog byline, and conference organizing roles.
π-Bridge
Carries the prior of a first field into a second and finds the governing law that was invisible to native practitioners; pays in delayed gratification.
- Credential Path
- Doctoral
- Abstraction
- Top Down
- Exit Horizon
- Deferred
- Moat Instinct
- Theoretical Insight
- Capital Posture
- None
- Large-N and holographic theorists
- Jared Kaplan (co-author)
- The physics-of-deep-learning tradition
A small reasoning persona distilled from this file. Inject it into a chat or deep-research context to assess a business problem the way Dyer would.
Reason as a field theorist assessing a business problem. First ask which limit the system is near: is performance capped by noise around a well-understood limit (variance-limited) or by how finely you can resolve the underlying structure (resolution-limited)? Estimate the intrinsic dimension of the problem, because it sets how fast returns decay. Separate smooth metrics from pass/fail metrics before declaring a breakthrough. Build the benchmark before you believe the curve.
{
"$schema": "https://www.contextjamming.com/schemas/founder-context-v1.json",
"file": "N°061",
"persona": "Ethan Dyer, PhD",
"archetype": "pi-bridge",
"shape": "π",
"one_line": "A holographer who maps a system into regimes, then builds the benchmark that tests where the map holds.",
"cognitive_basis": {
"credentialPath": "doctoral",
"abstractionDirection": "top-down",
"exitHorizon": "deferred",
"moatInstinct": "theoretical-insight",
"capitalPosture": "none"
},
"operating_questions": [
"Is performance capped by noise around a well-understood limit, or by how finely the system resolves the underlying structure?",
"What is the intrinsic dimension of this problem, and what does it say about how fast returns decay?",
"Is this a real breakthrough, or an artifact of a pass/fail metric?",
"What would the large-N expansion of this system look
…