Scaling Laws
The empirical observation that model capability scales predictably with data and compute, the foundational insight behind the frontier AI wave and the structural driver of demand for both compute and data infrastructure.
The empirical observation
Language model capability scales predictably as a power law with the amount of training data, compute, and parameters, with few hard ceilings visible at current scale. More data plus more compute produces a more capable model, which transforms model building from a clever-architecture contest into a resource-allocation problem: acquire more high-quality data, acquire more compute, and scale the training run. Alexandr Wang puts the downstream implication for data demand plainly: "the need for data will basically grow to consume all available information and knowledge that humans have."1
The inflection as felt by a practitioner
Wang offers a first-person account of when scaling became undeniable rather than merely theoretical. Self-driving teams before 2019 rarely thought in scaling terms at all, since their models had to run on constrained on-car compute. An early large language model in 2019 felt "mildly cool" but not qualitatively significant. Its 2020 successor made scaling feel real for the first time; a widely told anecdote about a friend getting genuinely frustrated arguing with the model marked a qualitative shift, even though the broader public had not caught on yet. A visual generative model in 2022 convinced people outside language entirely, and a public chatbot later that year was, in his telling, the defining moment for the company and for the world. A more capable model the following year was the undeniable signal, achieving near-zero hallucination in limited domains and opening what he calls an astronomically large opportunity.1
He draws a direct comparison to the historical pace of semiconductor improvement: "the analogies between AI and Moore's Law are pretty clear. You'll get on different technical curves, but if you zoom way out, it'll just feel like this smooth improvement of models."1 Each new scaling curve, pretraining then reasoning and reinforcement learning, is its own S-curve, but the overall envelope reads as continuous improvement.
The shift to reasoning and reinforcement learning
By Wang's more recent account, gains are no longer coming primarily from pretraining: "we're moving to a new scaling curve of reasoning and reinforcement learning. It's shockingly effective." The data required for reinforcement-learning fine-tuning, environments, rubrics, human feedback on reasoning chains, differs from and complements pretraining data rather than displacing the demand for it. If the pattern holds, the specialized fine-tuned model becomes a firm's core intellectual property the way a codebase is today, with the underlying alpha coming from data sets and environments specific to a given business problem.
At a later venue, from inside a different lab, Wang described a more procedural variant of the same idea: rather than simply scaling and observing the result, a lab can try to predict the capability level a given scale of training will produce, release a model as a checkpoint against that prediction, verify it held, and proceed to the next rung.2 One deliberately smaller model built under this approach was positioned as validating the calibration for a larger flagship model still to come, treating each release as a hypothesis test rather than simply a bigger run.
A three-axis version of the same compound
Masayoshi Son extends the scaling argument into a three-factor compound: rather than a single axis running from compute to capability, he separates chip-volume growth from infrastructure investment, per-chip efficiency growth from hardware generations, and algorithmic-efficiency growth from model research, each independently improving by roughly an order of magnitude per cycle and multiplying together into a combined improvement far larger than any single axis alone.3 The distinction matters because it shows infrastructure capital spending compounding with the other two factors rather than only supplying more raw compute on its own.
Open questions
When, not if, scaling laws hit a ceiling remains unresolved, with some suggesting diminishing returns on pretraining data specifically even as the reasoning curve keeps producing gains. Whether compute or data quality is the more binding constraint going forward is also unsettled, and the rapid diffusion of frontier capability through openly available models raises the further question of how quickly any single lab's lead actually erodes once a rival can approximate the same capability at a fraction of the training cost.
Practiced by
Connections
Loading connections…
References
- 01
Alexandr Wang: Building Scale AI, Transforming Work With Agents & Competing With China (Lite Cone)
Alexandr Wang · podcast · 2025
- 02
Meta AI Chief Wang on Winning the Race in AI
Alexandr Wang · interview · 2026
- 03
Masayoshi Son at FII Miami 2026
Masayoshi Son · interview · 2026
Related