Principle

Unbounded Demand for Frontier Intelligence

Demand for the highest-capability AI models has no visible ceiling in high-stakes, high-complexity domains, because in those domains you would always upgrade your staff engineers to distinguished engineers if you could. The apparent bear case, cheap open models eating the frontier, misreads a demand curve that keeps expanding.

The explanation

The apparent bear case on frontier AI models runs like this: open-weight models keep getting cheaper and more capable, so the set of tasks that genuinely require a frontier model should shrink toward zero over time. The counter is that this misreads the demand side entirely. The test offered is simple: ask any software company whether it would like to upgrade its staff-level engineers to principal or distinguished-level engineers, and essentially all of them would say yes. Intelligence is not a fixed-size job that gets filled once and is then done, it is an input whose marginal use keeps expanding as it becomes better and cheaper, through invention, discovery, new products, and work that runs around the clock.1

The bifurcation, not the substitution

The market splits by task rather than shrinking toward a single winner, and both halves of the split grow simultaneously. Commodity tasks migrate to cheap, once-frontier models, since a routine customer-service interaction does not require the most capable model available, and effective capability per dollar keeps falling fast enough that there is a real capability overhang for easy work, forming something like an assembly line where yesterday's frontier gets fine-tuned down into open weights and pushed toward high-volume, low-stakes workloads while companies mix and match by task. High-stakes, high-complexity tasks, coding, science, materials research, and law, keep pulling toward frontier models, because an incremental step up in capability is worth a large premium in domains like these, and demand there is effectively unbounded. The shrinking-frontier-market intuition confuses the set of tasks a cheap open model can now handle, which is genuinely growing, with the ceiling on demand for the very best available intelligence, which is argued to be systematically underestimated rather than approaching any limit.1

Why the ceiling is underestimated

It is hard to picture how much intelligence that can work around the clock, inventing, building, and discovering, could actually be used, and how much of it could be absorbed. This is framed as an invention and discovery frontier rather than an automation frontier: the binding limit is the scope of human imagination about what to point the intelligence at, not the underlying supply of tasks available to it.1

Why it matters

This is the demand-side twin of the idea that cheaper intelligence leads to more total intelligence being consumed rather than less, sharpened here into a specific claim that cheap commodity intelligence can coexist comfortably with unbounded demand for frontier intelligence, so efficiency gains at the bottom of the market do not cap total spend at the top. It also underwrites forecasts that developer-level token spend will keep converging toward a meaningful share of total compensation over time, not as a bubble but as the demand curve simply asserting itself, and it sets a rough floor under token prices, since unbounded frontier demand paired with a genuinely scarce physical resource, compute and energy, implies persistent scarcity rather than commoditization at the frontier. It also helps explain the economics of open-model releases, since if frontier demand is both bottomless and margin-rich, the labs with the most frontier capability have little incentive to ship frontier-equivalent open weights that would cannibalize their own highest-margin business, the economic half of the argument for withholding frontier weights on safety grounds.1

The inference market itself makes the underlying demand curve visible in hard numbers. Inference is described as being on the path to becoming the largest market in the world, and one major AI company's token throughput grew to several quadrillion tokens processed per month, roughly three hundred times its level two years earlier, a concrete measure of unbounded demand made real, and it is exactly this kind of demand curve that is pulling significant capital into purpose-built inference hardware distinct from general-purpose chips.2

Tensions

Whether demand is truly unbounded or simply very large is partly a rhetorical distinction rather than an empirical one: the real claim is that the ceiling sits far higher than current markets are pricing in, and separating no ceiling from a very high ceiling matters directly for whether present AI valuations are justified. Concrete forecasts of developer token spend as a share of compensation are explicitly offered as bets rather than settled facts, openly stated as figures the forecaster would not want to be held to years later, so real usage data will ultimately settle the question either way. And unbounded latent demand can still be throttled for years by integration and trust bottlenecks rather than by any shortage of underlying capability, since the share of enterprise tasks that are fully automated today, even where a capability gap does not exist, remains a rounding error next to the tasks technically eligible for automation.1

Practiced by

Connections

Loading connections…

References

  1. 01
  2. 02

    How to Win the Largest Market in AI

    Sarah Wang · article · 2026

Related