Framework

RL Environments

The successor to RLHF as the primary data type for training frontier models: complex, multi-tool, long-horizon task environments where humans build the framework and verification layer so a model can learn to replicate an expert workflow rather than just rate an output.

What they are

Where reinforcement learning from human feedback asked people to rate or compare model outputs, a labeling task, an RL environment asks a human to define and verify a complete task trajectory, the kind of work a professional would do over hours or days, using real tools. An environment includes a task description specifying what the model should accomplish, the full tool context such as email, calendar, file storage, chat, and any custom software the task requires, and a verification or reward signal scoring whether the model actually accomplished the task correctly. A model operates inside the environment, takes actions across the available tools, and receives feedback from a human-authored rubric, which researchers then use to improve model capability. The shift mirrors what models can now attempt: earlier human-feedback data was appropriate when undergraduate-level problems were near the frontier, while a full RL environment is appropriate once the task looks more like multi-year professional work, where a researcher can no longer personally verify correctness and needs an expert human who can.1

The addressable-market paradox

Brendan Foody's framing of the ceiling on this market is that its total addressable size is limited by the amount of work humans are still better at than models, but that ceiling is not fixed. Early single-tool environments converge quickly, since once a model has closed most of the gap on a narrow task, only a small share of contributors can still stump it. Adding more tools and much longer task horizons resets the bar, restoring a large pool of contributors who can meaningfully challenge the model again. This is the same dynamic as Jevons Paradox in Intelligence playing out on the supply side of data: as models close the gap on simpler tasks, the frontier of RL environments simply expands to harder ones, keeping demand for expert human contributors roughly constant or growing rather than shrinking toward zero.1

Market position

One company claims roughly fifty to sixty percent market share in RL environments as of mid-2026, a sharp shift from the prior era, when older-generation crowdsourcing companies dominated. That shift is largely explained by the difference between RLHF and RL environments themselves: RLHF rewarded low-to-medium caliber human raters working in minutes, while RL environments reward high-caliber experts working in hours to days, backed by a rubric plus task completion rather than a simple preference signal, which requires categorically different matching and vetting infrastructure than sourcing workers to rate short outputs. Multiple foundation-model lab leaders reportedly believe RL environments will eventually subsume the entire economy, not remain a niche data type but become the primary vehicle for encoding human expertise into models.1

A separate account of the broader data-selling market to frontier labs places these environments, described as worlds simulating a real business function such as a bank, a hospital, or an enterprise software workflow, at the current top of the value chain, while warning that they depreciate by succeeding: as a model masters a given environment, that environment's training value decays, and at least one major lab is reportedly spending roughly a billion dollars a year building environments in-house rather than buying them, which shifts the truly scarce asset toward whatever is hardest to verify at the frontier rather than the environment itself.2

The paved-with-humans thesis

The fuller argument for why this is not a temporary bottleneck is that today's models can win an math olympiad gold medal and out-reason most PhDs, yet cannot draft a routine email or schedule a meeting, cannot do many of the basic multi-tool tasks that take a person only a few hours, and that the entire path toward automating the economy and building agents for everything runs through humans creating eval and environment workflows first. Every new capability the economy wants an agent to have requires an environment authored by a human who already possesses that capability, so automating a large economy at scale requires creating a very large number of these environments, and the humans authoring them function as the forward line of that deployment effort rather than as a transitional inconvenience.1

Connection to the agentic economy

RL environments are the production input that makes reliable, broad agent deployment possible in the first place, since agents can only be trusted with tasks that have a verified environment behind them, and the environments effectively define the current scope of reliable automation. The same logic extends into the physical world: reinforcement-learning post-training on pretrained embodied models is named as one of the most consequential recent developments in physical AI and the likeliest path to the reliability threshold real-world production demands, except that in the embodied case the reward signal comes from physics and simulation rather than from a human-authored rubric, and the same compute-scaling dynamics that took language models through several capability generations are beginning to show up in the embodied domain as well.3

Practiced by

Connections

Loading connections…

Related