Evals for Everything
Because reinforcement learning can now solve almost any eval you can write, the binding constraint on applying AI across the whole economy is creating evals for every task, a process that inherently requires humans in the loop.
The argument in steps
"RL is so effective that you can create almost any eval and it will be able to solve that eval, and so the barrier to applying AI throughout the entire economy is just creating evals for everything, which inherently requires humans in the loop," argues Brendan Foody.1 The reasoning runs in four steps. Reinforcement learning closes the capability gap on any well-specified target, so capability becomes a function of having a defined goal and a way to score success rather than of raw model progress. That pushes the bottleneck upstream, to defining targets: automating a role requires enumerating and encoding an eval for everything that role does. Defining those targets requires expert humans, since writing a faithful eval for expert work requires people who can do, and judge, that work. And because eval-writing capacity gates deployment, the economy gets automated one eval-able task at a time, constrained by the supply of experts able to author the evals rather than by compute alone.
Evals are definitionally outside model capability
Adarsh Hiremath supplies the structural reason this bottleneck resists being automated away: "evals definitionally have to be outside of model capability, in order to see whether a model is doing well at a particular task you need an eval set created by humans that is better than the model at that task."2 An eval a model could author itself would, by construction, not test anything beyond what the model already does. So as long as there are capabilities models do not yet have, the measuring stick has to be built by a human who does have them, which makes expert human judgment the permanent leading edge. He extends the same logic to data quality: low-quality human data does not push models forward, only high-quality data does, and finding people capable of producing it is a talent-assessment problem rather than a crowdsourcing one.
The labor shape it implies
At a later venue, Foody sharpened the deployment-side claim: generalization from reinforcement learning is weaker than commonly assumed. "Reinforcement learning is becoming so effective that once we have an eval, the models saturate it very quickly, but the generalization is not as strong as most people realize. So the barrier to applying agents to the entire economy is how do we build evals for every workflow." One company's own customer-support operation illustrates the pattern in practice: roughly 80 percent of the support team now builds evals and updates context for the support agent rather than answering tickets directly, a redundant task done once, an agent trained on it, then the team moves to the next task, which is the same fixed-cost amortization pattern that let software eat the world, now applied to agents. Foody's summary: "the future of AI is actually very human," with humans the bottleneck on getting the right context in, especially in enterprise settings where the underlying data was never online in the first place.
Eval as PRD
Foody's sharpest reframing casts the thesis as a product-management failure: "if we think about the model is the product, then the eval is the PRD. So many people have been just like vibe spending on AI without actually writing the PRD of what do they want to implement and how do they measure that it's going to be successful."3 Organizations that deploy AI tools without defining success criteria, watch a large share of pilots fail, and conclude that AI does not work for their business are, on this reading, skipping the eval step rather than discovering a capability ceiling: capability was not the bottleneck, specification was.
Why it matters
If the thesis holds, demand for expert human labor rises with the breadth of AI adoption rather than falling with it, a bull case for human data that runs against the more common assumption that AI adoption is straightforwardly labor-replacing. It also reframes what looks like a missing capability breakthrough as closer to logistics: the blocker is the grind of encoding human judgment into millions of task-specific evals rather than a wall the models themselves cannot cross. Not every capability yields a clean eval; subjective, taste-laden, or safety-critical tasks resist the clean task-plus-score structure, and the claim is made by people whose businesses sell the humans who write the evals, which makes it directionally credible without being disinterested.
Practiced by
Connections
Loading connections…
References
- 01
Brendan Foody · interview · 2026-06-11
- 02
Adarsh Hiremath @ Mercor: The Fastest Growing Startup in Silicon Valley (20VC, E1261)
Adarsh Hiremath · podcast · 2025
- 03
Mercor CEO & Co-Founder, Brendan Foody: How They Grew from $1M to $500M in 17 Months (20VC)
Brendan Foody · interview · 2026
Related