Framework

Agent Time Horizon

The METR curve read as a ladder of delegation, from a seconds-long command to an hours-long task to a months-or-years-long mission an agent runs alone.

The curve as a ladder, not just a number

An AI two years prior could sustain roughly ten to twenty seconds of unassisted, human-equivalent work before a person had to intervene, correct it, or catch a mistake. Cognition co-founder Scott Wu cites that figure from the AI research organization METR, and notes the duration has been doubling roughly every couple of months since, reaching hours by 2026.1 His contribution is not the curve itself but a ladder read off it. At the seconds rung, the human hands over a command and keeps everything else: sequencing, judgment, context. At the hours rung, the human hands over a task, choosing and scoping it while reviewing the result. At the months-to-years rung, the human hands over a mission, and the only thing still owned is what the human cares about.

Why Edwin Land is the reference case

Wu's first attempt at describing what a year-long agent mission should look like was a broken email formatter, which podcaster David Senra rejected as a waste of the capability. Senra's replacement example is Edwin Land, who hired a researcher at Polaroid and set him one question, how to take instant photography from black and white to color, on which the man reportedly thought for two years before solving it. The anecdote sets the register for the mission rung. It is not a long to-do list, it is a single, underspecified, multi-year question handed to someone trusted to work on it without supervision. Wu's revised examples match that register: a societal problem that needs to be understood and coordinated on, a video game that fuses mechanics from two others, a novel materials construction where the agent runs its own experiments for months. Senra adds a further layer, an agent whose mission is choosing which missions to send other agents on.

What changes at each rung

The ladder converts a capability benchmark into a product and organizational question. A tool built to be excellent at commands is not a task tool with a longer timeout, and a task tool is not a mission tool with more patience; each rung implies a different interface, a different unit of trust, and a different failure mode. It also inverts which resource is scarce. At the command rung, execution is the scarce thing. At the mission rung, execution is abundant and knowing what is actually wanted is scarce, the same inversion described from the human side as commissioning replacing steering.

What the framework leaves open

Wu's own move, extrapolating the doubling curve out to months and years, is a challenge to the pessimist rather than an argued mechanism; nothing in his account explains what would make the curve hold or what would make it break. Task duration is also not the same as task difficulty. The question Land set took two years because the search space was unstructured and feedback was nearly absent, not because it contained many sequential steps, while the METR curve measures duration under a verifiable objective and Wu's mission-length examples mostly are not verifiable. Wu's own timeline for reaching the mission rung, for most of what he describes, is about five years, with the stated caveat that five years of AI progress is roughly a century of ordinary time.

Practiced by

Connections

Loading connections…

References

  1. 01

    The Future of Software & AI

    Scott Wu · podcast · 2026

Related