Framework

APEX Benchmark

Mercor's benchmark suite for economically valuable AI capabilities in investment banking, law, medicine, and software engineering: the industry's shift away from academic benchmarks toward measuring what professionals actually do in their jobs, now adopted by top AI labs as a standard for enterprise capability evaluation.

Mercor's benchmark suite for economically valuable AI capabilities, the deliberate counterweight to the academic benchmark tradition: IMO medals, GPQA PhDs, IOI coding. Where academic benchmarks measure models against competitive exam performance, APEX measures models against what professionals actually do in their jobs.1

Origin

The concept originated in an early conversation Brendan Foody had with Sean, an OpenAI researcher, at Bar Gemini, across from the OpenAI office, about IMO medalists. The early benchmark ecosystem was organized around academic excellence: "How could we get IMO medalists to measure Olympiad math? How could we get PhDs in GPQA? All the way to how could we get IOI medalists for various coding evals?"1

Foody and Mercor identified the transition early: the market was moving from academic capabilities to economically valuable capabilities, meaning how models perform on the things that professionals do in their real jobs.

What it covers

APEX covers the distribution of work across the most economically valuable professional domains:

  • Investment banking, financial modeling, deal analysis, pitch work
  • Law, contract review, legal research, memo drafting
  • Medicine, clinical reasoning, documentation, diagnosis
  • Software engineering, the APEX-SWE slice, built with Cognition
  • Plus other high value consulting and analytical domains1

How it works

Expert contributors are recruited from the top of each domain, not median practitioners. The distribution of tasks is weighted to reflect what customers and enterprises actually care about, for example management consulting versus operational consulting, with explicit weighting per sub-category.

Two use cases for the resulting eval data:

  1. Benchmark. Measure model versus model performance; labs use the scores when announcing new models and to determine which model is improving on enterprise relevant tasks.
  2. Reward function. The rubrics and unit tests can serve as the reinforcement learning reward signal, training the model to hill climb the economically valuable eval.

This dual use is key: APEX is simultaneously a measurement tool for labs and enterprises and a training data source for model improvement.1

Adoption

Taken up by "a lot of the top labs," in Foody's words, as an industry standard for determining whether their models are doing well at capabilities enterprises care about, and for helping enterprises make buying decisions about which model to use.1 This is Mercor building the very evals it argues the economy needs.

The academic to economic transition

The academic benchmark era had a structural flaw: Humanity's Last Exam, Olympiad math, and GPQA are, in Foody's consistent framing, "wholly disconnected from outcomes that consumers and enterprises actually care about."2 A model that beats a human on the International Math Olympiad may still fail to draft an email or schedule a meeting.

APEX is designed around the inverse principle: start from what a professional does in a real job and work backwards to the benchmark design. The frontier is no longer defined by academic test scores but by how well the model's performance in benchmark conditions predicts performance in actual professional workflows.

Why it matters

APEX also serves as Mercor's answer to a broader enterprise AI data gap: the problem that search and coding AI broke through because the training data existed, the open internet for search and GitHub for code, but investment banking, law, and consulting have no comparable corpus.3 APEX is a systematic effort to manufacture that missing distribution through expert human labor, and it puts Mercor at the center of how the industry measures enterprise grade AI capability.

Practiced by

Connections

Loading connections…

Related