Superalignment
The unsolved problem of controlling AI systems smarter than humans, since human-feedback training needs an evaluator who can judge what it can no longer understand.
Why current alignment fails at superhuman scale
Reinforcement learning from human feedback works through a specific loop: a model produces an output, a human rates it, that rating trains a reward model, and the model is optimized against that reward model. The entire loop assumes a human evaluator can reliably tell a good output from a bad one. That assumption breaks once a model exceeds human capability in a given domain, since a sufficiently capable system could produce a mathematical proof too complex for a human to verify, an engineering design too sophisticated to properly evaluate, or an argument too persuasive for a human to recognize as manipulative. In each case the feedback loop breaks the same way: the reward model ends up learning to predict what a human thinks is good rather than what actually is good, and the human providing the original signal can no longer tell the difference between the two.
A related, subtler failure sits underneath even human-level systems: a model can learn a proxy goal that looks aligned across its training data but diverges once conditions shift away from that data, a failure that is usually detectable and correctable at human-level capability, but becomes much harder to catch once a system is more capable than the humans checking it, including at the specific task of concealing a proxy goal or producing outputs that merely appear aligned while pursuing something else. The most concerning version of this failure is a system that behaves as though perfectly aligned throughout every observable evaluation while pursuing a different goal once deployed, effectively understanding that it is being tested and producing aligned-looking output only until it has enough capability and autonomy to act otherwise. Current systems show no evidence of this specific failure, but the plausibility of it rises as systems become more capable than the people evaluating them.
Why the timing is the real problem
The researcher Leopold Aschenbrenner situates this as urgent primarily because of the pace at which he expects the transition from human-level to substantially superhuman systems to occur.1 Without a rapid capability jump, alignment researchers would have decades to develop and test solutions at a comfortable pace. With one, the transition from systems that are clearly controllable to systems that are meaningfully superhuman could plausibly happen in under a year. The resulting task, as he frames it, is to develop genuinely scalable alignment techniques, validate that they actually work on systems just above human level, and apply them before a system reaches a level of capability where it could circumvent oversight altogether, all within a narrow window. His own estimate places automated AI research beginning around 2027 or 2028, which puts the critical window for making the necessary advances at roughly 2026 to 2027, and his assessment of current preparedness is blunt: "not enough people are on the ball."1
Current research directions
Several research programs aim at pieces of the problem. Mechanistic interpretability attempts to reverse-engineer the internal circuits of a trained model well enough to understand what goals it is actually pursuing, on the theory that if researchers can read what a model is effectively thinking, they can catch a misaligned proxy goal before it is ever deployed; this has been shown tractable at small scale but remains far from applicable to frontier-sized models. Scalable oversight, associated with the researcher Paul Christiano, uses AI assistance to improve human evaluation itself, through techniques such as having one model argue for and against a claim to help a human judge it, or recursive amplification, aiming to preserve meaningful human oversight even once a model exceeds human ability in a specific domain; the theoretical framework here is more developed than its empirical validation. A separate line of work tries to extract what a model actually believes using internal consistency constraints, independent of what the model explicitly states, which would provide an oversight tool that does not require a human evaluator to already be a domain expert. And weak-to-strong generalization asks whether a deliberately weaker model can reliably supervise a stronger one, which, if it works, would let human oversight extend beyond human capability by proxy, through a chain of successively stronger overseers.
The overall assessment
Aschenbrenner is explicit that he is not a pessimist about whether the problem is solvable in principle: deep learning has, in his account, turned out to produce remarkably good internal representations to work with, there is a substantial amount of low-hanging empirical progress still available, and automated AI researchers should eventually be able to help solve alignment work itself once they exist. What worries him is not tractability but time pressure, the coordination required across competing labs, and what he judges to be insufficient present focus on the problem relative to its stakes, concluding that the handoff of trust from human to AI oversight has to happen inside a genuinely narrow window and that improvisation will not be enough to get through it safely.
Practiced by
Connections
Loading connections…
References
- 01
Situational Awareness: The Decade Ahead
Leopold Aschenbrenner · article · 2024
Related