Principle

Engineered Failure Rate

Treating failure as a scheduled output of an organization rather than an accident to minimize: deadlines set at roughly fifty percent achievability, a mandate to add back ten percent of whatever was deleted, and a hiring filter built around naming the ways a candidate has screwed something up, on the logic that failure is irrelevant unless it is catastrophic.

The governing claim

Most organizations try to drive failure toward zero. Elon Musk's companies instead set a target rate for it, on the premise that failure is essentially irrelevant unless it is catastrophic. The economics behind that premise are about volume: a company about to build ten thousand of a single engine will replicate any inefficiency it fails to catch ten thousand times, so the stated ambition is to fail as many as 500 times on a single design decision before scaling it, because the failure that reveals the right answer pays off a million times over.1

The instruments

Several familiar operating rules turn out to be mechanisms for hitting that target rate rather than separate ideas. The rule that a team must add back at least ten percent of whatever it deleted sounds like a safety margin but functions as the opposite: it instructs the team to keep deleting past the point of correctness, effectively ordering them to build something that does not work so they can find out which part was actually load-bearing. Deadlines are deliberately set at roughly fifty percent achievability: "I want to pick a deadline I'm 50% likely to make. We're going to miss half our deadlines and I'm totally fine with that."1 The logic is that a team will always consume at least as much time as it is given, so hitting one hundred percent of deadlines is itself evidence they were set too loosely; missing half of them is the price of the other half being genuinely surprising. The hiring filter follows the same logic at the individual level: if a candidate cannot name the four ways they screwed something up, they were not the one doing the real work, since a failure inventory is difficult to fake and functions as proof of proximity to the actual work.

Musk also draws a distinction between timing errors and directional errors, describing himself as rarely right about when something will happen and almost always right about which direction it moves, and treats being early by a year or several years as an acceptable category of failure. And the failure has to stay cheap: early Starship prototypes shipped with no doors because the first ten-plus ships were never coming back from orbit anyway, so anything not required to clear the current bar was complexity being paid for needlessly.

Why the opposite bias wins by default

The force being counteracted is not laziness but self-protection: if a job is on the line, or someone wants to look good to peers, or simply wants the company to succeed, nobody wants to be seen failing, and engineers in particular hold this at the level of their own component, not wanting their specific part to be the one that fails. Each of those instincts is locally rational, and added together they produce a system that never learns, which is why these rules have to be repeated to the point of parody rather than stated once. The failure also has to be cheap and small for any of this to work, which is the entire reason the rest of the operating system, fast iteration, colocated design and production, vertical integration, exists: they are what turn a failure into something that costs hours rather than quarters.

Michael Dell's independent version

Michael Dell states a version of the same rule from a completely different industry: he wants mistakes to happen, because there is no existing playbook for a genuinely new business model, but he wants them small and he does not want to repeat the same one twice.2 The convergence is notable because it comes from a hardware and direct-sales operator working in an unrelated sector, arriving independently at the same core distinction between failure as acceptable noise and failure as unacceptable repetition.

The open edge

The stated escape hatch, unless it is catastrophic, is never defined. Rockets have an unambiguous catastrophe threshold, which is what makes this doctrine safe to run there. It is much less clear where that line sits for a medical device, a financial system, or an autonomous vehicle operating on a public road, and a doctrine that works because its failure mode is legible does not obviously transfer to domains where the failure mode is not.

Practiced by

Connections

Loading connections…

References

  1. 01

    How Elon Thinks

    Eric Jorgenson · podcast · 2026

  2. 02

Related