Repeated Success With a Known Flaw Doesn't Raise the Alarm. It Lowers the Bar.

Repeated Success With a Known Flaw Doesn't Raise the Alarm. It Lowers the Bar.

fragilityprocessreliability

Diane Vaughan spent a decade studying why NASA launched Challenger on a freezing January morning in 1986. Her conclusion, published in The Challenger Launch Decision in 1996, was not that managers were reckless or that anybody broke a rule. It was that the signals had gradually stopped counting as signals.

She called it the normalization of deviance, and it is the single most useful idea I know for understanding why competent teams ship systems they privately don't trust.

How a flaw becomes a finding

The mechanism runs on small, individually defensible steps.

Problems with the O-ring joints on the shuttle's boosters showed up as early as 1981, on some of the earliest flights. Each time erosion appeared and nothing catastrophic followed, the observed damage was reinterpreted as being within the bounds of acceptable risk. The engineers weren't hiding anything. They documented the erosion, discussed it, and signed it off, which is precisely the problem. Damage that counted as an anomaly in 1981 had become a documented, expected, signed-off finding by 1985. The bar never jumped. It slid, one reasonable meeting at a time, until a January morning cold enough to finish the job.

So the lesson people usually take is "listen to your engineers," which is true but incomplete. The engineers were listened to. Their findings were in the record. What failed is that repeated success with a known flaw doesn't raise the alarm. It lowers the bar. Every flight that came back safely was treated as evidence the flaw was fine, when it was only evidence the flaw hadn't been fatal yet.

Seventeen years later, the same agency did it again

If normalization of deviance were a lesson you could learn once, the story would end in 1986. It doesn't.

On January 16, 2003, a piece of insulating foam broke off Columbia's external tank about 81.7 seconds after launch and struck the left wing. Foam had been shedding on flights for years, and like the O-ring erosion before it, each strike that failed to kill anyone moved foam loss further into the category of maintenance issue. Columbia broke apart over Texas on February 1. The investigation board's August 2003 report concluded that the culture had as much to do with the accident as the foam did, and one board member described a silent safety program with "echoes of Challenger."

Same agency. Same mechanism. The organization that wrote the book on this failure mode, that trained on it and memorialized it, reran it within a generation. Which is why "we learned our lesson" is not a control. The drift doesn't announce itself, so the defense can't be memory. It has to be structural.

Data teams run the same mechanism weekly

The stakes are lower in analytics. The mechanism is identical.

The Tuesday report breaks and someone fixes it by hand. It breaks again the next Tuesday, and the fix becomes a step in the process. Six months later a new hire asks why that step exists and gets told that's just how it works. Nobody made a bad decision anywhere in that sequence. Each step was the cheapest available fix on the day it was taken, and the report shipped on time every single week, which is exactly the record of repeated success that keeps anyone from escalating. You don't escalate a problem that stopped looking like one.

The same slide happens with numbers. A dashboard total drifts two percent from the source system and someone writes it off as timing. The gap becomes "the reconciliation difference," gets its own cheerful nickname, and eventually appears in the onboarding doc. The day the gap is eleven percent, nobody can say when that became possible, because the alarm that should have caught it was recalibrated months ago, one defensible meeting at a time.

The tell is language

You can hear normalization before you can measure it. Listen for the phrases teams use about their own systems: "that's expected," "it usually settles by noon," "you have to run it twice, everyone knows that," "ignore the first email, the second one is right."

They get said cheerfully, which is what makes them easy to miss. Cheerful is the sound of a workaround that has finished turning into a procedure. Every one of those phrases is a documented, expected, signed-off finding, and each was an anomaly once.

The order to check things in

The practical defense is not vigilance, because vigilance is exactly what the mechanism erodes. It is writing things down while they're still true, so drift has something to be measured against.

First, write down what the system is supposed to do while it still does it: the report arrives at 7am, the totals match the source ledger, the pipeline runs without a rerun. This takes an afternoon, and it is the whole foundation.

Second, inventory the workarounds. Every manual step, every "run it twice," every known gap with a nickname goes on a list with a date. You're not fixing them today. You're refusing to let them become invisible.

Third, when something breaks and the fix is manual, log that the deviation happened even after it's patched. Then a flaw that recurred nine times reads as nine occurrences of one problem instead of nine successful Tuesdays.

A team that has to say out loud what changed will notice the drift. A team that only reacts to failures keeps grading itself against last week, and last week's bar was already lower than the week before.

Your data, our problem.

Work With Us