Blame Post-Mortem Alt

We changed the questions in the postmortem, and the recurring outage stopped recurring.
The same pipeline had broken 4 times in 6 weeks. Every review found a person, none found the cause. Each round produced a name, an apology, and a promise to be careful. Careful is not a control. Everyone in the room knew the fifth break was already scheduled.
We ran the sixth one differently. Same people, same hour, 4 questions:
- What did the person see? Reconstruct their screen as it was, from their seat.
- What made the wrong action easy? Usually a missing guardrail, a confusing name, a deploy with no dry run.
- What made the failure invisible for hours? That’s the detection gap, and it’s often the more expensive half.
- What would have to be true for this to be impossible? That’s the fix.
I’ve run this badly too. My first attempt still opened with a timeline of who did what, so the new questions landed as a politer interrogation and everyone stayed defensive. The order of the questions does more work than their wording.
Question 2 found it: a manual step in the deploy nobody had written down, easy to skip under time pressure. We automated it that week.
That pipeline’s run 5 months without an incident. The engineer everyone had been naming wrote the automation that fixed it.
When your last incident was reviewed, did the meeting produce a guardrail or a name?
Fractional Data Architect helping startups and scaleups build data platforms that scale.
More about Thomas Nys →