FractionalDataArchitect
Book a discovery call

Blame Post-Mortem Alt

Blame Post-Mortem Alt
Blame Post-Mortem Alt

We changed the questions in the postmortem, and the recurring outage stopped recurring.

The same pipeline had broken 4 times in 6 weeks. Every review found a person, none found the cause. Each round produced a name, an apology, and a promise to be careful. Careful is not a control. Everyone in the room knew the fifth break was already scheduled.

We ran the sixth one differently. Same people, same hour, 4 questions:

  1. What did the person see? Reconstruct their screen as it was, from their seat.
  2. What made the wrong action easy? Usually a missing guardrail, a confusing name, a deploy with no dry run.
  3. What made the failure invisible for hours? That’s the detection gap, and it’s often the more expensive half.
  4. What would have to be true for this to be impossible? That’s the fix.

I’ve run this badly too. My first attempt still opened with a timeline of who did what, so the new questions landed as a politer interrogation and everyone stayed defensive. The order of the questions does more work than their wording.

Question 2 found it: a manual step in the deploy nobody had written down, easy to skip under time pressure. We automated it that week.

That pipeline’s run 5 months without an incident. The engineer everyone had been naming wrote the automation that fixed it.

When your last incident was reviewed, did the meeting produce a guardrail or a name?

Written by Thomas Nys

Fractional Data Architect helping startups and scaleups build data platforms that scale.

More about Thomas Nys →

Recognise the problem? Let's talk about it.