MTTR Reduction Tactics

A broken pipeline costs you an hour. A wrong number nobody catches costs you a quarter.
By 2015, DORA was using four software delivery metrics. Data teams borrowed the vocabulary and mostly kept measuring uptime.
I use time to restore for data incidents, with one change to the definition. It runs from the moment a wrong number is visible to a user, to the moment it’s corrected and that user knows.
The pipeline gets fixed quietly, the dashboard heals overnight, and nobody tells the person who already made a call on the bad figure. That gap is usually most of the number.
I’d start here, in this order.
Put freshness and row-count checks on the source table, so detection doesn’t wait for an analyst to open a dashboard and frown at it.
Give each dataset an owner who can be paged.
Then the one people argue about. As soon as you understand the impact, put someone on drafting the correction note while the fix is still running. Teams resist it because it feels like admitting fault in writing mid-incident. It’s also the only step that makes you name who was affected, which is the part that shortens the loop.
Send me your last data incident and I’ll tell you where the time actually went.
Fractional Data Architect helping startups and scaleups build data platforms that scale.
More about Thomas Nys →