MTTR, MTTD, and the incident metrics worth tracking

Incident metrics are easy to collect and easy to misuse. Here's what MTTD, MTTR, and friends actually mean, which ones drive improvement, and how to avoid gaming them.

Spectra Team

Once you’re tracking incidents, the acronyms arrive: MTTD, MTTA, MTTR, MTBF. They’re useful — but only if you know what each measures and resist turning them into a scoreboard. Here’s a practical guide to the incident metrics worth your attention.

The four that matter

  • MTTD — mean time to detect. How long from when something breaks to when you know. This is the metric monitoring most directly improves: faster, multi-region checks shrink it.
  • MTTA — mean time to acknowledge. From alert fired to a human owning it. Reveals alerting and on-call health — high MTTA means alerts aren’t reaching the right person, or are being ignored.
  • MTTR — mean time to resolve (or recover). From detection to service restored. The headline number, but also the fuzziest — be clear whether you mean “recovered” (users are fine) or “fully resolved” (root cause fixed).
  • MTBF — mean time between failures. How often incidents happen at all. Rising MTBF means your reliability work is paying off.

Detection is the cheapest lever

Total downtime is roughly detect + acknowledge + fix. Teams pour effort into fixing faster, but the fastest wins often come from MTTD — you can’t start the clock on recovery until you know there’s a problem. If a chunk of your incidents are found by customers rather than monitors, MTTD is where to invest first.

Don’t let metrics get gamed

Metrics change behavior, not always for the better:

  • Chasing a low MTTR tempts teams to close incidents early or under-declare severity. Guard against it by defining “resolved” clearly and consistently.
  • Averages hide the tail — one ugly six-hour incident and fifty quick ones average out to something reassuring and meaningless. Track the distribution (p90, worst case), not just the mean.
  • Never tie these to individual performance. The moment they’re used to judge people, people stop declaring incidents.

Use them to ask better questions

Metrics are a starting point for a conversation, not a grade. High MTTD? Improve coverage and check frequency. High MTTA? Fix alert routing and on-call load. High MTTR on a specific service? Maybe it needs better runbooks or a rollback path. The number tells you where to look, not whether you’re “good.”

The bottom line

Track MTTD, MTTA, MTTR, and MTBF — but define them precisely, watch the distribution instead of just the average, and never weaponize them. Used honestly, they point you at the highest-leverage fix, and more often than not that fix is detecting faster.

Measure and shorten every incident. Explore incident management →

Start monitoring for free today!

Free forever plan No credit card required
Start for free