slh

Probe frequency algorithm

The probe frequency is the longest gap between health checks that still lets you detect and repair your expected incidents without going over the downtime budget. It depends on your mean time to repair (MTTR), how many incidents you expect, and how many failed probes it takes to confirm an incident.

probe frequency = buffer ÷ (incidents × probes)
buffer = downtime budget − (MTTR × incidents)
  • Incidents: how many times you expect the service to become unavailable in the period.
  • Probes: how many consecutive failed checks it takes to confirm an outage and fire an alert.
  • Buffer: what's left of the downtime budget after repairing every expected incident. That's the time you can spend detecting them.

Worked example

A 99% service level, 20 minute MTTR, 3 incidents per period and 2 failed probes to alert:

The minimum probing frequency is 6m 48s for a week, 1h 3m 3s for a month, 3h 26m for a quarter and 14h 26m for a year. The daily budget is 14m 24s, but three 20 minute repairs take an hour, so no probing frequency can keep a 99% daily target. For the week, the budget is 1h 40m 48s. Take away the hour of repairs and 40m 48s remains, split across 3 incidents × 2 probes, so you need to probe at least every 6m 48s.

Open this example in the calculator.

This formula is experimental. There's no guarantee it fits your service, so treat the result as a starting point.