Three questions come up whenever a service is being planned or is in trouble. How much downtime can we afford? How many requests are we holding at once? And are things failing fast enough that someone should be woken up? Each has a short formula. Change any number below and the answer follows; every figure is worked out in your browser, and nothing you type is sent anywhere.
What an availability target allows
Formula downtime = (1 − availability) × period
Example 99.9% over 30 days allows 0.1% of 43,200 minutes: 43.2 minutes. Each extra nine allows ten times less.
Why it matters: a target stricter than the time it takes a person to notice and act is a wish, not a target. At 99.99% there are about 4 minutes a month.
How much work is in flight
Formula in flight = requests per second × seconds each takes (Little's law)
Example 200 requests a second at half a second each means about 100 in progress at once. If each slows to 2 seconds, the same traffic needs 400.
Why it matters: worker and connection pools sized for average latency run out exactly when latency rises, which is during an incident.
How fast failures spend the error budget
Formula burn rate = share of requests failing ÷ share allowed to fail; a 30-day budget lasts 30 ÷ burn rate days
Example a 99.9% objective allows 0.1% of requests to fail. If 1.44% are failing, that is a 14.4× burn: the month's budget is gone in about 2 days, so someone should be paged.
Why it matters: alerting on how fast the budget burns pages people when users are being hurt quickly, and stays quiet otherwise. Each alert waits for a long window and a short one to agree, so a brief spike doesn't page and a fixed problem stops paging soon. The windows and rates are the starting point from the Google SRE Workbook's chapter on alerting on SLOs.