NOFire.ai

Error budgets

NOFire AI

What is an error budget and how does a team spend one?

An error budget is the amount of failure a service level objective permits. A 99.9% target over 30 days allows about 43 minutes of error. The budget exists to settle one recurring argument in advance: when reliability work should outrank shipping features.

At a glance

Objective over 30 daysBudget as timeBudget as requestsWhat that feels like
99%About 7 hours 18 minutes1 in 100A bad afternoon each month is within target
99.5%About 3 hours 39 minutes1 in 200One significant incident a month
99.9%About 43 minutes1 in 1,000One short incident, or several blips
99.95%About 21 minutes1 in 2,000Little room for a single bad deploy
99.99%About 4 minutes1 in 10,000Human response is already too slow
99.999%About 26 seconds1 in 100,000Only achievable with automated failover

How it works

An error budget is the inverse of a service level objective. If the objective says 99.9% of requests must succeed over a rolling 30 days, the budget is the remaining 0.1%, and it is a quantity you are permitted to spend rather than a threshold you must never cross.

That reframing is the entire contribution. Reliability arguments usually run as a contest between engineers who want to stop and fix things and product owners who want to ship, and neither side has a shared unit to argue in. A budget provides one. Below the line, shipping continues. Above it, an agreed policy takes over. The decision is made once, in calm conditions, rather than repeatedly during incidents when everyone is tired and defensive.

The spending mechanism is worth taking literally. A budget that finishes the month untouched is not a triumph, it means you were more conservative than your own objective required and probably shipped less than you safely could have. Teams that internalise this start using the budget deliberately, spending it on a risky migration or an aggressive rollout because they have quantified headroom.

Burn rate is how you watch it in real time. Spending the month's budget evenly is normal. Spending a third of it in an hour is an incident whether or not anyone has declared one. Alerting on burn rate rather than on absolute consumption is what makes a budget operationally useful instead of a monthly retrospective artefact, and the usual implementation pairs a fast window against a slow one so a brief spike does not page anybody.

The policy is where most implementations fail. A budget with no consequence attached produces guilt, not change.

How it is measured

Compute consumption from the same signal as the objective, at request level, over the same rolling window. Mixing sources between the two, for instance an objective measured at the load balancer and a budget computed from application logs, produces two numbers that disagree and a quarterly argument about which is right.

Alert on burn rate with at least two windows. A common arrangement pages when a fast window shows a high multiple of steady spend, and raises a lower-priority signal when a slower window shows sustained moderate overspend. The first catches the outage, the second catches the slow leak that would otherwise exhaust the budget invisibly.

Attribute spend where you can. Knowing that 60% of last month's budget went on one dependency, or on deploys from one service, turns the number into a decision about where to invest. Undifferentiated consumption tells you that things went badly and nothing about what to do.

Track how often the policy actually triggers. If the budget has never been exhausted, the objective is too loose to inform anything. If it is exhausted every month and releases never pause, the policy is not real and everyone knows it. How service level objectives are set covers correcting the target when either happens.

Common misconceptions

That the budget is a failure allowance to be minimised. It is a resource with an optimum spend near full. Consistently finishing with most of it unused means the objective and the engineering effort are mismatched.

That the policy has to be a release freeze. A freeze is the best-known response, not the only sensible one. Redirecting a fixed share of the next sprint to reliability work, or requiring extra review on risky changes, can work better and is easier to sustain politically.

That budgets replace judgement. A single incident that harmed a small number of users badly may matter more than a diffuse one that consumed more budget. The number informs the decision, it does not make it.

That every service needs its own budget. Budgets follow objectives, and objectives belong on user-visible journeys. Internal components with no user-facing failure mode generate numbers nobody acts on.

That exhausting the budget means somebody was careless. Sometimes it means a dependency failed, or a deliberate risk was taken and did not pay off. What change failure rate measures is the better metric for asking whether the problem is in how changes are made.

Frequently asked questions

How much failure does 99.9% allow?
About 43 minutes over a rolling 30 days, or roughly 0.1% of requests depending on whether the objective is time-based or request-based. 99.99% cuts that to about four minutes, which is why each nine costs so much.
What is burn rate?
How fast the budget is being consumed relative to a steady spend. A burn rate of 10 means you are spending ten times faster than the window allows, so a 30-day budget disappears in three days. It is the number worth alerting on.
What should happen when the budget is exhausted?
Whatever you agreed in advance, and it should be specific. The standard policy pauses feature releases in favour of reliability work until the budget recovers. The value is in having decided before the argument, not in the pause itself.
Can an error budget be spent deliberately?
Yes. An unspent budget means you were more cautious than the objective required. Deliberately spending it on a risky migration is a legitimate use, not a failure.

What to take from this

The mechanism only works if the policy has teeth. A budget with no agreed consequence when it runs out is a chart that makes engineers feel worse without changing anything.

Go deeper: how service level objectives are set

Back to SRE reference