NOFire.ai
Resources/SRE reference/Service level objective

Service level objectives

NOFire AI

What is a service level objective and how is it set?

A service level objective is a target for how often a service must meet a measured quality bar over a window, such as 99.9% of requests succeeding in 30 days. It turns reliability from an opinion into a number, and its main job is deciding when to stop shipping features and fix things.

At a glance

TermWhat it isWho it is forTypical form
SLIThe measurement itselfEngineersProportion of requests served under 300ms
SLOThe internal target for that measurementEngineering and product99.9% over a rolling 30 days
SLAA contractual promise, with penaltiesCustomers and legal99.5%, with service credits below it
Error budgetThe failure the SLO permitsEngineering and product0.1% of requests, about 43 minutes per 30 days
Burn rateHow fast the budget is being consumedOn-call10x means the month's budget goes in three days
Availability targetInformal shorthand, often confused with an SLAEveryone, unhelpfully"Five nines", usually unmeasured

How it works

An SLO is built from three decisions, and the order matters because each constrains the next.

The first is what to measure. A service level indicator has to reflect something a user would notice, which usually means request success rate, latency at a percentile, or freshness of data. The common failure is choosing a signal because it is easy to collect rather than because it corresponds to experience. CPU utilisation is trivially measurable and tells you nothing about whether anybody was served.

The second is the target and the window. 99.9% over a rolling 30 days is the workhorse, and each additional nine costs roughly an order of magnitude more engineering effort for a difference most users cannot perceive. The window matters as much as the number: a quarterly window forgives a bad week that a weekly window would surface immediately, and rolling windows avoid the effect where everything resets helpfully on the first of the month.

The third, and the one that makes the exercise worth doing, is what happens on breach. An SLO is a decision rule about when reliability work outranks feature work. The standard mechanism is an error budget: the SLO permits a quantity of failure, and when that quantity is spent, the policy says feature releases pause until it recovers. Without an agreed consequence, an SLO is a chart.

The measurement side needs care too. Server-side success rates miss failures where the request never arrived, so a service can report a healthy SLI while users see errors. Measuring at the load balancer or from synthetic probes catches more, at the cost of attributing failures that were not yours.

How it is measured

Start from a small number of user journeys rather than from your service inventory. Three or four SLOs covering what users actually do beats forty covering what you happen to run.

Compute the SLI from request-level data over a rolling window rather than from uptime checks. A ping every minute cannot distinguish a service returning errors to 20% of real traffic from one that is entirely healthy, and most historical availability figures are built on exactly that mistake.

Track burn rate, not only the current position. A budget consumed slowly across a month is a different situation from the same budget consumed in four hours, and alerting on burn rate catches the second while ignoring the first. Multi-window burn alerts, typically a fast window and a slow one, are the standard way to avoid paging for a blip while still catching a sustained problem.

Review the target on a schedule. An SLO set two years ago against a system that has since changed shape is measuring something nobody chose. If you have never breached it, it is too loose to be informative. If you breach it constantly and nothing happens, it is too tight and it is training people to ignore it. How error budgets work and how they are spent covers the mechanism that turns the target into behaviour.

Common misconceptions

That more nines are better. Reliability past the point users notice is spent capital. If a service's users are on mobile networks that fail more often than the service does, further nines are invisible to them and expensive to you.

That an SLO is a promise. It is an internal target. The customer-facing promise is the SLA, and it should be set looser so that breaching the SLO is a warning rather than a contractual event.

That every service needs one. SLOs on internal components with no user-visible failure mode generate noise and dilute attention. Write them where a breach would mean something.

That uptime and availability are the same. A process can be running and serving errors. Uptime measures the former and availability should measure the latter, and conflating them is how organisations report figures nobody in support recognises.

That the SLO is the work. Setting one is an afternoon. Making it true is the year, and it is mostly about reducing how long failures last rather than preventing every failure. How mean time to resolution is measured covers the half of the equation SLOs tend to make visible first.

Frequently asked questions

What is the difference between an SLI, an SLO and an SLA?
An SLI is the measurement, such as the proportion of successful requests. An SLO is the internal target for that measurement. An SLA is a contractual promise to a customer, usually set looser than the SLO so you have margin.
How many nines should we aim for?
Fewer than instinct suggests. Each additional nine costs roughly an order of magnitude more, and 99.9% allows about 43 minutes of error per 30 days. Pick the level users would actually notice falling below.
Should every service have an SLO?
No. SLOs are worth writing for services whose failure users feel. Giving every internal component a target produces a wall of numbers nobody reads and dilutes attention from the handful that matter.
What happens when an SLO is breached?
Something specific, agreed in advance, or the SLO is decorative. The usual answer is that feature work pauses in favour of reliability work until the budget recovers. If nothing changes, do not bother setting one.

What to take from this

Set them from what users actually notice, not from what is easy to measure. An SLO nobody would act on when it is breached is documentation, not an objective.

Go deeper: how error budgets work

Back to SRE reference