Incident severity levels
NOFire AI
How should incident severity levels be defined?
Severity levels classify an incident by user impact so that the response scales with the harm. The definitions have to be decidable in under a minute by a tired engineer at 3am, which is why impact-based criteria beat technical ones and why most teams need three or four levels rather than five.
At a glance
| Level | Typical impact | Response | Communication |
|---|---|---|---|
| SEV1 | Core function unavailable to most users, or data at risk | Page immediately, all hands as needed, incident commander | Customer-facing status page, executive updates |
| SEV2 | Significant degradation, or full outage of a secondary function | Page the owning team, escalate if unresolved | Status page if externally visible, stakeholder updates |
| SEV3 | Limited impact, workaround exists, or a single customer affected | Handle in hours, business hours acceptable | Internal channel only |
| SEV4 | Minor, cosmetic, or internal-only | Ticket, next sprint | None |
| Not an incident | Alert fired, no user impact confirmed | Tune the alert | None |
How it works
A severity level is a routing decision compressed into a label. It determines who is woken, how many people join, whether customers are told, and how much of the organisation's attention the problem gets. Getting it right early is worth more than getting it exactly right eventually.
The most important design property is that the criteria must be decidable quickly by someone under stress with incomplete information. That single constraint rules out most of what teams instinctively write. Criteria referencing internal architecture fail it, because at the moment of declaration nobody knows which component is at fault. Criteria requiring a measurement that takes twenty minutes to gather fail it too.
What survives the constraint is impact framed from the user's side. Can users do the main thing the product exists for? Is there a workaround? Is data being lost or exposed? Those are answerable in seconds from what a responder can already see, and they map directly onto how much response the situation deserves.
Severity should be assigned by the person declaring, immediately, and it should be correctable. Teams that require agreement before assigning a level lose exactly the minutes that a SEV1 cannot afford. A culture where over-declaring is safe produces faster responses than one where declaring a SEV1 invites scrutiny, because the cost of an unnecessary page is far lower than the cost of a delayed one.
Keeping severity separate from priority avoids a recurring argument. Severity describes the harm and does not improve because you are handling it well. Priority describes the effort and can drop once a mitigation is in place. Merging them leads to incidents being downgraded while the impact continues.
How it is measured
Write the criteria as questions with yes or no answers, and put them where a responder will actually see them, which is the incident tooling rather than a wiki page nobody opens at 3am.
Test them against history. Take the last twenty incidents, hand the criteria to two engineers separately, and ask each to assign a level from what was known in the first five minutes. Disagreement rate is your measurement. Above roughly one in five, the definitions are ambiguous and the resulting severity data will not support any analysis you try to do with it later.
Track the distribution over time. A healthy organisation has many low-severity incidents and few high ones. If almost everything is SEV3, either the criteria are too strict or real incidents are being under-declared to avoid attention, and the second is a cultural problem that no wording fixes. If everything is SEV1, the levels are not doing their job of allocating attention.
Track re-classifications too. Incidents that start SEV3 and become SEV1 are the interesting ones, because the pattern usually reveals a class of failure whose impact is not obvious from the first symptom. How alert triage works and how it is measured covers the upstream half of this, where the decision about whether something is an incident at all gets made.
Common misconceptions
That severity should reflect technical difficulty. A one-line fix that restores a broken checkout is a SEV1. A complex problem affecting nobody is not. Severity is about harm, and effort is a separate axis.
That five levels are standard. Five is common and often two of them are dead. Unused levels are worse than absent ones because they make historical severity data inconsistent between the periods when people used them and the periods when they did not.
That declaring a high severity is an escalation against someone. Where it feels that way, incidents get under-declared and response gets slower. The fix is cultural rather than definitional, and it usually starts with leadership visibly treating an over-declaration as a good outcome.
That severity can be assigned once and forgotten. Impact changes as an incident develops, and a SEV3 that spreads is a SEV1. Re-assessing at a fixed cadence during long incidents catches this.
That severity determines whether you write it up. A near-miss with no impact can be the most instructive incident of the quarter. What a post-incident review should produce covers choosing what deserves a write-up, and severity is only one input to it.
Frequently asked questions
- How many severity levels should we have?
- Three or four for most organisations. Five is common and usually means the bottom two are never used, or are used inconsistently, which makes severity data useless for any analysis afterwards.
- What is the difference between severity and priority?
- Severity describes how bad the impact is and does not change as you work. Priority describes what you do about it and can change. Conflating them is why teams end up arguing about downgrading a SEV1 that is under control.
- Who decides the severity?
- Whoever declares the incident, immediately, from written criteria. Committee decisions waste the minutes that matter most. Severity can be corrected later, and a wrong initial call is much cheaper than a delayed one.
- Should severity be based on affected user count?
- Partly, but not only. A small number of users unable to do anything can be worse than many users experiencing mild slowness. Good criteria combine breadth of impact with depth of impact and whether a workaround exists.
What to take from this
Define levels by what users lose, not by which system broke. If two reasonable engineers disagree about the severity of the same incident, the definitions are the problem, not the engineers.
Go deeper: how alert triage is measured
Related answers