Site reliability engineering, answered
Look up a concept behind AI SRE, how it works, how you measure it, and what people get wrong about it.
14 pages
How can I improve the speed of root cause analysis when debugging complex cascading failures within a Kubernetes cluster?
Start with the layer, not the logs. Reaching for kubectl logs on a pod that never started is the most common wasted first move in Kubernetes debugging.
How does automated root cause analysis actually work?
Judge any RCA system on its first answer, on incidents whose true cause you already know. Every other measurement flatters the tool.
How does causal AI find the root cause of a production incident?
The distinction that matters operationally is not accuracy in the abstract. It is whether the output points at a change you can act on, or at a correlation you still have to investigate.
How should incident severity levels be defined?
Define levels by what users lose, not by which system broke. If two reasonable engineers disagree about the severity of the same incident, the definitions are the problem, not the engineers.
What are the four golden signals and what should they alert on?
Use them to decide what pages you, not what you graph. Most estates collect far more than four signals and would be better off alerting on fewer.
What counts as toil and how much of it is acceptable?
The defining property is that it scales with the service. Work that is tedious but does not grow as you grow is a chore, and treating every chore as toil makes the measurement useless.
What is a service level objective and how is it set?
Set them from what users actually notice, not from what is easy to measure. An SLO nobody would act on when it is breached is documentation, not an objective.
What is alert triage and how is it automated?
Suppression lowers the page count and hides the same failures. The test of a triage system is what it lets through, not what it silences.
What is an error budget and how does a team spend one?
The mechanism only works if the policy has teeth. A budget with no agreed consequence when it runs out is a chart that makes engineers feel worse without changing anything.
What is blast radius and how is it calculated?
Blast radius is only useful when it is a bound rather than a number on a dashboard. The question is what the system refuses to do when the figure comes back too high.
What is change failure rate and how is it measured?
Read it alongside deployment frequency, never alone. A team deploying twice a year with a low rate is not outperforming one deploying daily, it is taking a different kind of risk.
What is mean time to resolution and how is it measured?
Useful as a trend on a consistent definition, misleading as a benchmark between organisations. Publish the definition alongside the number or the number means nothing.
What makes an incident postmortem worth writing?
Blameless is the well-known half and the easier half. The hard half is whether anything changes, and most organisations measure the first and fail the second.
When should a runbook be automated and when should it stay manual?
Split every runbook into the decision and the execution before deciding. The execution is usually mechanical and worth automating; the decision usually is not, and conflating them is why automation projects stall.
NOFire AI appears on some of these pages. It names the change behind an incident, and scores 89% Top-1 on RCAEval, a public benchmark of 735 scenarios anyone can rerun. Before an agent acts, a gate checks the action against policy and its predicted blast radius. It then holds the action for a person, or refuses it. Every other figure on these pages comes from the vendor that published it.
Trusted by