Runbook automation
NOFire AI
When should a runbook be automated and when should it stay manual?
Automate a runbook when it is executed often, has a deterministic decision path, and fails safely. Keep it manual when the judgement is the hard part, when it runs rarely enough that the automation will rot, or when a wrong execution is worse than a slow one.
At a glance
| Runbook property | Automate | Keep manual |
|---|---|---|
| Frequency | Weekly or more | A few times a year |
| Decision path | Deterministic, no branching judgement | Requires weighing context each time |
| Failure mode | Safe. A failed run leaves things as they were | Destructive or hard to reverse |
| Blast radius | Bounded and known | Unbounded or unknown |
| Underlying stability | The procedure has been stable for months | The system is changing under it |
| Verification | Success is machine-checkable | Success requires a human to look |
How it works
A runbook is a written procedure for a known situation. Automating one means turning that procedure into code that executes it, and the interesting question is never whether automation is possible but whether it is a good idea for that specific procedure.
The most useful move before deciding is to split the runbook in two. Almost every operational procedure contains a decision, whether to act at all, and an execution, the steps that follow once the decision is made. These have very different automation profiles. The execution is usually mechanical, tedious and safe to automate. The decision often depends on context a script cannot see, such as whether a customer migration is running or whether this is the third occurrence this week.
Teams that automate both together tend to stall, because the decision half is genuinely difficult and it holds up the mechanical half that would have delivered most of the value. Teams that automate the execution and leave a human to approve it usually ship in a fraction of the time, and end up with a system that is easier to get approved as well.
Frequency drives the payback, and it drives the risk too. A procedure that runs weekly is exercised constantly, so drift gets caught. One that runs twice a year is untested code by the time it is needed, which is more dangerous than a stale document because people trust automation more than prose.
The failure mode is the third gate. Automation that fails safely, leaving the system as it was, can be attempted freely. Automation that can leave things half-done needs the same care as a deployment, including a way to bound what a single execution may affect.
How it is measured
Count how often each automation fires, and treat a rising count as a defect report rather than a success metric. Automation that restarts a leaking service twenty times a month is doing its job and telling you something more important, which is that nobody is fixing the leak. This is the single most useful measurement here and the one most often skipped.
Measure success rate and, separately, silent failure. Automation that reports success while achieving nothing is worse than no automation, because it removes the human who would have noticed. Verify the outcome rather than the exit code.
Track time from trigger to resolution against the manual baseline. Some automations turn out to be slower than a practised engineer, particularly where they include conservative waits.
Audit for rot on a schedule. Any automation that has not run in six months should be exercised deliberately, in a safe environment, before it is next relied on. Where that is impractical, consider whether it should exist at all.
Record what each execution did and what it affected. What blast radius is and how it is bounded covers expressing the limit on effect as something enforceable, which is the property that makes automated action approvable in organisations with a change process for humans. What governing an AI agent in production requires covers the same question where the actor is an agent rather than a script.
Common misconceptions
That automating a runbook solves the problem. It makes the symptom cheaper, which reduces the pressure to fix the cause. This is a reasonable trade if it is a decision rather than an accident, and it usually is an accident.
That a documented runbook is nearly automation. The gap is the undocumented judgement. The steps are written down because they are the easy part to write down. The reason a human is still needed is usually absent from the document.
That automation should be unattended by default. Proposing an action for human approval captures most of the time saving with a fraction of the risk, and is often the version that gets deployed at all.
That rarely-used automation is harmless. It is worse than a document, because it carries authority it has not earned. If it cannot be exercised regularly it should be simple enough to read and verify by eye.
That runbook automation reduces toil automatically. It does when the underlying task is genuinely repetitive. Where it becomes another system to maintain, it can shift toil rather than remove it. What counts as toil and how much is acceptable covers making that distinction before committing.
Frequently asked questions
- What makes a runbook a good automation candidate?
- Frequency, determinism and safe failure. A procedure run weekly with no branching judgement and a harmless failure mode pays back quickly. One of those three missing usually means it should not be first in the queue.
- Why do automated runbooks rot?
- Because they are exercised rarely and the system underneath keeps moving. A runbook that runs twice a year is untested code by the time it is needed, which is worse than a document because people trust it more.
- Should automation ask for confirmation?
- For anything with a large effect, yes. A human approving a proposed action is a different risk profile from an agent acting unprompted, and keeping the decision human is often what makes automation approvable at all.
- Does automating a runbook fix the underlying problem?
- No. Automating a restart makes the symptom cheaper and removes the pressure to fix the leak. Track how often automation fires, because a rising count is a defect report.
What to take from this
Split every runbook into the decision and the execution before deciding. The execution is usually mechanical and worth automating; the decision usually is not, and conflating them is why automation projects stall.
Go deeper: what counts as toil
Related answers