NOFire.ai

Toil

NOFire AI

What counts as toil and how much of it is acceptable?

Toil is operational work that is manual, repetitive, automatable, tactical and devoid of lasting value, and which grows in proportion to the service. The standard ceiling is half an engineer's time, and the reason to measure it is that toil crowds out the engineering that would reduce it.

At a glance

PropertyToilNot toil
RepetitionHappens again and again in the same shapeHappens once
AutomatabilityA machine could do itRequires human judgement each time
Value left behindNone. The system is as it wasLeaves an improvement in place
GrowthScales with the number of services or usersConstant regardless of scale
TriggerReactive, driven by an event or a schedulePlanned, driven by a decision
ExamplesManual failover, restarting a stuck job, hand-applying a config changeDesigning a schema, writing a postmortem, capacity planning

How it works

Toil is a specific category, not a synonym for unpleasant work, and the specificity is what makes it worth measuring.

The five properties usually cited are that it is manual, repetitive, automatable, tactical rather than strategic, and devoid of enduring value. In practice the property that does the most work is the sixth, which is that toil scales with the service. That is what turns it from an annoyance into a structural problem. A task that takes fifteen minutes per service per week is trivial at five services and is a full-time role at two hundred, and nobody ever decides to create that role. It accretes.

The reason to care is a feedback loop rather than the discomfort. Toil consumes exactly the capacity that would be spent building the automation that removes it, so a team past a certain toil load cannot climb out by working harder. Every hour spent restarting jobs is an hour not spent fixing the reason they need restarting, and the load grows while that continues. This is why the standard remedy is a ceiling rather than an aspiration: half an engineer's time, with the rest ring-fenced for engineering that reduces future toil.

Judgement work is the usual boundary case. Deciding whether to fail over is not toil. Executing a failover once decided, through fourteen manual steps, is. Splitting a task into the decision and the execution usually reveals that the tedious part was mechanical all along, which is exactly the part that can be automated without removing the human from the loop.

Not all toil should be automated. Rare toil, or toil attached to a system being decommissioned next quarter, can be cheaper to endure than to engineer away.

How it is measured

Ask engineers to categorise their time rather than trying to infer it from tickets. A simple weekly split between toil, project work and everything else, self-reported, is more accurate than an elaborate scheme nobody completes honestly.

Count instances alongside hours. Twenty minutes of toil occurring forty times a week is a different problem from a thirteen-hour monthly task, even though the totals are similar, because the first fragments attention and the second can be scheduled.

Track the growth rate, which is the number the ceiling is really protecting against. Toil at 30% and rising five points a quarter is a more urgent situation than toil at 45% and flat, and a snapshot cannot distinguish them.

Attribute it to a source. Toil is usually concentrated: a handful of services or one fragile subsystem generates most of it. Ranking sources turns the measurement into an investment decision, and it typically shows that automating the top two sources removes more than a broad programme would. How runbooks get automated and when they should not be covers the mechanics of doing that.

Set a threshold that triggers something specific. A toil percentage that gets reported and never acted on is an audit, not a control.

Common misconceptions

That toil means work engineers dislike. Plenty of toil is mildly pleasant and plenty of essential engineering is miserable. The definition is about growth and residue, not about enjoyment.

That the goal is zero toil. Some operational load is the cost of running things, and the last few percent is usually the most expensive to remove. The ceiling exists to prevent the loop from closing, not to reach zero.

That automating toil always pays back. Automation carries a build and a maintenance cost, and automation that breaks silently creates a worse failure mode than the manual step it replaced. Rare toil often should stay manual.

That toil is an SRE-only concept. It applies to any team carrying operational load. Where a product team owns its own on-call, it accumulates toil the same way and usually measures it less.

That alert response is not toil. Repeatedly triaging alerts that require the same action is textbook toil, and it is often the largest single source. What alert fatigue is and what causes it covers the related failure where the volume stops being actioned at all.

Frequently asked questions

Is all manual work toil?
No. The test is whether it scales linearly with the service and leaves nothing behind. Designing a system is manual and not toil. Restarting a process every Tuesday is toil. A one-off migration is neither.
What is the usual toil ceiling?
Fifty percent of an engineer's time is the widely used cap, with the remainder reserved for engineering work that reduces future toil. The number matters less than having a threshold that triggers a decision.
Why does toil grow?
Because it scales with the thing it supports. Ten services generating one manual task a week each becomes a full-time job at a hundred services, without anybody deciding to create a job.
Is automating toil always worth it?
Not always. Automation has a build and maintenance cost, and toil that occurs rarely or is about to be designed out may be cheaper to endure. The calculation should include how fast the toil is growing, not just its current size.

What to take from this

The defining property is that it scales with the service. Work that is tedious but does not grow as you grow is a chore, and treating every chore as toil makes the measurement useless.

Go deeper: how runbooks get automated

Back to SRE reference