What to gather first.
Five numbers, all of which already exist somewhere in your tooling. Incidents per month that paged a human. Average people pulled into one incident. Average hours from page to root cause identified. Fully loaded hourly cost of an engineer, which finance has and you should not estimate yourself. And open or planned SRE and platform roles.
One rule runs through the whole worksheet: every number should be one you can defend to the person who approves the spend. If a figure needs an assumption, write the assumption next to it.
What incidents cost you now.
Incidents per month, times people per incident, times hours to root cause, times hourly cost. That gives a monthly figure and a yearly one, and the rest of the worksheet works against it.
Two costs are deliberately left out of that total. Context switching is real, but the recovery multiplier for your team is not something a worksheet can tell you. Escalation concentration is real too: if the same two people are in most incidents, that is a retention exposure that does not show up as an hourly cost until they leave. Both belong in the conversation, not in the arithmetic.
One sanity check. If the yearly figure looks small, recount the people. Teams routinely undercount everyone who joins an incident channel, looks at a graph, and says something useful.
Three reclaim scenarios.
Work all three at 20, 40 and 60 percent rather than picking one. The point of running three is that the decision should hold at the conservative end.
Which band is yours depends on where the hours actually go. If most of the time is spent finding and assembling information, you are in the 40 to 60 band, because that is the work a live model of production removes. If most of it is waiting, for a rollback, for a vendor, for a slow query, then you are at 20 or below and the honest read is that this category is not your bottleneck. A different purchase would help more.
Commit to the 20 percent figure and hold the 40 percent figure as the expectation. A case that overshoots on the first review is harder to recover than one that lands conservative and beats it.
Hire, or buy.
The comparison most business cases actually turn on, as two columns you fill in yourself: fully loaded yearly cost against yearly licence, months to hire against weeks to useful output, recruiting cost against trial cost, and how many timezones each one actually covers.
Three things the totals will not show:
- A hire is a permanent commitment and a licence is not. Weight that asymmetry in favour of trying the tool first, not in favour of the tool being better.
- A hire adds judgement. A tool removes the retrieval work currently consuming your senior engineers' judgement. If your actual gap is judgement, hire.
- Coverage is not linear. One more engineer does not give you a second timezone.
Governance exposure.
This page of the worksheet has no arithmetic on it, on purpose. Anyone who hands you a dollar figure for these is inventing it.
Four questions instead. How many agents in your environment hold credentials today. Whether you could reconstruct what one of them did last Tuesday, from a record the agent itself did not write. Whether an auditor has asked, or will within twelve months. And whether a named person owns this.
The cost of an ungoverned agent action is either zero or very large, and which one you get is not something a spreadsheet predicts. Present it as an unresolved risk with an owner and a date. That is more credible in front of a finance committee than a figure you cannot source, and harder to argue with.
The numbers tell you whether the purchase is worth making. They do not tell you which vendor clears the bar. For that, the AI SRE Buyers Guide has the gate questions and the evaluation worksheet with its eight must-pass rows. For the path decision behind it, see Build vs Buy.