Cleric vs Resolve AI
NOFire AI
Should we choose Cleric or Resolve AI for AI-driven incident investigation?
Resolve AI puts teams of agents into the on-call rotation and encodes your runbooks as Skills, so institutional knowledge gets applied consistently. Cleric is read-only by default with every investigation auditable, and accumulates operational memory so a resolved incident informs the next one.
VerdictResolve AI if you have real runbooks worth encoding and want agents inside the rotation. Cleric if the blocker to adoption is what an agent might do to production, because read-only by default answers that by construction.
At a glance
| Cleric | Resolve AI | |
|---|---|---|
| Core idea | Read-only investigation with an auditable trail and accumulated memory | Teams of agents in the on-call rotation, carrying encoded runbooks |
| Default authority in production | Read-only by default | Scoped through integration permissions, SAML SSO and RBAC |
| Encoding team knowledge | Not the centre of the product | Skills, the headline capability, plus MCP and an API |
| Learning across incidents | Operational memory, so a resolved incident informs the next | Not published as a distinct mechanism |
| Auditability | Every investigation auditable, and the product leads on this | Not published as a distinct mechanism |
| Published accuracy | Not published against a public benchmark | Not published against a public benchmark |
| Published outcome claim | Not the primary claim | Up to 5x faster MTTR, self-reported |
| Compliance posture | SOC 2 Type II, states customer data is never used for training | SAML SSO and RBAC, commits that customer data is not used to train models for others |
| Funding and availability | Smaller than Resolve AI | Best funded in this group. US and Canada |
How the two differ in practice
These two answer the same brief in opposite directions, which makes them unusually easy to separate once you know what you are optimising for.
Resolve AI is organised around the responder. The product's claim is that agents join the rotation, work alerts and incidents alongside engineers, and apply the runbooks your team has already written through Skills. That addresses a specific and real failure: the knowledge existed, somebody documented it, and at 2am under pressure nobody found or followed it. Encoding it so an agent applies it consistently is a genuinely different bet from trying to reason the answer out from first principles.
Cleric is organised around trust. Read-only by default is the design decision everything else follows from, and it changes the conversation with whoever has to approve the rollout. Instead of demonstrating that permissions are scoped correctly, the answer is that the product cannot write to production at all. Every investigation being auditable is the second half of the same posture, and SOC 2 Type II with a statement that customer data is never used for training is the third.
The memory model is the other real difference. Cleric accumulates operational memory so a resolved incident informs later ones. That compounds with volume, which means it is worth substantially more to a team having incidents weekly than to one having them quarterly, and it is a slow-burning benefit that a two-week trial will understate.
Neither publishes accuracy against a public benchmark. Resolve AI publishes an outcome figure from its own deployments, up to 5x faster MTTR, which is a claim about what it measured rather than a number you can compare against another vendor.
Where each one is stronger
Resolve AI is stronger where institutional knowledge is the asset and the rotation is the process. A team with real runbooks, a real on-call schedule and a specific set of recurring failure modes gets the most out of Skills, because the encoding work has a clear return. It is also the best funded product in this group, which matters more than teams like to admit when buying from an early-stage vendor, and the operational automation piece takes on toil that sits entirely outside incidents.
Cleric is stronger wherever the obstacle is organisational rather than technical. In an environment with a change advisory process for humans, the question that stalls an AI rollout is not accuracy but authority, and read-only by default retires that question before it is asked. The auditable trail matters for the same reason, because it turns the agent's work into something reviewable after the fact. The memory model is the third advantage and the one that grows.
The shared limit is worth stating plainly. Both reason over the telemetry, deploy history and incident records you already keep, so a thin observability stack produces a thin investigation in either product. Neither can see what nothing emits, and no amount of model quality substitutes for a signal that was never collected.
How to choose
Decide first whether your blocker is knowledge or authority, because that separates these two immediately. If your incidents are mostly recurrences of things somebody already understands, Skills is the direct answer. If your incidents are novel but your security review is the thing actually preventing adoption, read-only by default is the direct answer.
Then run whichever you pick against incidents whose true cause you already know, and score the first hypothesis rather than the eventual one. Neither vendor publishes a comparable accuracy number, so the only measurement available is the one you take yourself. The AI SRE Benchmark sets out how that scoring works on a public dataset if you want a method to copy rather than inventing one.
Write down the authority bound before the trial rather than after. What governing an AI agent in production requires covers what a testable version of that requirement looks like. Products that cannot express it will not acquire the ability during a pilot.
If neither shape fits, the field is wider than these two. The wider field, with where each tool is blind names the rest of it honestly.
Frequently asked questions
- Can an AI investigation tool change anything in production?
- It depends on the product and it is the first question to ask. Cleric is read-only by default, so the answer is structural rather than a configuration you have to get right. Others scope this through integration permissions and RBAC.
- What are Resolve AI Skills?
- Skills encode a team's runbooks so agents apply them consistently. They solve a real problem, which is that the knowledge usually existed and nobody applied it at 2am. They are also specific to Resolve AI and do not transfer.
- Does either publish accuracy against a public benchmark?
- Neither does. Resolve AI publishes an outcome figure of up to 5x faster MTTR from its own deployments. Cleric leads on auditability rather than on an accuracy number. The two claims are not comparable.
- What does operational memory actually change?
- It means a resolved incident informs the next investigation rather than every one starting cold. The benefit compounds with incident volume, so it is worth more to a busy estate than to a quiet one.
Go deeper: the AI SRE Benchmark
Book a demo