If an AI agent is going to act in production, it needs to know what it is acting on, why, and what it is allowed to do next. That sounds obvious until you look at how many incident workflows still depend on someone piecing together an answer from dashboards, Slack threads, runbooks, and whatever changed in the last hour.
An experienced SRE can work through that ambiguity. They know which signals to trust, which assumptions to question, and when to stop. Giving an agent access to the same tools does not give it that judgment. It gives it the ability to make mistakes much faster.
At NOFire.ai, we are working on both sides of this problem: the production context an agent needs to investigate an incident, and the controls required before it can take action.
The context an agent needs
The context starts with a continuously updated view of production. What is running? Which services depend on each other? What changed? Which alerts, logs, metrics, traces, deployments, and infrastructure events are relevant to this incident? A useful investigation needs these facts connected in time, with their sources preserved. A spike in errors shortly after a deployment is a lead, not proof that the deployment caused it. The system needs to make that distinction clear.
We call this the Production Graph. Its job is to give an investigation a concrete account of the environment and the evidence available at the time. It should let an engineer inspect the path from an alert to a service, a recent change, a possible failure mode, and the users or systems that might be affected. It should also show where the evidence is incomplete.
This is what I mean by deterministic context. Production itself changes constantly, and an LLM will not become deterministic because we ask it nicely. The point is to give each investigation a versioned, inspectable set of facts. We should be able to answer: what did the agent know when it proposed this action, where did that information come from, and what changed since? Without that, it is difficult to review a decision, let alone trust an agent to make the next one.
Context makes an agent more useful. It does not make the agent safe to execute a plan.
The control system runs outside the agent
For that, the control system has to sit outside the agent. Before an action runs, it must be checked against an explicit set of allowed operations, the current state of production, and the scope of the incident. A credential should grant only the access needed for that action and for a short time. Actions that exceed the permitted scope need to be denied or sent to a human for approval. We also need to observe what actually happened from outside the agent's execution environment. An agent's explanation of its own behavior is evidence of what it said, not proof of what it did.
Runtime policy enforcement for AI agents covers how that check is built in practice, and the Context and Control Model covers how the two halves fit together.
What this looks like in an incident
Consider an agent investigating elevated latency. It finds a recent rollout, compares the timing with service metrics, checks whether the affected instances share that version, and proposes a rollback. The graph supplies the evidence and the likely blast radius. The control system checks whether rollback is an allowed action for this service, whether the target version is valid, and whether the incident meets the conditions for autonomous execution. If any of those checks fail, the agent can still help the engineer investigate. It cannot proceed with the rollback.
Stopping a run that has gone wrong
The same principle applies during execution. If the agent's actions diverge from the approved plan, we need a way to stop them before they turn one incident into two. That is why we care about isolating agent execution in a microVM, monitoring it across that boundary, and being able to pause it and escalate. The boundary gives the control system an independent vantage point. It does not replace good reasoning; it limits what bad reasoning can do.
Why we wrote our own microVMM explains the runtime behind that boundary, and how to sandbox an agent with production write access covers the controls that go around it.
Earning autonomy in stages
This also gives us a sensible way to earn trust. An agent can begin by collecting evidence and drafting an investigation. Then it can handle bounded, repeatable tasks with clear outcomes and an engineer in the loop. As we learn where the evidence and controls hold up, we can allow more actions within defined limits. The scope of autonomy should follow what we can verify, not what a demo makes look convincing.
I have spent a lot of my career working on production readiness, incident response, and SLOs. One lesson keeps coming back: operational confidence comes from knowing the system, defining the limits, and testing what happens when something goes wrong. AI does not change that requirement. It makes the need for it more urgent.
Our goal is an agent that can help an SRE get from an alert to a well supported decision, and eventually carry out suitable actions under enforceable controls. The Production Graph provides the context for the decision. The control system governs the action. We need both before "autonomous SRE" means something I would be comfortable putting in front of production.



