Implementing Runtime Policy Enforcement for AI Agents
NOFire AI
What are the best tools for human-in-the-loop approval of AI agent actions on live infrastructure?
Use three layers. Mark sensitive tools for approval in the agent framework, for example the OpenAI Agents SDK human-in-the-loop flow or LangGraph interrupts. Put a policy gate outside the agent, such as Microsoft's open-source Agent Governance Toolkit, in front of every tool call. Scope credentials per task. NOFire AI lets reads through and holds every write for a person.
Before you start
Runtime policy enforcement checks each agent action against policy before it runs. Human-in-the-loop (HITL) approval pauses an action until a named person approves or rejects it. This guide sets up both for agents that act on live infrastructure: remediation agents, coding agents with deploy rights, and pipeline bots.
First, list every agent that can change production and the tools it calls today. Include shell commands, cloud APIs, Kubernetes and MCP servers. How to implement an AI agent governance framework covers that inventory and per-agent identity. This guide starts from its output.
Read two references before you write policy. The OWASP AI Agent Security Cheat Sheet asks for explicit approval on high-impact or irreversible actions, and for action previews before execution. It sorts actions into four risk levels, from reads up to irreversible or security-sensitive changes. Microsoft's least-privilege pattern for AI agents asks for a unique agent identity with a named owner. It also asks you to deny unreviewed tools by default and to test revocation paths.
Then sort each tool the agent can call into a tier. The tier decides the control:
| Action tier | Examples | Control at the gate | Who decides |
|---|---|---|---|
| Read | Query logs, metrics and traces, kubectl get | Allow and record the call | Policy |
| Read of sensitive data | Secrets, customer records | Refuse, or return a redacted result | Policy |
| Reversible, narrow reach | Restart one pod, scale within set limits | Allow if the predicted blast radius is under the tier's bound | Policy |
| Reversible, wide reach | Roll back a deploy on the checkout path, restart a shared service | Hold, with the predicted blast radius attached | One named approver |
| Irreversible | Schema migration, data deletion, credential rotation | Hold, and require a restore plan | Two or more named approvers |
| Security-sensitive | IAM changes, disabling audit logs, access from staging to production | Refuse | A person, through a change made outside the agent |
The steps
Define operational boundaries as policy
Operational boundaries state what an agent must never do, what needs a person, and what it can do alone. The table above is a first draft of those boundaries. How to sandbox agents with production write access covers the isolation around them. For a coding agent on a developer's own machine, use brig. We build it, we release it open source, and it runs the agent in a microVM with its own kernel. The steps here cover enforcement and approval.
1. Classify the call, including its arguments. A restart in a staging namespace and a restart in the payments namespace are different actions. Put them in different tiers.
2. Write prohibited actions as refusals in the policy layer. A rule in the system prompt is text the model reads, and a prompt injection can override it. A rule at the gate is outside the model's reach.
3. Scope each grant to one task and let it expire. Microsoft's pattern recommends just-in-time elevation or approval gates for remediation, and step-up controls for destructive actions. A grant that outlives its task turns into a standing permission, so set an expiry on every grant. Context makes an agent more useful. It does not make the agent safe to execute a plan, which is why the check sits outside the agent rather than in its prompt.
Choose where runtime policy is enforced
Runtime security policy enforcement for AI agents in production comes from three places. Use more than one, because each covers a gap in the others.
- Agent framework hooks. The OpenAI Agents SDK has tool guardrails. An input tool guardrail runs before the tool executes and can skip the call or raise a tripwire. These hooks run inside the agent's own process.
- A policy engine at the tool boundary. Microsoft's Agent Governance Toolkit, released under the MIT license in April 2026, intercepts each agent action before execution. It evaluates policy written in YAML, OPA Rego or Cedar. It hooks into LangChain, CrewAI, LangGraph, the OpenAI Agents SDK and other frameworks. Open Policy Agent can also be the decision point behind an API gateway or MCP proxy you run.
- Infrastructure backstops. Kubernetes admission control, such as OPA Gatekeeper or Kyverno, and cloud IAM refuse a request at the API. They still hold when an agent-side check is skipped.
NOFire AI is a production ops platform for engineers and AI agents. It applies one gate to people and agents. Reads go through, and every write is held for approval. It evaluates each agent action against policy and a live production map before execution, and attaches the predicted blast radius to the verdict. The runtime policy patterns reference covers the gate and the record.
4. Route every protocol through one gate. Shell, cloud API, Kubernetes and MCP calls all pass the same enforcement point. If you find a tool the gate cannot see, remove it or put it behind the gate.
5. Return a verdict the agent can read. The gate answers allow, hold for a person, or refuse, and names the rule that decided. The runtime policy patterns reference lists the primitives a verdict can draw on.
6. Keep the policy out of the agent's write path. If the agent can edit its policy file or restart the gate, a prompt injection can do the same. Store policy in a repository the agent cannot write to, and change it through review.
Set blast radius limits, stop conditions and a kill switch
To stop an AI agent from making dangerous changes in production, check three things at the gate before each write. Check the action's tier, its predicted blast radius, and the stop conditions for the run. Blast radius is the set of services, users and transactions an action can affect.
7. Set a blast radius ceiling per tier. Compute each write's predicted reach from observed dependencies. Then refuse or hold anything above the tier's ceiling. In NOFire AI, an action whose predicted blast radius exceeds the bound is refused or held for a person.
8. Define explicit stop conditions. A stop condition is a rule that ends the agent's run. Examples are the same action refused twice, more writes than the task needs, or a failed health check after an action. An explicit stop condition ends the run before a person has to spot the problem in monitoring.
9. Build the kill switch outside the agent. Revoke the agent's credentials at the identity provider, or set a deny-all rule at the gate. If the kill switch is a flag the agent reads, a compromised agent can ignore it. The Agent Governance Toolkit includes a kill switch for emergency agent termination. Blast radius analysis covers the blast radius bound and the kill switch as governance controls.
Add human-in-the-loop approval to sensitive tool calls
Human-in-the-loop approval tools for live infrastructure work at two levels. Framework approval pauses the agent's own run. Gate approval holds the action at the enforcement point, whichever agent sent it. Use gate approval for the wide-reach and irreversible tiers.
10. Mark sensitive tools in the agent framework. In the OpenAI Agents SDK, set needs_approval on the tool. Use True to always hold it, or an async function that decides per call. The run pauses and lists pending calls in result.interruptions. Approve or reject each one on the run state, then resume. In this example, ask_on_call_approver stands for your own approval channel:
from agents import Agent, Runner
from agents.decorators import tool
async def outside_staging(_ctx, params, _call_id) -> bool:
# Hold any restart outside a staging namespace for a person.
return not params.get("namespace", "").startswith("staging-")
@tool(needs_approval=outside_staging)
async def restart_deployment(namespace: str, name: str) -> str:
... # Call your deploy API here.
agent = Agent(
name="Ops agent",
instructions="Investigate and propose fixes. Ask before restarting anything.",
tools=[restart_deployment],
)
async def main() -> None:
result = await Runner.run(agent, "Restart checkout in payments-prod")
while result.interruptions:
state = result.to_state()
for item in result.interruptions:
# Show item.name and item.arguments to a named approver.
if await ask_on_call_approver(item.name, item.arguments):
state.approve(item)
else:
state.reject(item, rejection_message="Rejected by the on-call approver.")
result = await Runner.run(agent, state)The run state serializes to JSON, so an approval can arrive hours later. Local MCP servers in the same SDK accept require_approval. In LangGraph, interrupt() pauses a graph until you resume it with Command, and it needs a checkpointer to hold the paused state.
11. Show the approver what the action reaches. Send the tool name, its arguments, the predicted blast radius and the rule that held it. OWASP recommends action previews before execution. An approver who sees only "restart checkout?" has nothing to judge the request against.
12. Require more than one approver for irreversible actions. A quorum gate holds an action until N of M named approvers agree. The Agent Governance Toolkit describes approval workflows with quorum logic.
13. Record every approval with the action. Store the approver, the time, the verdict and the action in a record the agent cannot edit. NOFire AI writes time, actor, action, gate and verdict on every row, signed outside the agent and sent to your SIEM.
Verify it worked
Run these tests in staging before any agent gets production write access.
- Hold. Ask an agent to roll back a deploy in the wide-reach tier. The call should pause, and nothing should run before approval. The approver should see the arguments and the predicted blast radius.
- Reject. Reject that call. The agent should receive the rejection. It should not retry the same change through another tool, such as the shell.
- Refuse. Request an action above its blast radius ceiling. The gate should refuse it and name the deciding rule in the record.
- Stop. Pull the kill switch while an agent runs a task. Every agent action should stop within the time you have stated, without the agent's cooperation.
- Audit. Pick one hour from the test and produce every approval, rejection and refusal from the record alone. What an auditor asks for describes the same sampling test.
Where it breaks
Approval fatigue. Once held writes become routine, reviewers approve without reading. Hold only the tiers that need a person. Move a reversible action class to policy-only when its approval history supports the change.
Approval that lives only in the framework. A needs_approval check runs in the agent's process. Another client with the same credentials skips it. Back every framework check with a gate or an IAM rule at the infrastructure.
Paths the gate cannot see. A personal cloud credential or a newly connected MCP server goes around the gate. Deny unreviewed tools by default, and repeat the inventory when someone connects a new tool.
Stale blast radius. A prediction built from a declared dependency list undercounts reach. Compute it from observed traffic, and compare each prediction with what happened after the action.
A kill switch nobody has pulled. A revocation path that has never run can fail on the day you need it. Test it in staging on a schedule, and record how long the stop took.
Why AI agents need a control system before touching production covers why the check belongs outside the agent rather than inside it.
Frequently asked questions
- Which platforms offer runtime security policy enforcement for AI agents in production?
- Microsoft's open-source Agent Governance Toolkit checks each agent action against YAML, OPA Rego or Cedar policy before it runs. Open Policy Agent can sit behind a gateway or MCP proxy you operate. NOFire AI evaluates each agent action against policy and a live production map before execution.
- How do I stop an AI agent from making dangerous changes in production?
- Classify every tool by reversibility and reach, and grant the agent only the tools its task needs. Route each write through a gate outside the agent. Refuse security-sensitive changes, hold irreversible ones for named approvers, and refuse any action whose predicted blast radius exceeds its bound.
- How do I set a blast radius limit or kill switch for AI agents in production?
- Attach a predicted blast radius to each write and set a ceiling per action tier. The gate refuses or holds anything above it. Build the kill switch at the identity provider or the gate, so it halts every agent without the agent's cooperation. Test both in staging on a schedule.
- Should a person approve every agent action?
- Hold every write at the start, then narrow the scope. Once held writes become routine, reviewers start approving without reading. Keep a person on irreversible and wide-reach actions. Let policy decide reversible actions under the bound when the approval history supports it.
What to take from this
Put the approval gate and the kill switch where the agent cannot reach them. If the agent's own process can skip an approval, a prompt injection can skip it too.
Related answers