Enterprise AI

Agentic SRE and the Autonomy It Earns

Magnus Hedemark 8 min read
A South Asian reliability engineer studies incident traces at her desk as a constellation-like agent offers evidence beside her.
The person on call stays responsible for the decision. Better evidence can give engineers more room to improve the service.

Give an agent more authority one capability at a time, and let evidence, not enthusiasm, set the limits.

Say a checkout service starts throwing errors right after a deploy. An agent tracks down the commit, charts the error spike, pulls the failing traces, and suggests pulling the new build out of the pool. When the on-call engineer opens a laptop, the diagnostic legwork is already done and waiting for review.

That checkout scenario is strictly hypothetical. But it tees up the real practical question: after an agent has helped with incidents like that, what must it prove before you let it pull the canary itself? And who gets to make that call?

Google’s introduction to site reliability engineering (SRE) describes SRE as software engineers designing operations teams and writing software to replace repetitive manual work. Reliability objectives give teams a way to balance change against the service reliability users need.

Here, agentic SRE means putting artificial intelligence (AI) agents to work gathering operational evidence, picking investigative paths, running tools, and adapting as new signals come back. You do not grant authority because an agent sounds smart in a prompt. You delegate specific operational chores only when the surrounding engineers, tools, and guardrails can safely support that autonomy.

Start with the work you want to improve

During an outage, high-value work usually means finding the right service owner, piecing together a deployment timeline, verifying dependencies, or ruling out a tempting red herring. An agent can assemble that picture long before anyone lets it touch production infrastructure.

An engraved left-to-right process diagram. Four source boxes at the left are labeled “Deployments,” “Traces,” “Metrics,” and “Dependencies.” Arrows converge on “Evidence packet,” which leads to a responder labeled “Human judgment,” then to “Improve the service.” The image shows the proposed sequence: operational signals become an evidence packet for human judgment before a service improvement.
An evidence packet can make investigation more useful before an agent receives permission to act.

The aim is to spare the responder some of the scramble for context while keeping their judgment in the loop. When routine diagnostics assemble themselves, engineers can turn recurring failure patterns into maintained runbooks, build better recovery tools, and feed incident findings back into core architecture. The time this saves can help reduce repeat incidents, but that is an outcome you have to measure in your own systems, not an automatic guarantee.

Meta notes that its AI Regression Solver prepares performance fixes for review. Their setup pulls context around an identified regression, draws on codebase-specific expertise, and drafts a proposed pull request for the original author. The write-up confirms a human review step in the pipeline, not an independently audited boost in site availability. Still, it highlights practical ground between typing every shell command by hand and blindly handing an agent keys to production.

Five levels for a particular capability

This ladder adapts the Groktopus earned autonomy concept to operational reliability. It is a proposed way to reason about delegated work, not an industry standard, certification, or vendor score. Apply it to one capability in one environment. Investigating a test cluster, adjusting production traffic, changing a database, and declaring an incident closed are different decisions. A level describes the evidence for that specific capability and environment; it does not transfer to the next capability. No Level 4 or Level 5 label is universal proof, and Level 5 is not a destination every system should seek.

A five-step ascending diagram labeled in order: “1 Assist,” “2 Human review,” “3 Bounded action,” “4 Escalate exceptions,” and “5 Sustained evidence.” A dotted arrow rises across the steps, and a human observer stands beside the fifth. The footer reads “One capability • one environment.” This is a proposed, capability-scoped ladder, not a certification or automatic promotion.
The proposed ladder applies to one capability and environment, not an entire SRE function.

Keep the checkout scenario hypothetical as it moves through the levels.

Level 1: Assisted investigation

The agent gathers deploy records, distributed traces, and service metrics. It flags the newly shipped checkout revision as a suspect, links every piece of evidence, and explicitly notes what it could not verify. A human engineer looks at the findings, weighs what they mean, and chooses what to do.

Track whether this actually helps under pressure. A fluent explanation that points a tired responder down the wrong rabbit hole makes an outage harder, not easier.

Level 2: Delegated with human review

The agent drafts a concrete recovery plan: the target deployment, the precise routing change, sanity checks on remaining capacity, and the exact telemetry that will verify success. A human reviews the plan before any command runs.

That review has to carry enough technical depth to catch subtle errors. If traffic shifts or system state drifts before someone clicks approve, the proposal has to be re-evaluated. Greenlighting a narrow traffic shift never grants blanket permission for whatever else the agent wants to poke at next.

Level 3: Autonomous within verified boundaries

Once evaluated and explicitly authorized, the agent can drain an unhealthy canary from a strictly bounded traffic slice without waiting on a click. In this proposed level, execution controls outside the model’s reasoning loop enforce the approved targets, actions, and resource limits.

Verification checks must confirm real customer requests succeed and downstream capacity holds steady. Test coverage needs to account for partial executions, concurrent human adjustments, and botched recoveries. If mandatory health signals cut out, the automated path stops cold and hands control back.

Level 4: Mostly autonomous with escalation

The agent handles an expanded set of pre-approved checkout failure modes, selecting among tested recovery recipes. It detects when live symptoms diverge from those patterns. Spiking database connection errors or replication lag, for example, trigger an immediate escalation with whatever diagnostic context was gathered so far.

The team has to verify that the handoff actually reaches a person who can respond in time. Paging a dead Slack channel or dumping tickets into an unread backlog does not count as a human handoff.

Level 5: Sustained autonomous operation

Treat a capability as sustained only when it demonstrates consistent, safe behavior across real operational variance: traffic spikes, routine releases, shifting dependencies, and degraded conditions. Engineers keep sampling outcomes, updating evaluation suites, and tracking what it costs to keep the setup dependable.

Authority has to pull back whenever operational evidence slips. Shipping a new model checkpoint or tool version requires fresh validation. A calm week with no incidents tells you almost nothing about how the setup handles failures that never occurred. Sustained autonomy never removes the need for accountable human owners.

People decide what the evidence permits

When agents take on more direct execution, human engineering work shifts toward designing, testing, and hardening the conditions under which they run.

A responsibility map flows left to right: “SRE evidence” points to “Human owners decide,” with “thresholds,” “grant,” and “revoke” beneath. “Governance records the why” supports the human decision. The flow continues through “Approved limits” to “SRE enforcement,” where an agent is visibly contained by a boundary. The diagram separates human authority, governance records, operational evidence, and enforcement.
People set authority; governance records the decision; SRE supplies evidence and enforces the approved limits.

Service owners, incident commanders, security teams, and governance leads must hash out acceptable consequences, promotion thresholds, demotion triggers, and exact authority to grant or revoke operational access. Accountable people make those calls. An artificial intelligence (AI) governance skill can help teams structure and record them; it does not grant authority. SRE supplies operational evidence, tests recovery behavior, and implements approved limits through controls outside the agent’s reasoning loop.

A written prompt or skill can describe an operational boundary. The runtime environment has to enforce it mechanically. That matches the broader architectural reality explored in QA Is Becoming the Control System: assurance must govern the operational harness surrounding the agent, not just inspect the final artifact it outputs.

That takes disciplined engineering. Teams have to refine service level objectives, build realistic incident datasets, tighten verification checks, drill escalations, and write recovery scripts that survive an interrupted agent. Engineers also need practice taking over when automated paths fail.

The engineering payoff comes from making future failures less likely and triage decisions faster. Shaving off a few confirmation clicks is the least interesting part of the project.

Earn autonomy with evidence that survives scrutiny

Start by establishing an honest baseline for the exact task you plan to delegate. How much responder time does it eat up today? How often does the initial triage guess turn out wrong? How fast do end users actually recover, and how often does the issue flare back up?

An evaluation scorecard titled “Measure the whole run.” Four columns read “Coverage” with “eligible incidents”; “Actions” with “attempts,” “abstentions,” and “escalations”; “Outcomes” with “verified recovery” and “harmful effects”; and “Operating cost” with “runtime,” “review effort,” and “maintenance.” The footer reads “State the denominator for every rate.” The figure contains no performance values.
Meaningful evaluation tracks coverage, actions, outcomes, and ongoing human costs with explicit denominators.

Replaying historical incidents is valuable, but only if the replay is restricted to the messy signals visible when the alert first fired. As incident.io points out in their account of building Investigations, later fixes and postmortems can leak the solution into a test. Let that information leak through, and a brittle system looks deceptively brilliant in benchmarks.

For the checkout scenario, build tests that include bad rollouts, third-party API outages, broken metrics pipelines, misleading correlation spikes, and completely healthy systems that need to be left alone. Move from replay to tightly controlled dry runs and narrowly scoped production trials. Make sure to test scenarios where a human engineer or another automation pipeline makes changes at the exact same moment.

Log eligible incidents, attempted mitigations, deliberate abstentions, escalations, confirmed recoveries, and negative side effects. State the explicit denominator for every rate you quote, and keep diagnostic accuracy separate from remediation success. Factor failed and abandoned runs into total runtime and infrastructure spend, alongside human review time and test harness maintenance. Faster triage reports are helpful, but what matters to customers is whether checkout works.

Pin down the exact evaluated configuration: model and version, tools, permissions, instructions, and recovery checks. Establish clear, pre-agreed criteria for when authority will be audited or revoked. An agent declaring itself successful should never close the loop. We do not propose a fixed promotion count or quiet period; the evidence should fit the consequence and scope of the work.

Useful deployments already have different boundaries

These three real-world cases offer distinct snapshots of where organizations draw boundaries today: a company account, product documentation, and a retrospective. They illustrate different operating choices, not a head-to-head safety ranking.

Google reports that AI Operator requires human approval for critical operations while autonomously mitigating some minor incidents. Google’s account also describes Actus, an execution layer that checks proposed actions and can route elevated-risk requests back for approval. The report does not provide the exposure, severity, or harm denominators needed to transfer its safety judgment to another service.

Cloudways documentation describes a customer confirmation step before SmartFix executes a change. It does not provide effectiveness or harm rates, and it notes that dashboard users cannot undo a fix themselves.

In an August 2026 Microsoft retrospective, the authors recount an agent deallocating a virtual machine (VM) after the logging service it relied on became unavailable partway through required safety checks. The agent relied on a remembered safe pattern instead of stopping. The account gives no failure frequency or customer-impact estimate, and it should not be generalized to every Azure SRE Agent path.

A three-column source comparison. “Google report” describes “Human review: critical actions” and “Autonomous: some minor mitigations.” “Cloudways docs” says “Customer confirms before execution.” “Microsoft retrospective” says “VM deallocated when a safety check became unavailable.” The footer reads “Different boundaries • not a safety ranking.” The columns distinguish a company report, product documentation, and a retrospective, not comparable safety results.
Company accounts describe different workflow boundaries, not comparable safety results.

These sources describe very different operational boundaries backed by completely different kinds of evidence. They cannot be mashed into a comparative leaderboard, and they offer no proof that any system is ready for blanket, unattended production management. Every engineering team has to build its own evidence base for each specific workflow it chooses to delegate.

Begin with one useful capability

Pick a specific, recurring task that consumes engineering time and has an identifiable service owner. Define the desired outcome, lock down the permitted blast radius, and establish the exact evidence required to advance. Begin with assisted diagnostics, study every correction your engineers make, and harden the operational harness before handing over more control.

A five-stage loop reads from left to right: “Recurring work,” “Name the capability,” “Set owner and scope,” “Gather local evidence,” and “Expand or hold authority.” A human decision-maker appears at the final stage. A return arrow leads to “Improve the service,” then back to the recurring work. No other text appears in the illustration.
Let local evidence decide whether a capability earns wider authority, should hold, or needs to step back.

Some capabilities may stay at human review indefinitely. Others may earn sustained autonomy under defined conditions. Either path can improve the service and free engineers to work on recurring problems.

In the hypothetical checkout incident, real success looks much simpler than an abstract autonomy metric: customers can complete purchases, responders get useful diagnostic support, and engineers have more time to address the conditions that caused the page. Don't evaluate your agentic investments simply by how many approval clicks you managed to eliminate.