Software developer analyzing code on a tablet in a modern office workspace.

Photo by Jakub Zerdzicki on Pexels

A pull request from a coding agent can show that a change is plausible, but it cannot prove why an outage happened. Diagnosis requires evidence that connects the observed failure, the proposed mechanism and the effect of the change.

The distinction matters as coding agents move closer to incident response. A public preview now lets Copilot Business and Enterprise users start shared coding sessions from Slack, investigate failures, implement changes and open pull requests within existing GitHub permissions. That shortens the path from alert to proposed code. It does not settle the causal question.

A working patch can support several explanations

Suppose an agent examines an error, finds a suspicious code path and opens a pull request. The change may pass tests. Reviewers may see a credible explanation in the pull request description. The service may even recover after deployment.

Each observation is useful. Together, they can still leave the outage’s cause unproven.

A patch may remove the original trigger. It may also suppress a symptom, change timing, reduce load or bypass the failing path. Recovery could coincide with a cache expiry, dependency restoration, traffic drop or configuration change. When several conditions move at once, the successful deployment offers weaker causal evidence than it first appears.

Speed makes this easy to miss. A complete-looking pull request arrives as a visible object: changed lines, test results and a proposed explanation. The diagnosis often remains scattered across logs, traces, deployment records and conversations. People naturally give more weight to the finished artifact.

That creates a dangerous substitution. The team begins reviewing whether the code change looks correct when it still needs to establish whether the stated failure mechanism matches the evidence.

Separate observation, inference and verification

A useful incident record keeps three layers distinct.

Observed facts include error rates, affected requests, deployment times, trace output and dependency responses. These should preserve timestamps, scope and the queries used to retrieve them.

Inference connects those facts into a possible mechanism. For example, a team might infer that a recent change caused requests with a particular input to enter a failing branch. That explanation should remain a hypothesis until the evidence rules out credible alternatives.

Verification tests the hypothesis. The strongest checks reproduce the failure under the suspected conditions, remove or reverse one variable, and observe the predicted result. A regression test can help, provided it recreates the production failure rather than a simplified behavior that merely resembles it.

The pull request belongs primarily in the verification layer. Its description can cite observations and explain the inference, but the existence of the change proves neither. Reviewers should be able to trace each causal statement back to evidence outside the agent’s narrative.

This is the same control problem behind Friday Green, Friday Broken: a passing check answers only the question encoded in that check. Confidence grows when the test matches the conditions that produced the failure.

Review the diagnosis before merging the remedy

Teams can make agent-generated incident fixes safer without slowing every change to a crawl. The first step is to split the review into two decisions:

  • Does the evidence support the proposed cause?
  • Does the change safely address that cause?

Those questions need separate answers. A reviewer may approve the code while marking the diagnosis uncertain. Another may accept the likely cause but reject the patch because it adds risk elsewhere.

The pull request should identify the failing behavior, the evidence supporting the suspected mechanism, alternative explanations considered and the test that would disprove the claim. If production reproduction would be unsafe, the record should say so plainly and describe the next-best evidence.

Permission boundaries also deserve careful treatment. The public preview operates within existing GitHub permissions, which limits what the coding session can do. Permissions govern authority. They do not establish that the agent inspected every relevant system or received complete telemetry. A technically valid investigation may still have a narrow evidence base.

That gap resembles the issue examined in The Safeguard Nobody Can Demonstrate. A control carries more weight when someone can show what it covered, how it ran and what result it produced.

Preserve uncertainty after service recovery

Incident pressure rewards closure. Once the service recovers and a pull request looks convincing, leaving the root cause open can feel like unfinished work. Sometimes “probable cause” is the most accurate conclusion available.

That label should trigger follow-up rather than embarrassment. Preserve the relevant telemetry, document competing explanations and assign a verification task with an owner. If later evidence contradicts the original theory, update the incident record and the regression test.

Coding agents can reduce the time required to inspect a repository, propose a change and prepare review material. That is valuable during an outage. The gain becomes risky when proposal speed is treated as diagnostic certainty.

Before merging the next fast fix, ask for one concrete item: the observation that would have looked different if the proposed cause were false. If nobody can name it, the team has a plausible patch and an open investigation.

Comments

No comments yet.