A man working on a laptop in a cozy, modern office space with a focus on technology.

Photo by Matheus Bertelli on Pexels

A coding agent should be called production-ready only when its runtime record shows what it attempted, which tools it used, what it changed, where it stopped and how operators recovered. A polished demo, successful benchmark or green test suite cannot answer those questions.

The useful reporting begins at the failure boundary. When an agent stops during a release, the relevant evidence sits below the final chat message: session events, tool-call arguments, command output, file diffs, permission decisions, model and configuration versions, retry behavior, resource limits and the state left behind.

GitHub has released a production-ready SDK for embedding Copilot’s agent runtime, including planning, tool use, file edits, streaming and multi-turn sessions. That description establishes the scope of the runtime. It does not, by itself, establish how reliably a particular implementation behaves under production pressure.

Start with the complete session record

Ask for the full event stream from the failed run, with sensitive values redacted but event order and timestamps preserved. A screenshot of the last error removes the sequence that often explains it.

The record should show the initial instruction, relevant context supplied to the agent, each plan revision, every tool request, every tool response and the final termination event. For multi-turn sessions, it should also show which state carried across turns and whether older context was summarized, dropped or replaced.

Streaming deserves separate scrutiny. The user-facing output may stop while work continues in the background, or the interface may continue showing progress after the runtime has stalled. Reporters should ask how the product distinguishes generated text, confirmed tool results and status messages inferred by the interface.

Request correlation identifiers too. Can an operator trace one session across the agent runtime, model provider, shell, repository host, deployment service and application logs? If each system assigns an unrelated identifier, reconstructing the failure may depend on approximate timestamps and guesswork.

A credible demonstration should make one failed session traceable from instruction to terminal state.

Inspect what the agent changed

The second evidence set is the mutation record. Ask for the exact repository state before the run, the patch the agent produced and the state after the session ended. “The agent edited three files” is too coarse.

The diff should reveal whether the agent changed generated files, dependency locks, test fixtures, deployment configuration or files outside the requested scope. It should also show whether edits were committed, staged, left untracked or partly reverted.

Then inspect tool permissions. Which commands could the agent execute without approval? Could it write outside the workspace, access environment variables, call network services, modify continuous integration settings or initiate a deployment? A product can constrain file edits while leaving a powerful shell tool broadly available.

Reporters should request the policy configuration used during the run, plus a log of permission checks and operator approvals. The distinction between proposed action and executed action matters. So does the distinction between an approval requested in the interface and one actually enforced by the runtime.

This is the same measurement problem described in The Self-Grading Ad Platform Problem: evidence supplied by the system being evaluated needs an independent check.

Find the exact stopping condition

“Agent stopped” can describe several different events. The model may have declared completion. A tool may have timed out. The runtime may have exhausted a step, token or cost limit. A permission request may have gone unanswered. A network call may have failed. The process may have crashed while the interface displayed the last buffered message.

Ask for the termination code and its documented meaning. Then request the configured limits, actual usage and retry record. If a command timed out, did the runtime kill the child process? If the model response failed, did the system retry with the same context? If a tool returned an ambiguous result, did the agent treat it as success?

The important question is whether the system failed closed. A stopped agent should leave a legible state that an operator can inspect before resuming, reverting or deploying. Automatic retries can hide intermittent faults and repeat side effects, especially when a tool call creates resources, sends messages or changes remote configuration.

Tests also need provenance. Ask which tests ran, who selected them, what commit they ran against and whether the deploy artifact came from that same commit. Green Checks, Dead Application examines why passing indicators can coexist with a failed system.

Require a reproducible failure and recovery

The strongest evidence is a replayable case. Request the smallest safe reproduction, with the prompt, repository fixture, runtime version, model identifier, tool configuration and dependency versions recorded. Where deterministic replay is impossible, the vendor should say which components vary and how repeated runs are compared.

Recovery evidence matters as much as failure evidence. Can an operator resume from the last confirmed state without repeating completed tool calls? Can the system produce a clean rollback? Does it identify uncertain side effects? Who has authority to approve the next action?

Before publishing “production-ready,” ask the vendor to demonstrate one controlled failure and one recovery from it. Preserve the raw trace, compare it with the interface narrative and document any missing fields.

At 4:47 PM, the valuable artifact is not the agent’s apology. It is the record that lets the person on call determine what happened before anyone presses Run again.

Comments

No comments yet.