An agent’s proposed production security change should be approved only after the on-call engineer verifies the incident, the control’s purpose, the change’s exact scope, and a tested rollback path. At 2:13 a.m., speed matters, but an agent’s confidence is not evidence.
In April 1970, Apollo 13’s crew faced rising carbon dioxide after an oxygen tank failure forced them into the lunar module. The command module had square lithium hydroxide canisters; the lunar module system used round openings. The crew had canisters they could not fit into the system keeping them alive.
On the ground in Houston, engineers had to produce a procedure using materials available aboard the spacecraft. NASA’s Apollo 13 records document the improvised adapter, assembled from items including plastic bags, cardboard, a hose and tape. The astronauts followed the transmitted instructions, and the device worked.
That response succeeded because Mission Control treated every constraint as real. A plausible sketch was insufficient. The materials had to exist onboard, the instructions had to survive transmission, and the crew had to be able to perform each step under pressure.
The same discipline belongs in a production approval at 2:13 a.m., even when the consequences are smaller.
Turn the proposal into a falsifiable claim
The agent proposes changing a production security control. Perhaps it recommends disabling a policy, widening an allowlist, changing a firewall rule or granting a service broader access. The on-call engineer’s first task is translation.
“Change this control to restore service” is a conclusion. A reviewable proposal identifies the affected resource, the current control, the proposed state, the error or event it should stop, and the evidence connecting that control to the failure.
Start with the observed production symptom. Which request failed? Which deployment, configuration update or traffic change preceded it? Do runtime telemetry, authentication logs or policy decisions show the security control blocking a legitimate action? Could the same evidence also fit an expired credential, a bad release or a dependency failure?
This distinction matters for tools that combine agents, integrations and generative assistance. Sysdig Secure AI, for example, uses runtime telemetry to investigate and remediate cloud risks. Runtime evidence can make a recommendation more relevant, but the engineer still needs to inspect what the evidence demonstrates and what remains inferred.
Ask the agent to expose its chain of evidence in operational terms:
- Show the events that identify the control as the cause.
- State which resources and identities the change affects.
- Separate observed facts from inferred causes.
- Name the permissions or protections removed.
- Describe the expected signal if the change succeeds.
If the proposal cannot be tested against an observable result, it is not ready for approval.
Bound the change before touching production
Once the control appears causal, the next question is scope. The safest useful change is the smallest one that can restore the required behavior while preserving the rest of the boundary.
A request to disable a control globally should trigger resistance. Can the exception apply to one workload, identity, route or short-lived session? Can traffic remain restricted by source, destination or action? Can the engineer add a temporary exception with explicit ownership instead of rewriting the baseline policy?
This is where automation can create false comfort. A syntactically valid change may still grant more access than the incident requires. The agent should produce the exact diff, including inherited effects, rather than summarize it as “low risk.”
The engineer also needs to check dependencies outside the immediate alert. Does the control satisfy a contractual requirement? Does another detection assume it remains enabled? Will the proposed exception persist after the incident? A production fix that quietly becomes permanent creates a second incident with a slower clock.
This evidence-first approach also applies before deployment. AI Coding Agent Risk Thresholds: How Priya Narrowed a Production Rollout shows why permissions and production reach should expand only when the evidence supports them.
Require a rollback that can actually run
“Rollback available” is too vague for an overnight approval. The engineer needs the exact command or configuration change, the person authorized to execute it, and the signal that will trigger reversal.
A useful rollback plan answers four questions. What state will be restored? How will the engineer confirm restoration? What happens to requests or data created while the exception is active? Could rollback fail because the agent changed another dependency?
Test the reversal outside production when the environment permits it. If that is impossible, inspect the previous configuration, confirm it remains valid, and preserve a known-good copy in the approved operational system. Do not rely on the agent remembering its earlier output.
Then define the observation window. Watch the original failure signal, the security events the control was meant to catch, and any new access enabled by the change. Restored availability alone does not prove the fix is safe.
This is the same approval gap examined in Maya’s missing audit trail. One weak answer could end the trial.: a system’s answer matters less when nobody can reconstruct what it did and why.
Make the temporary decision expire
If the evidence supports the change, record the approving engineer, the exact diff, the supporting telemetry, the rollback procedure and the expiration condition. Assign a daylight owner before closing the incident.
The approval should expire automatically where possible. Otherwise, create a tracked removal task with a named owner and a clear trigger. “Review later” tends to become undocumented production policy.
Apollo 13’s improvised adapter was designed around verified constraints: the available materials, the incompatible fittings and a procedure the crew could execute. At 2:13 a.m., the on-call engineer needs the same shape of reasoning. Verify the failure, constrain the intervention, prove the reversal and leave enough evidence for the next person to challenge the decision.
Comments
No comments yet.