← All stories

The Friday a security lead is asked to approve an agentic coding tool before the vendor has published enough evaluation detail: a practical checklist for separating demonstrated safeguards from reassurance.

Three professional women collaborating in a business meeting with documents and laptops.

Photo by Vlada Karpovich on Pexels

Do not approve broad access when the vendor cannot show how its safeguards were evaluated. Offer a narrow, reversible trial only if your team can independently enforce permissions, observe every consequential action, and stop the tool without relying on the vendor.

Define the approval boundary before reviewing claims

Write down exactly what the tool would be allowed to reach. “Use the coding agent” is too vague for a security decision.

Specify the repositories, branches, developer accounts, credentials, networks, package registries, CI systems, issue trackers and production resources in scope. Record whether the agent can edit files, run shell commands, install dependencies, open pull requests, merge code or trigger deployments.

Then classify each action by consequence:

  • Read-only actions expose information but do not change systems.
  • Reversible actions create changes that your team can inspect and undo.
  • Irreversible or high-impact actions can disclose secrets, alter production, delete data or weaken security controls.

This boundary matters because a useful evaluation for a read-only code assistant may say little about an agent that can execute commands with developer credentials.

Ask for evaluation evidence, not a list of controls

A vendor may say the product uses sandboxing, permission prompts, policy checks and human approval. Those are design claims. Evaluation evidence shows what happened when testers tried to defeat those controls.

Request enough detail to understand:

  • Which product version and configuration were tested.
  • Which tools, permissions and network connections were enabled.
  • What attack classes the evaluation covered.
  • How many scenarios were run and how they were selected.
  • What counted as a pass, failure or partial failure.
  • Whether testing included repeated attempts and multi-step attacks.
  • Which failures remain unresolved.
  • Whether independent assessors reproduced any results.

Aggregate language such as “strong performance on internal safety tests” provides little basis for approval. A useful report connects a defined threat to a test procedure, observed result and known limitation.

OpenAI slowing Astra model development over security concerns offers a relevant principle: delaying capability can be a rational response when safety evidence has not caught up. Your approval process should preserve that option.

Separate prompts from enforceable boundaries

Human confirmation screens can reduce accidental actions. They offer weaker protection when the agent controls how an action is described, splits one risky operation into several harmless-looking steps or overwhelms the user with repeated requests.

Check where each safeguard operates. A restriction enforced inside the model can fail when the model behaves unexpectedly. A restriction enforced by an external policy layer, operating system permission or isolated execution environment can still hold after the model makes a bad decision.

For each important control, ask your team to complete this sentence:

> Even if the agent attempts the prohibited action, it cannot succeed because ______.

Accept answers tied to independently configured controls. “The model was instructed not to” is reassurance.

Test the failure paths that match your environment

Vendor evaluations rarely reproduce your repository structure, secrets handling or deployment permissions. Run a local assessment before granting meaningful access.

Use disposable infrastructure and synthetic credentials. Include repositories that resemble your real projects without containing customer data or reusable secrets. Test whether the agent:

  • Reads files outside its assigned workspace.
  • follows malicious instructions embedded in source files, issues or documentation.
  • sends code or secrets through allowed network channels.
  • installs an unexpected package or executes its setup scripts.
  • changes security checks, CI configuration or dependency pins.
  • disguises a consequential command behind a vague approval message.
  • continues after a user denies permission.
  • leaves enough evidence to reconstruct its actions.

One successful refusal proves little. Repeat important tests with altered wording, file placement and task order. Agent behavior can vary between runs.

For more on choosing permissions that match observed risk, see AI Coding Agent Risk Thresholds: How Priya Narrowed a Production Rollout.

Verify that logs support investigation and rollback

A useful audit record should show the user request, model-visible context, tool calls, command arguments, files changed, network destinations, approval decisions and resulting outputs. Timestamps and actor identities should survive export into your own logging system.

Confirm how long records are retained, who can alter them and what disappears when a session ends. Redacted logs can protect sensitive data, but excessive redaction may make incident review impossible.

Rollback also needs testing. Reverting a Git commit does not undo a published package, disclosed credential, database migration or external API call. Map each allowed action to a recovery procedure and an owner.

Choose a decision that preserves optionality

When evidence is incomplete, avoid forcing the choice into “approve” or “reject.” Use a tiered decision:

  • Approve read-only use in selected repositories with no secrets and no outbound network access.
  • Permit code changes only on isolated branches, followed by mandatory human review and existing CI checks.
  • Withhold shell execution, credential access, merges and deployments until the relevant evaluations exist.
  • Set an expiry date so temporary access does not become permanent by neglect.

Document the missing evidence as conditions for expansion. That converts vendor assurances into testable gates.

The concrete next step is a 30-minute review with engineering, security and the proposed users. Produce one page containing the allowed resources, prohibited actions, external enforcement points, required logs, trial end date and named approver. If any high-impact action lacks an independent boundary or recoverable failure path, remove it from the trial.

Comments

No comments yet.