A programmer working on code with a laptop and monitor setup in an office.

Photo by Jakub Zerdzicki on Pexels

An unapproved agent skill usually appears in CI as a chain of ordinary events: a package or plugin is discovered, instructions are loaded, and a tool or command runs under the job’s existing credentials. Teams miss it for days because those events sit in different log sections and rarely share one clear label.

Start with the build where behavior first changed, then trace backward from the command, network request, or file modification that produced the visible effect. The useful evidence is often present before the skill has a recognizable name.

Reconstruct the invocation from fragments

A skill invocation may leave no line saying “unapproved skill started.” The build log might show a dependency arriving during setup, a configuration directory being scanned, an instruction file being opened, and a shell process starting several minutes later. Each line looks routine on its own.

The reconstruction depends on sequence.

First, find the earliest build with the suspicious behavior. Compare it with the last clean build using full logs, including setup output that the CI interface collapses by default. Look for changes in downloaded packages, cache keys, environment preparation, executable paths, and configuration discovery.

Next, identify the first action the normal build definition cannot explain. That could be an unexpected command, an outbound request to a new host, a write outside the usual build directory, or access to credentials unrelated to the job. Record its timestamp, process, working directory, arguments, and parent process where the CI system exposes them.

Then work backward. A command executed by the agent may have been selected because an instruction file described when to use it. That file may have entered through a plugin installation, a restored cache, a checked-in configuration change, or an image update. The invocation is the whole chain, not merely the final command.

This resembles the debugging problem in The 2 AM Stack Trace That Names a Method That Was Never Real: the visible name may be generated, wrapped, or detached from the component that caused the behavior.

Separate installation, discovery, and execution

These are three different events, and treating them as one creates blind spots.

Installation means the skill’s files became available. Discovery means an agent or compatible client found and read them. Execution means the agent acted on those instructions. A plugin can be installed without being discovered, discovered without being invoked, or invoked through a general-purpose shell tool whose logs never mention the skill.

That distinction matters during incident review. Finding a package in a build image does not prove it ran. Finding its instruction file in an access trace does not prove a command followed. Conversely, failing to find the package name in command output does not clear it.

Build a small evidence table with one row per event:

  • Record the UTC timestamp, build ID, job, and runner identity.
  • Capture the exact file path, process, network destination, or API call.
  • Label each event as installation, discovery, instruction loading, tool selection, or execution.
  • Mark whether the evidence comes from the CI log, audit log, runner telemetry, repository history, or cache metadata.
  • Keep inference separate from observed fact.

That final point prevents a plausible narrative from hardening into a false conclusion. “The agent read this file 14 seconds before starting this process” is evidence. “The file caused the process” remains analysis until another record connects them.

Explain why approval controls failed

Once the sequence is visible, inspect the approval boundary. The failure may have happened before the build began.

A repository review can approve workflow code while missing content restored from a mutable cache. A pinned workflow can still run inside an updated image. An agent may also inherit broad credentials from the job, allowing a newly discovered skill to use permissions granted for another purpose. The same risk appears when an agent uses the right login for the wrong job.

Check who could change every input in the chain: repository files, plugin manifests, package versions, container tags, cache contents, organization-level agent settings, and remote instruction sources. “No one changed the workflow” covers only one of those paths.

Approval should bind a reviewed skill identity to a fixed version or digest. It should also define permitted tools, network destinations, file paths, and credentials. A name alone provides weak control because the contents behind that name can change.

Make the next invocation obvious

Preserve the affected logs before rerunning anything. A rerun can rotate records, refresh dependencies, or replace the cache state needed to explain the original build.

Then add telemetry at the boundaries the investigation exposed. Log the resolved plugin version or digest, the paths of loaded instruction files, the selected skill identifier, and the tool called as a result. Redact secret values while retaining credential names or scopes. Send those records to storage with retention longer than the CI interface’s default window.

Finally, deny unknown skills by default in protected jobs. Run approved skills with the smallest available permissions, pin their contents, and alert when discovery or execution falls outside the allowlist.

The practical test is simple: on the next build, an investigator should be able to connect a reviewed skill digest to an instruction load and then to a specific tool call without assembling the answer from five collapsed log panels.

Sources

No source links were supplied.

Comments

No comments yet.