a group of people sitting around a table with laptops

Photo by Lyubomyr Reverchuk on Unsplash

An AI agent should not be deployed until the team has defined its authority, escalation path, and the actions it may never take. A Friday deadline does not turn an undefined workflow into a safe one.

The pressure usually arrives as a simple instruction: put an agent into the workflow this week. The request sounds concrete until someone maps the work itself. Does the agent draft a customer reply, send it, issue a refund, change a record, approve an exception, or notify a person who can decide? Those are different jobs with different consequences.

Recent reporting on OpenAI’s move beyond coding points to a broader gap: employee use of agentic software is growing faster than customer use. That gap matters. Internal experiments can absorb uncertainty in ways customer-facing systems cannot. A bad internal draft can be corrected. A wrong account change, payment decision, or customer promise can be much harder to unwind.

Start with the decision, not the agent

The first workflow review should identify the decision at the center of the process. “Handle support tickets” is too broad. “Classify incoming password-reset requests and draft a reply” is a bounded task. “Reset customer accounts” is an authority decision.

That distinction is where rushed deployments fail. Teams describe a workflow as if it were one action, then discover it contains a chain of judgment calls:

  • Is the request legitimate?
  • Does the account meet the policy?
  • Is there an exception?
  • Does the action create a financial, legal, security, or customer-relationship commitment?
  • Can the result be reversed?

An agent can help at every point in that chain. It does not need permission to complete every point.

Write the workflow in verbs. Retrieve. Summarize. Compare. Draft. Recommend. Route. Execute. Each verb should have an owner. If a team cannot name the person or system accountable for an action, the agent should not take it.

Define a boundary that survives a messy case

A useful authority boundary is specific enough to hold when the obvious case becomes an awkward one.

For example, an agent may retrieve an order record, compare it with a published return policy, draft a response, and route the case to a support specialist. It may not issue a refund, alter an address, disclose account information, or make an exception without a defined approval step.

The test is not whether the happy path works. The test is what happens when a customer’s request is incomplete, contradictory, urgent, or outside policy. The agent needs a clear place to stop.

That stop condition should be written in plain language. “Escalate when confidence is low” sounds responsible but leaves the hardest question unanswered: low by whose definition? A better rule names the trigger: missing account verification, a request involving payment, conflicting customer records, a policy exception, or any instruction that would change access, pricing, or contractual terms.

The boundary should also cover tools. An agent with read access can investigate. An agent with write access can change the world. Treating those as a minor configuration detail is how a drafting tool becomes an action system without a deliberate decision. Friday, 4:47 PM: The Agent Can Approve Payments explores the same problem from the point where approval authority enters the workflow.

Make human review a designed step

“Human in the loop” is often used as a comfort phrase. It only protects anyone when the human knows what they are reviewing, has enough context to challenge the recommendation, and can act before the consequence occurs.

Put the review step next to the consequential action. If the agent recommends an account change, show the evidence it used, the proposed change, the relevant policy, and the reason it chose that route. A reviewer should not need to reconstruct the case from a chat transcript.

Then decide what the human can do:

  • Approve the proposed action.
  • Reject it and select a reason.
  • Edit the action before it is carried out.
  • Send the case to a different queue.
  • Stop the workflow altogether.

Those choices create feedback that can improve the process later. They also leave an audit trail when a decision needs to be examined. A review button that appears after an email has gone out or a record has changed is documentation, not oversight.

Treat Friday as the start of a controlled release

A deadline can still be useful. It can force a team to choose a narrow first use case instead of trying to automate an entire department. The release by Friday might be an agent that gathers information, prepares a recommendation, and routes edge cases. That is a real deployment with a limited blast radius.

Before it goes live, run representative cases through the workflow: ordinary requests, incomplete requests, requests that appear valid but violate policy, and requests that require a person to exercise judgment. Record where the agent stops, what it shows the reviewer, and whether the system prevents forbidden actions.

The most important launch artifact may be a short authority map: what the agent can read, what it can write, what it can recommend, what requires approval, and what it must never do. Keep it beside the workflow, not in a planning document nobody opens after launch.

On Friday afternoon, the useful question is not whether the agent is live. It is whether every action it can take has a named owner, a visible limit, and a safe route for the cases it cannot resolve.

Comments

No comments yet.