Diverse business team in a meeting analyzing client testimonial on laptop screen.

Photo by Mikael Blomkvist on Pexels

At 8:07 on Monday morning, the agent can be cleared to handle support and billing work while responsibility for its first costly mistake remains undefined. That blank matters more than the launch approval because automation can execute immediately, while accountability often arrives only after damage appears.

OpenAI has introduced Presence as an enterprise product for deploying AI agents across customer-facing and internal workflows. The announcement puts a familiar operating problem in sharper focus: companies can approve what an agent may do without deciding who owns the consequences when it does the wrong thing.

Approval can hide an ownership gap

A launch meeting usually produces decisions that are easy to record. The agent may answer support requests. It may classify billing disputes. It may issue credits below a set amount or send larger cases to a person.

Those permissions describe the system’s operating range. They do not establish who carries the decision when the agent acts within that range and still creates a costly outcome.

Consider the unresolved questions behind a billing workflow. Who reviews an incorrect refund? Who decides whether affected customers must be contacted? Who can pause the agent outside normal working hours? Which team absorbs the loss? Who explains the incident to finance, support leadership and customers?

If the launch document lists engineering, support, finance and legal as stakeholders, ownership may still be blank. A stakeholder receives updates. An owner has the authority and obligation to act.

This distinction becomes more important when an agent crosses departmental boundaries. A support response can change a customer relationship. A billing action can affect revenue and accounting. One automated decision may create work for three teams, each assuming another team controls the system.

Guardrails do not replace a named decision-maker

Teams often treat controls as a substitute for accountability. They set spending limits, require approval for selected actions, restrict data access and log agent activity. Those controls are useful, but none can decide what the company should do after an unexpected failure.

A threshold also creates a misleading sense of precision. An agent authorized to issue a small credit may repeat the action across numerous accounts. A response that appears harmless in isolation may become expensive at volume. A human approval step can fail too, especially when reviewers see a queue of similar cases and begin approving them by habit.

The key question is not simply, “Can the agent perform this action?” The launch team must decide who owns each failure class and what authority that person has.

That means assigning an accountable owner for customer harm, incorrect financial actions, privacy incidents and operational disruption. One person may own several categories. Different categories may belong to different people. Either approach can work if the names, escalation paths and stop conditions are written down before launch.

This is the operational lesson behind the ten seconds after you press go: deployment changes the job from evaluating a system to controlling a live one. The first abnormal result should trigger a prepared response, not a debate over whose queue receives it.

The first costly failure needs a runbook

A useful launch record should make the response executable. “Support and engineering will investigate” leaves too much room for delay. A stronger record identifies the first responder, the accountable decision-maker, the evidence to preserve and the conditions that require a pause.

The practical test is simple. Pick one plausible failure and walk it through from detection to resolution.

Suppose the agent makes an incorrect billing decision. The team should be able to identify who receives the alert, who checks the transaction history, who can suspend billing actions, who decides whether to reverse the decision and who approves customer communication. The record should also state how the incident is handed over if the owner is unavailable.

Logs matter here, but only if they show enough to reconstruct the action. The team needs the agent’s input, the relevant context, the action taken, any tool calls, the policy or rule applied and the resulting system change. A timestamped output without its decision context may confirm that something happened while leaving the cause unclear.

The same discipline applies before launch. The unapproved skill in the build log examines a related control problem: capabilities can enter an operating environment before their approval status is clear. An inventory of tools, permissions and owners gives reviewers something concrete to approve.

What to record before the agent goes live

The approval should name one accountable owner for every action that can move money, alter customer records, disclose information or interrupt service. It should also identify a backup owner. Shared responsibility becomes fragile when an incident begins outside office hours.

Record the stop conditions in plain language. Examples might include an unexplained rise in refunds, repeated action on the same account, use of an unauthorized tool or a mismatch between the agent’s stated decision and the system change it made. The exact conditions depend on the deployment, but they must be observable.

Then define the pause mechanism. The person accountable for an incident needs the practical ability to stop the affected action without waiting for a new meeting. If only a separate infrastructure team can disable the agent, the escalation route and expected response must be explicit.

Finally, schedule the first review around evidence, not reassurance. Examine actual actions, escalations, reversals and near misses. Ask which decisions lacked a clear owner. Update the runbook while the details are still available.

Before the Monday meeting ends, every high-impact action should have a name beside it, a backup contact and a tested way to stop it. The blank ownership field is itself a launch blocker.

Comments

No comments yet.