Your Next User Is an AI Agent

Products built for people need new identity, authority, observability and recovery controls when AI agents become active users.

Imagine an expense agent with access to a product’s API. In a demo, it reads receipts, matches transactions and drafts reports.

Then someone asks what else it can do with the same employee token.

Can it create a vendor? Change bank details? Release a payment at 2 a.m.?

Many SaaS workflows still assume a person is present to read a warning, notice something odd and stop before the final click. Persistent agents weaken that assumption.

OpenAI’s Dots are described as always-on agents with their own cloud computer and browser. OpenAI says they can connect to more than 4,000 apps through plugins.

Meta’s Enterprise Platform includes Muse, Meta Business Agent and Muse API in a stack built for business deployment.

Both announcements point to the same practical problem: software now receives requests from agents that can retain context and act while the human is elsewhere.

For persistent agents making consequential choices, adding an API is not enough.

PMs need to decide which actions an agent may take, how customers can inspect them, and who can stop or reverse them.

The first decision is the operating model. A product can block machine access, expose a fixed integration, support a supervised agent with approval gates, or permit bounded autonomous action.

Choose the least autonomous option that can complete the job.

If the task path, inputs and exception rules can be enumerated, a service account plus deterministic workflow may be all the product needs.

The account model was built for people

A familiar SaaS model assumes that a person signs in, receives a role and moves through a workflow.

The product can show a warning before a destructive action. If something looks wrong, the person pauses.

An agent may work for hours and cross several products to finish one task. It can inherit an employee’s broad permissions for a job that needs only a fraction of them.

Its next action depends on how a model reads the instruction. Nobody has to be reviewing each screen.

The agent can use the product, but the product cannot necessarily identify what acted, who owns it or what authority it was given.

Teams often try the shortest route first: expose an API, reuse an employee token and rely on existing audit logs.

That can be sufficient for fixed, low-risk workflows. It becomes brittle when the agent chooses actions or handles exceptions.

Return to the expense product. One token could let the agent:

  • read a receipt;
  • match it to an existing transaction;
  • draft a reimbursement;
  • create or edit a vendor;
  • change bank details;
  • release a payment.

The product should treat those actions differently. Reading a receipt is low risk; changing bank details or releasing a payment is not.

The permission model should specify the action and its conditions:

  • draft reimbursements, but require approval before submission;
  • use vendors from an approved list, but do not create new ones;
  • prepare payments, but never release them;
  • spend up to a defined amount within a named budget;
  • pause when account details change during the workflow.

These rules have to be enforced where the action occurs. A sentence in the agent prompt is not a permission boundary.

Give the agent its own identity

An agent acting through an employee account creates an attribution problem. The audit trail says the employee changed the record, even when an automated process chose the action.

Give the agent its own identity. Microsoft Entra Agent ID describes distinct identities for agents, registered owners or sponsors and lifecycle controls.

Combined with delegation records and policy logs, this helps answer four questions: what acted, who authorized it, on whose behalf and under which policy?

A separate identity does not solve authorization. It gives the product something precise to authorize.

Products supporting enterprise agents should give administrators an inventory of active agents, their owners, connected systems and recent actions.

Each production agent needs a customer-side business owner, a technical or security owner and a review or expiry date. If nobody owns revocation or incident response, keep the workflow out of production.

GitHub’s enterprise policies for default AI-feature enablement show another part of this admin work: each new capability creates a rollout decision for the customer as well as a feature decision for the vendor.

Define what belongs in your product

Product teams should not rebuild identity, secrets management, policy engines, rate limits or telemetry pipelines. Existing platforms already provide much of that foundation.

The product-specific work sits closer to the workflow: which actions are sensitive, what approval means, which result counts as accepted, how an action is reversed and what customers need to administer it safely.

For the expense product, an identity platform can verify the agent and a policy service can evaluate a rule.

The expense product still has to know that changing a vendor’s bank details is different from reading a receipt. It owns the approval state, the business mutation and the recovery path.

This boundary matters commercially. Generic controls are likely to move into cloud, identity and observability platforms.

A SaaS company has a reason to build its own layer when it owns distribution into the workflow, domain-specific verification, useful outcome data, regulated auditability or responsibility for execution.

If the controls are identical across workflows, integrate with a platform rather than trying to become one.

Operators need a readable record

Human users make mistakes, but they leave clues: a screen path, a ticket, a conversation or a memory of what they intended.

An agent working in the background may leave only a trace.

GitHub’s OpenTelemetry support for Copilot includes model requests, tool use and step-by-step execution traces.

GitHub says prompt and response content is excluded by default. Its telemetry can help teams reconstruct a run without copying every sensitive input into another system.

Meta describes a related architecture for Muse with credential mediation, action-level approvals, policy checks and monitoring layers.

These are vendor designs, not proof that agents are safe. They show which controls are moving into the product surface.

A useful record should connect:

  • the task or instruction reference;
  • the identity and delegated authority;
  • the policy decision and approval;
  • the tools called and business records changed;
  • the accepted, rejected or reversed outcome.

The final link matters. A trace can explain a wrong action without proving that the result was correct.

Support and security teams need to search this record without reading raw model logs.

Product teams also need retention rules, access controls and a budget for storing and inspecting traces. Observability that no owner can afford to review is not a control.

Run the agent readiness test

Before allowing agent access, a PM should be able to answer six questions.

1. Identity

Does the agent have its own identity, owner and lifecycle, or does it borrow a person’s credentials? Can the company disable the agent without locking out the employee?

2. Authority

Which actions are allowed? Which require fresh approval? Are limits attached to amount, data type, destination, time or confidence?

3. State

What context does the product receive, persist, trust and return? Where did it come from, when does it expire, and can stale context authorize an action?

4. Observability

Can an operator reconstruct the task, policy decision, attempted mutation and accepted outcome without exposing sensitive content by default?

5. Economics

What does an accepted outcome cost after model and tool calls, retries, latency, infrastructure, human review, rework, reversals, support and incidents? Compare that total with the human or deterministic workflow, not with token cost alone.

6. Recovery

Can the company pause the agent, revoke access, reverse an action and hand the case to a person without support rebuilding the state by hand?

The first integration can start with gaps, but its boundaries cannot remain implicit.

If the team cannot answer identity and authority, keep the workflow deterministic. If it cannot answer observability and recovery, keep a person at the approval point.

Existing controls are the foundation

The strongest counterargument is straightforward: enterprise identity stacks already provide OAuth scopes, service accounts, role-based access and audit logs. Why invent a new model?

Often, teams should not.

A fixed integration executes a predefined workflow.

An agent may interpret delegated intent, select actions, carry state across tools and adapt when conditions change. The extra work is justified only when those choices are necessary.

If fixed rules can cover the task, use them.

If the agent must make consequential choices across changing conditions, assemble the existing controls around a distinct identity, bounded authority and recoverable actions.

Start with one accepted outcome

“Support AI agents” is too broad to design, price or test.

Start with one bounded action: let an expense agent prepare a reimbursement package from an existing receipt and transaction. A finance operator accepts or rejects it.

The agent cannot create a vendor, alter bank details or release payment.

Name the owners before the pilot starts. The finance operations lead owns the workflow, security or IAM owns access policy, and the product team names an incident owner. That group makes the continue, narrow or stop decision at the end of the pilot.

The buyer may be the finance leader. The daily reviewer is the person whose time the workflow must actually save.

Run the pilot for four weeks across 200 reimbursement packages. Its primary outcome is accepted packages per reviewer hour, compared with the preceding four-week baseline.

Also track reviewer minutes per package, false passes, false blocks, unauthorized-action attempts, rollback or rework rate, completion time, support cases and cost per accepted package.

Set the stop conditions in advance. Fall back to deterministic automation or draft-only assistance if reviewer throughput does not improve, accepted-output cost stays above the current workflow, or a high-impact action bypasses approval.

Stop as well if false passes exceed the workflow’s tolerance or each customer requires custom policy work.

The same evidence tests the business. If every customer needs bespoke connectors and policy rules, the company may be selling integration services under an agent label.

If no team owns policy changes, trace review and incidents, the workflow is not ready for production.

Start with the reimbursement draft. Give the agent permission to prepare it, not release the payment.

Do not widen the scope unless a finance operator can inspect the run, stop it and recover without support rebuilding the trace by hand.