When should an AI agent ask a person before it acts?
Imagine an agent asking a person to approve every small step: a customer preference update, a draft message, a routine reversible change. After a week of prompts, the person starts clicking through without reading them. That is not meaningful oversight. The useful question is which actions need a person before they happen, which can be reviewed afterward, and which should never be available to the agent at all. A practical starting point is the action’s impact and reversibility.

Agents are moving from answering questions to taking actions: updating records, sending messages, issuing refunds. Each new capability raises the same question for the team running it: which of these actions should wait for a person? The usual answer is to keep a human in the loop.
The phrase sounds reassuring, but it does not tell an operator where to put the human, what they should see, or what they are deciding. A person can be in a workflow and still have no useful control over it: the approval arrives after the action, the preview leaves out the affected account, or the prompt gives them one button and two seconds to understand a complicated request.
The other failure mode is asking for permission constantly. If an agent pauses for every low-impact step, the approvals become routine. People stop treating each one as a decision. Adding a prompt does not automatically add oversight.
NIST’s AI Risk Management Framework treats human oversight as something organizations define and document for the system and its context. It also calls for clear responsibilities across the people who use, manage, and oversee an AI system. That points to a more useful design question than “Do we have an approval step?”: who is responsible for which decision, and what information do they need to make it?
Start with what the action can change
Two properties are useful to assess before deciding whether the agent should pause: the possible impact if the action is wrong, and how easily the action can be reversed. The following is a starting model, not a universal policy:
| Action | If it goes wrong | A reasonable default |
|---|---|---|
| Read a record or draft a response | No external state changes | Let the agent proceed; log the action where appropriate |
| Change a reversible internal field | A person can inspect and restore it | Allow, then review a sample or flag exceptions |
| Send a message, issue a refund, or change access | Another person or system is affected | Show the exact target and change, then ask before execution |
| Delete data, move significant funds, or make a change with hard-to-reverse effects | Harm may be serious or recovery uncertain | Require a named approver or a separate control; some actions may remain off limits |
The labels depend on the organization. A “reversible” change may not be reversible in practice if other systems immediately consume it. A routine message can become consequential if it goes to the wrong recipient or contains sensitive information. Classify the effect in context, not just the tool name.
Make the approval about one specific action
A useful approval should let someone answer a few plain questions: what is the agent about to do, to which record or person, with what exact value or message, and what happens if they approve? If the action is part of a longer workflow, the reviewer should also see what has already happened and what the approval will allow next.
“Allow the agent to continue?” is a weak prompt. “Send this message to the customer attached to case 1842?” gives the reviewer something concrete to check. For a payment or access change, show the amount or permission being added, not a generic description of the workflow.
The approval should also have a clear expiry. If the underlying details change after the person approves, the old approval should not silently authorize the new action. An approval is a decision about the action presented at that moment, not a reusable blank cheque.
Decide what must be blocked by policy
Some actions are too consequential to leave to a model’s judgment about whether it should ask. The agent should not be able to approve its own escalation, change its own access, or bypass a required separation of duties. A policy can block those actions outright or route them to an authorized person outside the agent’s control.
This is different from a confirmation prompt. A confirmation is a step in a workflow. A policy is a boundary on what the workflow can do. If the agent can ignore the boundary whenever a prompt or instruction tells it to, the boundary is not doing much.
Treat the protocol prompt as a mechanism, not a policy
The current MCP specification supports a tool call that needs more information from the user, including a confirmation, to pause and request input through a multi-round-trip interaction. That gives clients and servers a way to build an approval moment into a call. It does not decide which actions require approval, who is allowed to give it, or what the person needs to see. Those remain product and organizational policy decisions.
It is also worth testing what happens when the person declines, does not respond, or changes the underlying request. A safe workflow should fail closed for the gated action, preserve the reason for the decision, and make it clear whether other steps already ran.
Check whether the human step is working
Review a sample of approvals and declines. Can the reviewer explain what they authorized? Do approvals happen before the side effect? Are people declining because the action is wrong, or because the prompt is too vague? Do approvals time out, and what happens next? These questions reveal more than the raw number of prompts.
The right number of approvals is not zero or one for every tool call. It depends on the consequence of each action and on whether a human can meaningfully judge it. Ask when the person’s decision can change what happens. Let safe, reversible work proceed within a clear boundary. Block actions that should not be available. That is a more useful definition of oversight than putting a person somewhere in the diagram.
Sources
- NIST, Artificial Intelligence Risk Management Framework 1.0: Core (human roles, oversight, and lifecycle risk management; accessed 2026-10-05)
- Model Context Protocol, The 2026-07-28 Specification (multi-round-trip requests and user confirmation; accessed 2026-10-05)