Guides15 min read
Human-in-the-Loop AI Agent Approval: What Needs Your Yes
Which agent actions should wait for your approval, how to design the approval step, and how to approve from your phone without writing framework code.
By The Unorderly team
Quick Answer: Human-in-the-loop AI agent approval means your agent can research, draft and prepare on its own, but anything that spends money, speaks for you, deletes data or goes public waits for a person to say yes. Gate a short list of high-stakes actions, show the exact action being approved, and make sure nothing executes until the decision is recorded.
Most people who run agents seriously land here the same way. The agent does something slightly too confident, like emailing the wrong list or tidying up a folder that was not meant to be tidied, and suddenly "let it run" feels less charming. You do not need to become an engineer to fix this. You need a clear rule about which actions need your yes, and a place to give it.
This guide is written for founders, operators and marketers who use agents like Claude, Grok, Cursor or OpenClaw day to day. It stays technically accurate, so if you do write code, the framework notes near the end map it to the patterns you already know.
What is human-in-the-loop AI agent approval?
A human-in-the-loop approval is a checkpoint between your agent deciding to do something and that thing actually happening. The agent proposes, you approve or reject, and only then does the action run. Everything before that checkpoint (reading, searching, drafting, planning) can stay fully automatic.
The idea is not new, and the people who write the security guidance take it seriously. The OWASP Top 10 for LLM Applications lists "Excessive Agency" as a core risk, and one of its recommended mitigations is to "utilise human-in-the-loop control to require a human to approve high-impact actions before they are taken." OWASP breaks excessive agency into three root causes: too much functionality, too many permissions, and too much autonomy. An approval step is the fix for the third.
The agent frameworks have built this in too. The OpenAI Agents SDK describes a run that "records an approval interruption instead of executing the tool," then waits for your application to approve or reject it. LangChain's human-in-the-loop middleware lets developers choose, per tool, whether a call runs automatically or pauses for a person. Microsoft's Agent Framework marks individual functions as needing approval so that the agent returns a request instead of a result.
If you use Claude Code, you have already met a version of this. Its permission modes decide which actions Claude can take without asking; in Manual mode it "stops and asks you before most actions that edit files, run shell commands, or reach the network." Same principle, applied to a coding agent.
Why this matters for people running agents
Agents act on what they read, and what they read can be wrong
OWASP names two common triggers for harmful agent actions: the model getting things wrong (hallucination) and prompt injection, where text the agent reads (a web page, an email, a document) quietly tells it to do something else. You cannot fully prevent either. What you can do is make sure the damaging step cannot happen on the agent's word alone.
OWASP's own example is useful here. An agent that drafts social media posts should have "a user approval routine within the extension that implements the 'post' operation." Drafting is fine. Posting waits for you.
The summary is not the action
A good point from Jake's write-up on human-in-the-loop agents is that agents describe their plans in prose, and the prose can drift from what actually runs. His example: the agent says it will email the three customers who churned, then calls the email tool with thirty recipients. If you approve the sentence rather than the actual action, you approved the wrong thing.
This is why the better guides, including StackAI's approval workflow guide, insist on showing the reviewer the real action and its evidence. Their phrase for it: "the evidence pack is the difference between a 15-second approval and a 15-minute investigation."
Governance frameworks expect a person to be accountable
The NIST AI Risk Management Framework organises AI risk work around four functions: Govern, Map, Measure and Manage. It is voluntary, and it was not written with a solo founder's inbox agent in mind. But the direction is clear: someone should know what the system can do, which risks matter, and who is responsible when it acts. A written list of "actions that need my yes" is a small, practical version of that.
What you actually need: which actions should wait for your yes
Here is the heart of it. You do not want to approve everything, because you will stop reading. You want to approve the things that are costly, public, or hard to undo. Cloudflare's agent docs put it simply: "Only require confirmation for actions with meaningful consequences (payments, emails, data changes)."
The approval tiers
This table is a starting point. Adjust it to your business, but keep the shape.
| Action type | Examples | Can you undo it? | Recommended gate |
|---|---|---|---|
| Reading and research | Searching the web, reading your docs, pulling analytics | Nothing to undo | Fully automatic |
| Drafts and internal notes | Drafting an email, writing a brief, updating an internal dashboard | Yes, easily | Automatic, review when convenient |
| Reversible internal changes | Moving a task, tagging a lead, renaming a file | Yes, with a little effort | Automatic, with a visible log of who changed what |
| Messages sent as you | Emails to customers, replies to investors, DMs, support answers | No, it has been read | Approval required |
| Spending money | Ad budget changes, purchases, refunds, paid API calls at scale | Sometimes, often not | Approval required, show the amount |
| Public anything | Social posts, blog publishing, press replies, public repo changes | Partly, screenshots last | Approval required |
| Deletions and overwrites | Deleting records, files, contacts, or overwriting a sheet | Rarely | Approval required, show what disappears |
| Access and permissions | Inviting users, sharing folders, changing roles or API keys | Yes, but damage can happen first | Approval required |
A two-question test for everything else
When you are unsure about an action, ask:
- If the agent gets this wrong, who notices first? If the answer is "a customer", "the public" or "the bank", gate it.
- Could I fully undo it in five minutes? If not, gate it.
An action that fails both questions is a firm yes-required. An action that passes both can run on its own, as long as you can see afterwards what happened.
What the approval request should show
The request is only as good as what it shows you. Borrowing from StackAI, Cloudflare and Jake's piece, a good request includes:
- The exact action: the actual recipients, amount, record IDs or post text, not the agent's summary of them.
- Reach and cost: how many people, how much money, which account.
- The reason: one line on why the agent wants to do this, and what it looked at.
- Reversibility: can this be undone, and how.
- An expiry: when this request stops making sense (a reply to yesterday's thread probably should not go out next week).
If a request makes you open three other tabs to decide, it is badly designed. Ask the agent to include what you would otherwise go and look up.
Step-by-step: how to design an AI agent approval workflow
You can do all of this without touching a framework. The steps below work whether your agent is Claude, Cursor, Grok or something you built.
-
List the actions your agent can take. Write down every tool or connection it has: email, calendar, CRM, ad accounts, file storage, social accounts. OWASP's first two root causes (too much functionality, too many permissions) are often fixed right here by removing access the agent does not need.
-
Sort each action into a tier. Use the table above. Be strict with anything that sends, spends, deletes, publishes or grants access. Everything else can usually run automatically as long as it is logged.
-
Separate proposing from doing. This is the most important design choice. The agent should be able to prepare the action (draft the email, build the refund list) but not carry it out itself. Execution happens in a separate step that only runs after your approval. If the agent holds both the pen and the send button, your approval is a suggestion, not a gate.
-
Decide where approvals land. Pick one place you will actually check: a review queue, a dashboard, a chat channel you watch. One place beats three, because approvals scattered across tools get missed. Think about where you are when requests arrive; for a lot of operators that means a phone.
-
Write the request template. Tell your agent, in its instructions, exactly what every approval request must contain: action, parameters, reach, cost, reason, reversibility, expiry. Agents are good at following a fixed shape once you give them one.
-
Decide what happens after yes, no, and silence. Yes should trigger the execution step. No should tell the agent why, so it can revise. Silence should not quietly become yes; set an expiry and let stale requests lapse. StackAI and Cloudflare both recommend explicit timeouts and escalation rather than leaving requests waiting forever.
-
Review the log weekly. Look at what you approved, rejected and edited. If you approve 100 per cent of a certain action for a month, consider moving it down a tier. If you keep rejecting one, the agent's instructions need work.
What good looks like
A healthy approval workflow feels boring, which is the point. Some signs you have it right:
- You get a handful of requests a day, not dozens. Jake's piece suggests keeping each reviewer to roughly ten items a day so they do not start skimming. Your number might be different, but "a handful" is the right order of magnitude.
- Each request takes seconds to judge. Everything you need is in the request itself.
- Nothing high-stakes happens without a record. You can answer "who approved this, and when?" for every email sent or pound, dollar or euro spent.
- The agent still does most of the work. It researches, drafts and prepares. You make the final call on a short list of things.
- Rejections teach something. A rejected request comes with a reason the agent can use next time.
A concrete example. A marketer runs an agent that monitors ad spend. Every morning it pulls the numbers, updates a dashboard, and flags two campaigns that are over their cost per acquisition target. It proposes pausing one and cutting the budget on the other, with the current spend, the target, and the last seven days of results attached. The marketer reads the two requests over coffee, approves the pause, rejects the budget cut with "launch week, leave it", and the approved change goes through. Total time: about a minute.
Mistakes to avoid
-
Gating everything. If every file read and draft needs a tap, you will start approving on reflex. This is often called approval fatigue: the control still exists on paper, but it no longer protects you. Keep approvals for the actions in the red rows of the table.
-
Approving the summary, not the action. "Email the churned customers" is not an approval request. The list of addresses and the email body is. If you approve prose, the actual action can differ and you will not know until after.
-
Letting the agent execute its own approvals. If the same agent that proposes the action also checks whether you said yes, a prompt injection or a plain mistake can skip the check. Keep execution in a separate step, ideally one the agent cannot reach. OWASP's guidance allows approval "within the extension itself or in downstream systems"; downstream is often the sturdier choice.
-
No expiry. A request that sits for a week and then gets approved on autopilot can do something that made sense last Tuesday and makes no sense now. Let requests expire and have the agent re-propose if still needed.
-
No record of decisions. If you cannot see what was approved, by whom and when, you cannot learn from mistakes or explain them to a client or co-founder. Keep a log, even a simple one.
How the frameworks do it (for the developers reading)
If you write your own agents, every major framework now has a native pattern for this. The shared idea is: the run pauses before the tool executes, state is saved, a person decides, and the run resumes.
| Framework | How approval is declared | What happens at the pause | Source |
|---|---|---|---|
| OpenAI Agents SDK | Tools marked as needing approval | The run returns interruptions plus a resumable state; you approve or reject, then resume from that state | OpenAI docs |
| LangChain / LangGraph | Human-in-the-loop middleware with a per-tool interrupt policy | Execution halts; the reviewer can approve, edit or reject; a checkpointer is required to persist state | LangChain docs |
| Microsoft Agent Framework | Functions wrapped or decorated as approval-required | The agent returns an approval request containing the function name and arguments; you send back approve or reject | Microsoft Learn |
| Cloudflare Agents | Confirmation required for chosen tools | Pending approvals are stored in agent state; scheduled reminders and escalation handle waiting | Cloudflare docs |
| Claude Code | Permission modes that set what runs without asking | Claude stops and asks before gated actions in Manual mode | Claude Code docs |
Two details worth noting. OpenAI's guide is candid that the SDK gives you the pause, not the whole system: if review takes time, you "serialize state, store it, and resume later," and building the review interface is up to you. And Microsoft's framework, by default, only accepts an approval that matches a pending request recorded in the same session, and warns that turning this off "can let a fabricated or replayed response approve a privileged tool call." That is a good principle for any approval system: a yes should be bound to one specific request.
The OpenAI Agents SDK for JavaScript has its own human-in-the-loop guide if you work in TypeScript.
Approving without writing framework code
Most of the people reading this are not building agents from scratch. You use Claude in the app or Claude Code, Cursor, Grok or OpenClaw, and you connect them to your tools. You cannot add an interrupt node to someone else's agent. So how do you get a real approval step?
The pattern that works is the one from step three: the agent proposes into a queue, you decide in the queue, and something else executes. Concretely:
- The agent does the work and writes each high-stakes action into an approval list, with the exact details.
- You open that list on whatever device you have, and approve or reject each item.
- Your decision goes to a place you control, which triggers the actual action, or does nothing if you said no.
The agent never gets the send button for gated actions. It only gets the ability to ask.
This also answers the "approve AI agent actions from phone" question. You do not need a phone-specific framework. You need the queue to be readable on a phone, and the decision to land somewhere that can act on it. The part that does need some setup is the "something else executes" end: an automation tool you already use that accepts webhooks, or a small endpoint someone on your team maintains. That piece is worth getting right, because it is where the gate is actually enforced.
How Unorderly helps
Unorderly is an iPhone app (it adapts to iPad) where your agent publishes live dashboards, and one of its native widget types is an approval list. You add your Unorderly MCP URL (shown in the app) to your agent and sign in once; the remote MCP server explainer covers what that means, and the custom connector guide walks through it for Claude. Your agent can then put approval items next to the KPIs, tables and to-dos it is already tracking.
When you open the app, you approve or reject each item, and the decision is POSTed to your own webhook. Only you can set that webhook, and the agent never sees it, which fits the "agent proposes, something else executes" design above. Each view also shows which agent wrote it and when, so you know whether a request came from Claude ten minutes ago or Grok yesterday. Deleted views sit in a recycling bin for 30 days, and agents cannot delete whole workspaces.
To be clear about what it does not do yet: push alerts, background refresh and Home or Lock Screen widgets are coming soon, not available. Today you open the app, pull to refresh, and clear the queue. Unorderly is pre-launch with a waitlist, and the how-it-works section shows the flow end to end.
Frequently asked questions
What is human-in-the-loop AI agent approval?
It is a checkpoint where your agent proposes an action and a person has to say yes before it happens. The agent can still research, draft and plan on its own. The approval only sits in front of the actions that are costly, public or hard to undo.
Which AI agent actions should always need approval?
Anything that spends money, sends a message as you, deletes or overwrites data, publishes something public, or changes who has access to what. A useful test is to ask what happens if the agent gets it wrong and whether you can undo it in five minutes. If the answer is no, put it behind an approval.
Won't approvals slow my agent down?
Only if you gate everything. The agent keeps doing the reading, drafting and preparing, and only the final irreversible step waits for you. Batch approvals into a queue you clear a few times a day and most of the delay disappears.
Can I approve AI agent actions from my phone?
Yes, if the approval request lives somewhere you can reach on your phone and your decision is sent somewhere that can act on it. The key is that the phone shows the exact action, not a summary, and that the actual execution happens only after your decision is recorded.
Do I need LangGraph or the OpenAI Agents SDK to build an AI agent approval workflow?
Not necessarily. Those frameworks have built-in ways to pause a run and wait for a person, which is ideal if you are writing the agent yourself. If you use an off-the-shelf agent like Claude or Cursor, you can get a similar result by having the agent propose actions into a review queue and letting a separate step carry out only the approved ones.
What is approval fatigue?
It is what happens when an agent asks for approval so often that you stop reading and start tapping yes. The approval step still exists, but it no longer protects you. The fix is to gate fewer, higher-stakes actions and make each request quick to judge.
What should an approval request show?
The exact action and its parameters, who or what it affects, the cost or reach, the agent's reason, and whether it can be undone. A rejection should also let you say why, so the agent can do better next time.
How does Unorderly handle agent approvals?
Your agent can add approval items to a live dashboard in the Unorderly iPhone app. You open the app, approve or reject each item, and the decision is POSTed to a webhook that only you set, so the agent never sees where it goes. Unorderly is pre-launch with a waitlist, and push alerts are not available yet, so today you open the app to clear the queue. More answers are in the FAQ.