We had already done the hard part on credentials. The agent could only call a short list of tools. The broker bound the subject from the session. Secrets never sat in the prompt. Least privilege looked good on the review slide.
Then the agent proposed a refund on the wrong order.
The call was in scope. The amount was under the tool’s max. The customer id in the session was correct. The order id was not. A human caught it because someone was watching the pending queue that week. If the runtime had just executed, we would have had a clean audit line of an authorized mistake.
That is the gap this post is about. Sibling posts cover who holds the keys and how you route classify/gate decisions without asking a chat model to vote yes in prose. This one owns blast radius: AI agent approval gates when the tool can move money or change a citizen record.

Drafting is safe. Acting is not.
A chatbot that drafts a refund note, a benefits letter, or a form fill is useful and mostly reversible. Someone pastes, edits, or throws it away. The side effect never left the chat.
An agent that calls payments.refund, benefits.grant, or citizen.updateAddress is a different product. The durable state changes before anyone finishes reading the assistant’s summary. If you cannot undo that with a compensating journal entry and a support ticket, it belongs behind a gate by default.
Money, entitlement grants and revokes, identity or address changes on citizen systems, and bulk irreversible ops are our always-gate set unless a written risk tier says otherwise. Content filters and “are you sure?” prompts inside the same model conversation do not count. The model can talk itself past those.

Keys say can. Gates say should.
We already wrote about architecting agent tool access without handing over the keys. Keys answer: can this principal call this tool at all? Approval gates answer: should this specific call run now, under this policy, with whose signature?
Routing that decision with another free-text LLM vote is the wrong layer. Prefer typed rules or a decision model for classify/route. Keep the human (or dual control) when blast radius is high. One sentence, then we move on: the gate owns the pause, not the chat reply.
Three gate types that hold up

1. Confirm on irreversible
Mark the tool so the runtime pauses before execute. Show the reviewer the exact proposed call (name and args), not only a model summary. Reject returns a controlled message into the run. Approve executes once.
OpenAI’s Agents SDK documents this as needs_approval, interruptions, and state.approve / reject, with serialized run state for long waits (OpenAI Agents SDK HITL). Their API guide splits automatic guardrails from human review before side effects such as cancellations or sensitive edits (guardrails and approvals). Azure AI Foundry’s long-running HITL pattern suspends on a multi-turn task and resumes with the same task id after a human decision (Azure Foundry HITL).
Product examples: refund or void above a threshold; benefit eligibility change; citizen name or address merge; bulk notice send. Confirmation must live outside the model’s prompt loop (broker, workflow, or gateway), or it is theater.
2. Spend and threshold limits
Per-call caps are not enough. A loop of small transfers can still empty a budget. Session-level sums and counts matter.
Amazon Bedrock AgentCore temporal policies evaluate session history at the gateway with count / sum, deny-by-default, and examples such as a cumulative trade cap with per-trade approval that is consumed once (AgentCore temporal policies; AWS ML blog on temporal policies). Their docs also note that session rate limits reset with a new session id. Pair agent sessions with identity-bound design and org-level velocity limits if you need continuity across sessions.
Product examples: payouts under X auto; over X need a fresh human approval per transfer; new beneficiary always escalates; idle sessions lose write tools after a quiet period (progressive trust decay in the same AgentCore material).
3. Action allowlists (and dual control)
Closed tool catalog. Deny-by-default for anything not listed. Scan the tool call and tool response, not only chat I/O. Azure Foundry guardrails describe intervention points at user input, tool call, tool response, and output, not only chat I/O.
For high money and privileged citizen changes, add maker-checker. The agent is the maker. A distinct human role (or second control plane) is the checker. Payload frozen between propose and execute. No self-approval from the same run because the model “confirmed.” Citi’s dual-approval best practice for payments is the familiar bank shape: maker creates, checker authorizes, separate logins (Citi dual approval PDF). Temporal’s Approval pattern shows durable wait, Signals, multi-level L1/L2/L3, timeouts, and rich approval metadata in workflow history (Temporal Approval).
Product examples: only orders.refund and orders.lookup on the allowlist for a support agent; entitlement mass changes require dual control; role elevation on a citizen account is deny unless a named ops role signs.
Auto, human, dual, or deny
Put the matrix in config, not in the prompt.
| Risk signal | Typical path | Notes |
|---|---|---|
| Read, session-owned resource | Auto | Broker still binds subject |
| Small reversible write, known payee, under per-call and session caps | Auto | After shadow period on that action class |
| Irreversible, medium money, new beneficiary, citizen PII update | Human (single) | Show exact args; fail closed on timeout |
| High money, entitlement mass change, privileged status | Dual / multi-level | Agent is maker; distinct checker; no self-approve |
| Out-of-policy args, missing consent artefact, authz failure | Deny | Log as first-class; do not silent-retry into success |
Industry anchors for choosing guardrail vs HITL by use case sit in OpenAI’s approvals guide and in AgentCore’s under-threshold auto / over-threshold fresh approval patterns. Temporal is explicit that Approval is a poor fit for sub-second synchronous APIs: design async UX (“pending approval”), notify the reviewer, resume the workflow.
What the product must store
Before any side effect, freeze a pending intent:
- tool name
- canonical args (after broker subject binding)
- risk tier
- idempotency key
- policy version
- actor (which agent, which human session)
On decision, append: approve / deny / escalate / timeout, reviewer id or auto-rule id, reason, timestamps. On execute, append outcome and the same idempotency key. Keep denies and timeouts. Probing and mis-routing show up there.
OpenAI’s HITL security notes are worth taking literally: keep full run state server-side; authenticate the reviewer from session, not from a client-supplied body; authorize against stored pending ids; fail closed (HITL docs). Temporal’s best practices push the same idea for Signals: capture identity, reason, timestamp; validate Signal permissions; handle duplicates (Temporal Approval).
For citizen-facing flows in India, treat consent purpose, expiry, and revocation as hard gates before the mutating tool, not as prompt text. MeitY’s consent artefact framing and RBI Account Aggregator directions are useful analogies: authorization is an artefact plus an audit trail, not a chat affirmation. That is obligation framing for builders, not a compliance claim.
A thin reference architecture
Five steps. Keep them boring.

- Propose. The model names a tool and args. Nothing durable has changed.
- Policy check. Rules (and optional typed classifier) map to auto, human, dual, or deny. Rewrite or bind subject in the broker. Fail closed if the rule cannot safely inspect args.
- Approve or deny. Auto-rule id, or human / dual decision recorded against that exact pending intent. Sticky always-approve for the rest of a run is convenient and dangerous for payments; prefer per-call consumption on high tiers.
- Execute. Broker runs once with short-lived scoped creds. Idempotency key travels with the call.
- Audit. Append-only row for proposal, decision, and outcome. Denies count.
Shadow first when you can. AgentCore’s LOG_ONLY then ENFORCE, and Bedrock’s detect-only guardrail checks, are the same idea: observe would-be decisions on mutating proposals, then promote by action class. Do not flip all writes on at once.
agent proposes tool+args
|
v
policy check (tier, caps, allowlist)
|
+----+----+----+
| | | |
auto human dual deny
| | | |
+----+----+ +--> audit (deny)
|
v
execute once (broker)
|
v
audit (approve + outcome)
Failure modes we keep seeing
Rubber-stamp UX. If the queue is a thousand near-identical cards and the approve button is one click with no arg diff, reviewers will invent macros. Raise the auto tier for proven low-risk classes. Require a reason code on high tiers. Sample audit of approvals. Alert on reviewer approve-rate outliers (describe the practice; do not invent a percentage).
Silent retries. The agent fails, retries, and the second attempt skips the gate or reuses a spent approval. Bind approval to the intent hash and idempotency key. One approval, one execute.
Approval fatigue. Putting everything behind a human is not maturity. It is how you train people to stop reading. Controlled autonomy means some classes auto after shadow review, some always human, some dual.
Gates after the side effect. Logging “needs confirmation” after the payment API already returned 200 is not a gate. The pause sits on the tool boundary before the broker call. Temporal’s durable wait (and the same shape in Step Functions task tokens) exists for that reason.
Honest tradeoff: durable HITL adds latency. Sync chat UIs lie about that. Ship pending states and notifications, or do not ship mutating agents.
Build checklist
- Classify every tool into read / small reversible write / irreversible or regulated write. Irreversible behind an explicit gate by default.
- Freeze the intent: tool name, canonical args, risk tier, idempotency key, policy version before any side effect.
- Implement pause/resume outside the LLM (workflow Signal or task token, SDK interruptions with server-owned state, or gateway policy that requires a prior approval event).
- Risk-tier matrix in config: amount, counterparty trust, data class, irreversibility map to auto | single human | dual | deny. No self-approval on high tiers.
- Session and org limits: per-call and cumulative spend/velocity; new-beneficiary rules; idle trust decay for write tools where it fits.
- Shadow then promote mutating policies (
LOG_ONLY/ detect-only / proposal log) before ENFORCE on production writes. Promote by action class. - Audit approve, deny, escalate, timeout; reviewer or rule id; intent hash; execution result. Retain denies.
- Fail closed on high-tier timeout or reviewer-auth failure. Never treat client-supplied serialized run state as authoritative without server authz.
Optional for citizen work: check consent purpose, expiry, and revocation (or equivalent lawful-basis record) before the mutating tool. Missing consent is deny, not “ask the model.”
Where we put this in client builds
At Bluelupin we put AI agent approval gates in the same place we put brokers: on the path that can change state, not in the system prompt. We do not sell a magic autonomy switch. We ship allowlists, spend tiers, pending queues, and audit rows that name a person or a rule.
Brokers hold keys. Decision models help route. Gates own blast radius when money or citizens are on the line.