Appearance
Human-in-the-Loop
Human-in-the-loop (HITL) is the layer of agent engineering most often gotten wrong. Too little, and an agent teaches you a lesson with a single rm -rf; too much, and by the 30th confirmation dialog users have built the muscle memory of clicking "Approve" with their eyes closed — Anthropic's own data shows users approved 93% of Claude Code permission prompts. At that point, the "human in the loop" is no longer a safety control; it's a ritual of going through the motions.
This page won't recite the platitude that "human-machine collaboration matters." The goal is a design framework you can implement directly: which actions get auto-approved, which need approval, and which are hard-banned; how to design approval UX that doesn't fatigue; how to checkpoint and resume long tasks; how humans and agents hand control to each other; and a complete guardrail checklist for unattended scenarios.
Before continuing, read Agent Loop Explained and Agent Security & Alignment — every approval point on this page hangs off the loop's tool-call step.
1. The Autonomy Spectrum and Risk Matching
The three-tier model
Formalize the intuition first. Every agent action can be scored along two axes: reversibility (can it be undone) and blast radius (how bad if it goes wrong). Multiply the two and you get a three-tier handling policy:
| Tier | Action characteristics | Examples | Handling |
|---|---|---|---|
| L0 read-only | No side effects, idempotent | Reading files, searching code, checking logs, git status | Auto-approve, no prompt |
| L1 reversible writes | Mutates state but is rollbackable | Editing workspace files, running tests locally, installing dependencies | Approval by default; allowlist once trusted |
| L2 high-risk | Irreversible or externally visible | Deleting data, sending email, pushing to main, production deploys, payments, accessing secrets | Forced approval + some hard bans, regardless of mode |
Blast radius
↑
large │ L1 (caution) L2 (forced approval / hard-banned)
│ run local tests rm -rf / git push --force
│ edit files send email / payments / prod changes
small │ L0 (allowed) L1 (approval or allowlist)
│ read/search calls to external APIs / installs
└──────────────────────→
reversible irreversibleThe key principle: approval strength scales with the action's risk level, not with "how smart the model is." A stronger model gets things right more often, but the cost of a wrong git push --force origin main doesn't shrink accordingly. Claude Code's auto mode follows the same logic — no matter how confident the classifier is, actions like curl | bash, production deploys, and force-pushing to main are always hard-blocked.
An implementable tiered policy
Don't scatter approval logic across prompts (a prompt can be talked around via injection) — hang it at the tool-execution entry point as deterministic code:
python
from enum import Enum
class RiskLevel(Enum):
READ_ONLY = 0 # auto-approve
REVERSIBLE = 1 # approval by default; the allowlist can override
HIGH = 2 # forced human approval; deny rules take precedence
# Each rule: (matcher, risk level); match deny first, then allow, then default
TOOL_POLICY = [
("Bash(rm:*)", RiskLevel.HIGH), # deny-class: any rm gets a human look
("Bash(git push:*)", RiskLevel.HIGH),
("SendEmail", RiskLevel.HIGH),
("Bash(*)", RiskLevel.REVERSIBLE),
("Write", RiskLevel.REVERSIBLE),
("Read", RiskLevel.READ_ONLY),
("Grep", RiskLevel.READ_ONLY),
]
def gate(tool_call, allowlist: set[str]) -> str:
"""Returns auto / ask / deny. Deny rules are always evaluated before modes and the allowlist."""
level = classify(tool_call, TOOL_POLICY)
if level is RiskLevel.READ_ONLY:
return "auto"
if level is RiskLevel.HIGH:
return "ask" # high-risk actions get no exemption from any mode
if signature(tool_call) in allowlist:
return "auto" # rule-based approval
return "ask"Note that deny is evaluated before everything else — this matches Claude Code's actual semantics: disallowedTools and the deny rules in settings.json take effect under every permission mode (including bypass).
Irreversible actions have no "dry run"
Deletes, sends, payments, and deploys cannot be validated by "let the agent try once and see," because the side effect has already happened. For L2 actions, the only correct design is blocking before execution, not auditing after. Wrapping irreversible actions in a two-stage tool — "produce a plan first; once the human confirms, the agent receives the execution credential" — is far more effective than patching things up from logs afterward.
2. Approval-Point Design: Fighting Approval Fatigue
Approval fatigue is a security hole, not a UX blemish
"Confirmation fatigue is not a UX annoyance but a security vulnerability" — that line from researcher Changkun Ou deserves to be on the wall of every agent team. Anthropic's behavioral data from the "Measuring AI Agent Autonomy in Practice" study is more concrete:
- Users approved 93% of permission prompts — the vast majority of "approvals" were not effective reviews;
- About 20% of new users (fewer than 50 sessions) enabled auto-approval for everything, rising past 40% by 750 sessions — trust accumulates continuously;
- Unattended run lengths grew fast: the 99.9th percentile of continuous runtime rose from under 25 minutes to over 45 minutes within three months;
- Interestingly, senior users auto-approved more but also interrupted more proactively (novices stepped in on roughly 5% of turns, senior users about 9%). That isn't recklessness — it's a migration of supervision strategy, from "approve each action in advance" to "monitor during execution and intervene when something's wrong." On complex tasks, agents proactively stopped to clarify more than twice as often as humans interrupted them.
Conclusion: approval fatigue is inevitable over time, and your product must be designed for the reality that "no one carefully reads the 50th dialog," not for the ideal user.
Four ways to reduce the number of approvals
- Tighten the default set: read-only actions never prompt. If a 10-file refactor produces 30+ dialogs, your tiering model is leaking L0/L1 actions into the approval flow.
- Rule-based approval (allowlist): when a user approves an action, offer options like "always allow
npm test" or "always allow writes undersrc/" so that one-off approvals crystallize into rules. Thepermissions.allow/permissions.denyentries in Claude Code's settings.json are exactly this mechanism. Mind the allowlist granularity: "allowBash(npm run test:*)" is a good rule; "allowBash(*)" is running naked. - Batch approvals: bundle multiple actions of one logical unit into a single approval — "the agent will modify these 12 files, here's the diff, approve?" — instead of 12 single-file dialogs. Batch approval presupposes action preview (see Section 6): you have to know what it intends to do first.
- Tier promotion and demotion: give the agent a "trust budget." After N consecutive normal low-risk actions, automatically demote certain L1 actions from "approve each one" to "batch confirm"; one rejection or one failure, and they are promoted straight back. This turns the management instinct of onboarding a new hire into code.
What goes in an approval dialog
An effective approval needs three things: the action (down to the command and arguments — not "run a shell command"), the rationale (the agent's own statement of why), and the diff (writes show a diff; commands show the fully expanded string). Without rationale and diff, a high approval rate isn't supervision — it's a coin toss. Also provide a "reject with comment" path — the rejection reason is itself the highest-quality correction signal.
3. Interrupt, Pause, and Resume
Checkpoints are the foundation of HITL
Any "pause and wait for a human" design presumes that execution state can be persisted and resumed at any moment. If a long-running agent (tens of minutes, hundreds of tool calls) keeps state only in memory, the approval request itself becomes volatile — once the process dies, whether the human approves is moot.
Engineering requirements:
- Persist a checkpoint after every tool call, containing: message history, the current plan, the list of completed side effects, and the pending approval queue;
- Separate side effects from state: the checkpoint records "which irreversible actions have already executed," and recovery never replays them (an email already sent must not be sent twice);
- Resume = rebuild context from the checkpoint and continue the loop, transparent to the caller.
LangGraph's interrupt mechanism is currently the most mature reference implementation: interrupt() suspends execution anywhere inside a node, the checkpointer saves the graph state, and the outside world injects the human's decision back in via Command(resume=...). Key points (API per the 2026 LangGraph docs):
python
from langgraph.types import interrupt, Command
from langgraph.checkpoint.memory import InMemorySaver # swap in a persistent checkpointer for production
def execute_tool_node(state):
tool_call = state["pending_tool_call"]
if gate(tool_call, state["allowlist"]) == "ask":
# suspend: the payload must be JSON-serializable; this is what the human reviews
decision = interrupt({
"action": tool_call.name,
"args": tool_call.args,
"reason": state["agent_rationale"],
})
if not decision["approved"]:
return {"result": f"User declined: {decision.get('feedback', 'no reason given')}"}
return {"result": run(tool_call)}
graph = builder.compile(checkpointer=InMemorySaver())
config = {"configurable": {"thread_id": "task-42"}} # thread_id is the resume pointer
# first invocation: runs until the interrupt and suspends
graph.invoke({"messages": [...]}, config=config)
# after the human approves, resume with the same thread_id; the resume value becomes interrupt()'s return value
graph.invoke(Command(resume={"approved": True, "feedback": ""}), config=config)Three easy-to-hit pitfalls, all stemming from the semantics of "the node re-runs from the top on resume":
- Side effects before the interrupt get replayed — put email sends and database writes after the interrupt, or make them idempotent;
- Never wrap interrupt in a bare try/except — suspension is implemented by raising a special exception, which you would swallow;
- Multiple interrupts in one node must have a deterministic order — resume values match by index; conditionally skipping one misaligns the approvals.
For more framework-level detail, see the LangGraph deep dive.
Static breakpoints and debugging
LangGraph also supports compile-time/runtime static breakpoints, interrupt_before / interrupt_after — "always pause before entering a tool node." This serves two scenarios: stepping through node by node while debugging, and putting a uniform approval gate in front of every tool call — at the cost of coarse granularity, with no conditional skipping inside a node. A common production combination: static breakpoints for full approval in staging environments, dynamic interrupt() for conditional approval in production.
4. The Control-Handover Protocol
HITL isn't a one-way "humans review the agent" — it's a bidirectional handover of control. Design it as a protocol, with explicit signal formats in both directions.
Agent → human: requests for help
When an agent should hand over control proactively:
- Ambiguity that can't be resolved: the request has two or more reasonable readings, and the cost of choosing wrong is asymmetric (delete table A or table B?);
- Confidence below threshold: two consecutive tool calls whose results don't match expectations, or a plan that needs to be scrapped and rebuilt;
- Hitting capability/permission boundaries: the required credential doesn't exist, the target system is outside the allowlist;
- Budget alarms: tokens, time, or retries exceeding some proportion of the budget.
Don't bury the signal in the ordinary conversation stream — structure it:
json
{
"type": "escalation",
"reason": "ambiguous_requirement",
"question": "Does \"clean up old data\" mean archiving or physical deletion?",
"options": ["archive_to_cold_storage", "hard_delete", "abort"],
"context": {"affected_rows": 184320, "last_backup": "2026-08-20T03:00Z"},
"blocking": true
}blocking: true means the agent stops the clock and waits; a non-blocking blocking: false also exists ("I'll proceed with plan A; you have 10 minutes to veto"), suited to low-risk forks. Anthropic's data confirms the value of help signals: on complex tasks, agents proactively asked for clarification more than twice as often as humans interrupted them — a good agent knows when to ask.
Human → agent: course-correction injection
The reverse direction is subtler. Injecting corrections while an agent runs comes in three forms of increasing invasiveness:
- Soft prompt: append a user message to the stream; the agent naturally sees it next turn. Fits directional guidance ("ignore tests for now; fix the type errors first").
- Plan rewrite: directly edit the agent's plan/todo state and resume. Requires checkpoint support for state forking — LangGraph's time travel exists for exactly this.
- Hard interrupt + state rollback: kill the current execution, restore a historical checkpoint, and re-run with the correction attached. Cursor's Restore Checkpoint and LangGraph's
update_stateboth belong to this class.
The biggest engineering trap is injection timing: if the correction arrives while the agent is in the middle of a tool call, you have two options — wait for the current action to finish, then inject (delayed but clean state), or interrupt immediately (timely but may leave half a side effect). The principle: reversible actions can wait for completion; irreversible actions get interrupted immediately — better a half-finished product than letting it run to completion.
5. Product Case Dissections
Three representative products embody three different HITL forms: CLI permission modes, in-editor diff review, and asynchronous remote collaboration. Worth dissecting one by one.
Claude Code: a spectrum of permission modes
Claude Code implements HITL as switchable "permission modes," cycled with Shift+Tab during a session. The main tiers as of mid-2026:
| Mode | Behavior | Best for |
|---|---|---|
default | Read-only actions auto-approved; file edits, command execution, and network calls prompt one by one | Everyday interactive development |
acceptEdits | File edits pass automatically; Bash, network, and MCP calls still need approval | Refactor sessions where file changes are trusted but commands are watched |
plan | Read-only + reasoning; no file changes or command execution; produces a plan for human review first | Design review before high-risk changes |
auto | No prompts; a transcription classifier powered by Sonnet 4.6 rules on each action | Unattended runs that still need a backstop (launched March 2026, research preview) |
bypassPermissions | Skips all confirmations (YOLO mode, --dangerously-skip-permissions) | Sandboxes / containers / CI only |
Design details worth stealing:
- Deny rules precede modes:
permissions.denyand--disallowedToolswork under every mode, including bypass. Modes govern "how wide the gate opens"; deny rules govern "where the floor is" — two independent axes. - Protected-path exceptions: since v2.1.78, even in bypass mode, writes to
.git/,.claude/,.vscode/, and.husky/still prompt — these are "agent self-referential" paths where breaking things disables the guardrails themselves. - Hooks are the final deterministic gate: a
PreToolUsehook still fires in bypass mode and can hard-blockrm -rfpatterns and direct pushes to main. The classifier is the probabilistic line of defense; hooks are the deterministic one; you stack them. - The auto mode architecture: two layers of defense — prompt-injection probes on the input side, and on the output side a classifier ruling on every action; the classifier deliberately never sees tool result content, preventing malicious content in files or web pages from steering the verdict. Official benchmark numbers: a 0.4% false-block rate on real traffic and a 17% pass-through rate for "overly aggressive actions"; after 3 consecutive or 20 cumulative blocks, it automatically falls back to human approval. Simon Willison's criticism is also on record: an AI-based defense is non-deterministic, and
pip install -r requirements.txtsitting in the default allowlist means supply-chain attacks can bypass the classifier — his position is OS-level sandboxing for deterministic isolation, with AI classifiers as a supplement only. The debate itself is worth reading; the link is at the end. - Organization-level control: administrators can set
permissions.disableBypassPermissionsModein managed settings to ban bypass company-wide.
The cautionary tales come from here too. In October 2025 a user reported Claude Code running rm -rf from the root directory (GitHub issue #10077; every user file on the machine was destroyed — note that he had not enabled bypass; the permission system simply failed to block it). In December there was another incident in which trailing-argument expansion turned rm -rf tests/ patches/ plan/ ~/ into a wiped home directory. Two lessons: approval systems themselves can fail, so guardrails must be layered; and the fully expanded shell command string must appear in the approval dialog — nobody can spot the murderous ~/ in an incomplete display.
Cursor: the in-editor review loop
Cursor's HITL sits not "before the action" but "after the output" — the natural choice for an IDE form factor:
- After the agent (
Cmd/Ctrl+I) completes a round of cross-file changes, the changes appear as a diff with per-file, per-hunk accept/reject, plus a one-shot Review of everything; - Checkpoints: a codebase checkpoint is generated automatically for every request and every AI change; if the direction turns out wrong, one click of Restore Checkpoint rolls back to the state before that message — this is the hard-rollback injection from "human → agent corrections";
- Auto-Run settings determine the approval strength for terminal commands and similar actions (you can set it back to "Ask Every Time"), and
Inline Diffscontrols how diffs are presented.
Mind its limits: checkpoints are a local mechanism that excludes your own manual edits and gets cleaned up automatically — they don't replace git. The standard flow among experienced users is "commit first, then let the agent loose, use checkpoints for fine-grained rollback, and git as the final backstop." When the permission system works, checkpoints are the safety net; in bypass mode, git is the only one.
Devin: the asynchronous remote-coworker model
Devin stretches HITL to the scale of "asynchronous collaboration" — the agent runs in a cloud sandbox for tens of minutes to hours while the human is away from the terminal:
- Slack as the primary interface:
@Devinkicks off a task; Devin reports progress, asks questions, and delivers session artifacts in the channel/thread; the rest of the team is naturally "in the loop," seeing the whole process; - Plan view + live takeover: Devin maintains its own planning view, editor, terminal, and browser inside the sandbox; a human can jump into the session anytime to correct course or take over the controls;
- REST API + credential isolation: tasks can be triggered via the API as a machine user, and PRs can be configured to be created under your personal GitHub username (off by default) — identity attribution is itself part of governance.
This form answers "what do you do when no human can watch every action?": the answer isn't abandoning oversight but converting supervision from synchronous approval into an asynchronous, auditable collaboration stream. The cost is latency and context decay — by the time you see Devin's Slack question two hours later, its working context has gone cold.
None of the three forms is superior; they map to three working styles: pair programming (Claude Code default), code review (Cursor), and contractor management (Devin). Selection is really about choosing at which point in the loop, and at what granularity, the human steps in.
Further reading
Full dissections of the three products: Claude Code case study, Cursor case study, Devin case study.
6. Building and Calibrating Trust
Trust is calibrated, not accumulated
The most misread datum in Anthropic's study is "trust grows steadily with usage": novices start at 20% full auto-approval, past 40% after 750 sessions. That sounds like a trust-building success story, but flip the angle: the growth came from users self-calibrating, not from the system guiding them. After being burned a few times and watching the agent behave for hundreds more, users found their own appropriate level of intervention.
Good products should make the calibration process explicit and shrink the cost of "getting burned a few times":
- Tighter rules for novices: no bypass for the first N sessions, and allowlist rules default narrow. This isn't distrust of users — it's not letting users hand over the defenses before they understand the agent's failure modes.
- Partitioned trust: maintain different trust levels for "writing code under
src/" and "touching Terraform underinfra/." A single global trust switch will inevitably fail in the hardest domain.
Action previews beat after-the-fact logs
This is this page's most important judgment: "the agent states in one sentence what it's about to do, before doing it" contributes more to trust than any after-the-fact audit log.
Three reasons:
- Interveneability: a preview gives humans a veto window; a log only gives them grounds for blame afterward. A safety mechanism's value lies before the fact.
- Intent alignment: "I'm about to migrate the
userstable to the new schema" is itself a testable hypothesis — if the following actions don't match the preview, that's the real alarm signal. A log that records actions but not intent cannot distinguish "a planned dangerous operation" from "an injected malicious operation." - Cognitive load: reviewing one sentence of natural-language preview costs seconds; reviewing 50 tool-call logs costs giving up.
The engineering is plain: before every tool call, have the model emit a one-line intent (why, and the expected result), and put it into the approval dialog, the batch-approval summary, and the trace. Combined with the trace system from observability, you get both the preemptive defense and the post-hoc audit.
A cheap trust metric
Track "the divergence rate between stated intent and actual action." A high divergence rate means the agent is drifting or injected content is steering its behavior; a low, stable divergence rate is the evidence base for widening the allowlist. Let trust expansion follow data, not user mood.
7. The Guardrail Checklist for Unattended Agents
When a task must run unattended (code fixes in CI, overnight batch jobs, monitoring responses), "human in the loop" degrades to "a human once set rules." At that point every line of defense must be internalized into the infrastructure. A checklist to audit against:
Environment isolation (the deterministic line)
- [ ] Run in a container/sandbox with only the target directory mounted; home, system paths, and other codebases invisible
- [ ] Network egress allowlist: only the domains the task needs; arbitrary outbound connections forbidden (the main line against data exfiltration)
- [ ] Zero long-lived credentials inside the environment: no SSH keys, no production API keys; inject short-lived tokens for whatever is needed
- [ ] Disposable environments: destroyed at session end, so no damage survives across tasks
Action constraints (the rule-based line)
- [ ] The deny list exists independently of modes: hard-ban
rm -rfvariants, force-pushes to protected branches,curl | bash, production endpoints - [ ] Allowlists precise to command prefixes and paths, with stale entries audited regularly
- [ ] Irreversible actions (sending email, deploying, paying, deleting data) physically unavailable as tools to unattended agents, or forced into two-stage confirmation inside the tool
- [ ] Hard budget caps: max turns, token ceilings, per-action retry limits (3 recommended, with exponential backoff), terminate on timeout
Monitoring and backstops (the probabilistic line + humans)
- [ ] A behavior classifier or rule engine ruling on actions in real time (like Claude Code auto mode's transcription classifier), while admitting it has a miss rate — a 17% pass-through rate means it cannot be the only defense
- [ ] Full traces persisted, with action previews matched one-to-one against actual actions for post-hoc audit
- [ ] Escalate on anomalies: N consecutive blocks, divergence from stated intent, boundary touches — auto-pause and page a human (PagerDuty/Slack, not email)
- [ ] Rollback rehearsed: git checkpoints, database snapshots, or idempotent replay — "can roll back" only counts if it's been drilled, not just drawn in a design doc
Governance
- [ ] Every unattended agent has a named owner and a declared blast radius ("the worst thing it can break" written into the docs)
- [ ] Regular red-team drills: inject content into the agent's inputs and verify the defenses behave as designed
- [ ] Trust expansion driven by data: widening allowlists or loosening modes must be justified by operating metrics, not by "nothing bad happened lately"
Guardrails must defend against "the over-eager majority," not just attackers
Most agent incidents aren't injection attacks — they're agents doing something overly aggressive inside a normal task, inferring "so let's just delete it" from a vague instruction. Anthropic maintains an internal log of real incidents: agents deleting remote git branches based on ambiguous instructions, uploading an engineer's GitHub token to an internal compute cluster, running migrations against production databases. Your guardrail checklist should treat these "overly zealous" actions as the primary defense target; attackers are the secondary threat.
Closing
Mature HITL isn't "ask a human for every action" but a layered system: deterministic rules hold the floor (deny lists, sandboxing, isolation), probabilistic mechanisms cover the middle ground (classifiers, trust budgets), and humans appear only where judgment is genuinely needed (ambiguity resolution, high-risk approvals, anomaly escalation). Anthropic's own research points the same way: regulatory requirements that force "a human approves every action" manufacture friction, not safety — what matters is whether the human is positioned to monitor and intervene effectively, not the form the intervention takes.
When designing HITL, keep asking one question: if users never carefully read approval dialogs (the data says they don't), is my system still safe? Only when the answer is yes is your design finished. For implementation pitfalls, see Common Pitfalls & Anti-Patterns; for metric design around human-agent collaboration, see Agent Evaluation Methodology.
References
- The Human-in-the-Loop Illusion — Resilient Cyber — a detailed analysis of Anthropic's autonomy research (the 93% approval rate and other data) and the Auto Mode architecture, including the 0.4%/17% benchmark figures and Simon Willison's criticism.
- Claude Code --dangerously-skip-permissions Explained — TrueFoundry — the three-tier permission model, the boundaries of bypass mode, protected-path exceptions, and the timeline of real
rm -rfincidents. - Confirmation Fatigue and the Protocol Gap in Agentic AI Oversight — Changkun Ou — the argument that approval fatigue is a security vulnerability, not a UX problem.
- LangGraph Interrupts official docs — the authoritative API reference and usage rules for
interrupt()/Command(resume=...)/ checkpointers. - Cursor docs: Agent mode — the official description of reviewing changes and Restore Checkpoint.
- Devin docs: Slack integration — Devin's HITL pattern with Slack as the collaboration interface.
- How Cognition Uses Devin to Build Devin — Devin's REST API, in-Slack collaboration, and other engineering practices.