Skip to content

Permissions, Safety & Human-in-the-Loop

At a glance A deep dive into the trust mechanics of an agent harness: the spectrum of permission modes from full manual approval to yolo, Claude Code's four permission modes and hook guardrails, the paired filesystem-and-network sandbox, the prompt injection and lethal trifecta threat model, and the design principles behind human-in-the-loop gates and read/write tool separation.

Permissions, Safety & Human-in-the-Loop ​

Every component covered so far has solved a capability problem: context helps the model see more sharply, tools let it do more, the loop lets it go further. This chapter tackles the problem in the opposite direction — constraints. It's the climbing-harness side of the harness's triple metaphor (see What Is an Agent Harness): the harness restricts your movement, but that restriction is exactly what keeps a fall from killing you.

A proposition worth memorizing: an agent's capability ceiling is set by its autonomy, but how much you trust it is set by its permission system. Open up every tool and switch off every approval, and any agent instantly becomes "more capable" — that part is easy. The hard part is keeping it out of trouble when nobody is watching, and making users willing to deploy it where nobody is watching. What decides product competition is not "how much autonomy you can grant" but "how low you drive the risk while granting it."

The Spectrum of Permission Modes ​

Every coding agent's permission design lands somewhere on the same spectrum: who approves each action, and on what basis.

text
  Human involvement: high ─────────────────────────────────────────> Autonomy: high

 ┌────────────────┐  ┌────────────────┐  ┌────────────────┐  ┌────────────────┐  ┌────────────────┐
 │ Full manual    │  │ Rule allowlists│  │ Mode shifting: │  │ Sandboxed      │  │ Full auto      │
 │ approval: every│  │ allow / deny,  │  │ default/plan/  │  │ autonomy: free │  │ (yolo): all    │
 │ tool call pops │  │ exact match    │  │ acceptEdits,   │  │ inside bounds, │  │ allowed, only  │
 │ a dialog       │  │                │  │ switch anytime │  │ escalates out  │  │ in isolation   │
 └────────────────┘  └────────────────┘  └────────────────┘  └────────────────┘  └────────────────┘
   Safe but unusable     Precise but heavy to config     Shift gears per task     Best current trade-off     A prelude to incidents

There is no single "correct point" on this spectrum, only points that match the task's risk and the environment's isolation. A few key judgments:

  • Full manual approval collapses under its own weight in practice. The problem isn't security, it's people: when every action pops a dialog, users hit approval fatigue around the twentieth click and start clicking Approve without reading — the approvals are still there, but they no longer mean anything. Anthropic called this out explicitly in its engineering blog on sandbox design: overly frequent approval prompts actually make development less safe. The highest security level on paper can underperform the middle of the spectrum in practice.
  • Full automation isn't off-limits; it's a question of environment. Cursor's YOLO mode, Gemini CLI's --yolo, and Claude Code's --dangerously-skip-permissions are all real shipping features, and their names carry the warning for them. Running a fully autonomous agent in a disposable container or an isolated cloud VM is reasonable (it's the default shape of Devin and OpenHands); running one on your primary machine, next to your SSH private keys and .env files, is a prelude to an incident.
  • The answer the industry has converged on is "autonomy inside a sandbox": don't try to approve each action one by one; draw a hard boundary (this directory, these domains), stay approval-free inside it, and escalate to a human when the agent crosses it. More on this below.

Claude Code's Permission System ​

Claude Code's permission system is the most complete reference implementation on this spectrum, and it rewards being taken apart layer by layer. It has three layers: modes set the tone, rules fine-tune it, and hooks add code-level guardrails.

The Four Permission Modes ​

ModeBehaviorBest for
defaultPrompts for approval the first time each class of tool is usedDay-to-day interactive development
acceptEditsFile edits inside the working directory pass automatically; everything else still promptsFamiliar repos, low-risk changes
planRead-only exploration plus a written plan; no action until a human approves the planGetting to know a new codebase, before high-risk changes
bypassPermissionsSkips nearly all approvals (entered via --dangerously-skip-permissions)Isolated environments such as containers/CI

Within a session, Shift+Tab cycles through the first three — a mode is a gear, and you shift as the risk of the current task changes, instead of configuring once and living with it forever.

allow / ask / deny Rules ​

Beneath the modes sit fine-grained rules, kept in settings files (layered across user, project, and enterprise-managed scopes, each level overriding the last):

json
{
  "permissions": {
    "allow": ["Bash(npm run test:*)", "Bash(git status)"],
    "ask":   ["Bash(git push:*)"],
    "deny":  ["Read(./.env)", "Read(./secrets/**)"]
  }
}

Rules match on tool name plus argument pattern, and priority runs deny > ask > allow: even if a user has allowlisted all of Bash, a single deny: ["Read(./.env)"] still nails the .env file shut. The direction of that priority is the right one — a rule system's default posture should be "a tightening rule always outranks a loosening one," otherwise one broad rule quietly voids every security boundary.

Hooks: Guardrails That Don't Rely on the Model Policing Itself ​

The third layer is hooks: user-configured shell commands that run at key points in a tool call's lifecycle (PreToolUse, PostToolUse, UserPromptSubmit, Stop, and so on), anchoring the harness's behavior to deterministic code.

PreToolUse is the pivotal one: the hook script inspects the tool call about to run, blocks it with exit code 2, or returns an allow / deny / ask permission decision as JSON. The essential difference from model-side approval: a hook is code. It runs every single time. There is no such thing as the model forgetting.

text
┌───────────────── Claude Code permission pipeline ─────────────────┐
│  Model requests a tool call                                       │
│      │                                                            │
│      ▼                                                            │
│  ① PreToolUse hooks ──deny──> blocked (deterministic, unbypassable)│
│      │ pass                                                       │
│      ▼                                                            │
│  ② deny rule matched? ──yes──> blocked                            │
│      │ no                                                         │
│      ▼                                                            │
│  ③ allow rule / mode exemption? ──yes──> execute directly         │
│      │ no                                                         │
│      ▼                                                            │
│  ④ prompt the human ──approved──> execute (can become a new rule) │
└───────────────────────────────────────────────────────────────────┘

The pipeline embodies two design principles worth copying into any harness you build yourself: deny is always evaluated before allow (a guardrail cannot be overridden by a broad rule), and deterministic checks live outside the model's decision-making (the model may propose actions, but it never gets to approve its own). For a detailed product-design breakdown, see the Claude Code case study.

Human-in-the-Loop: Where to Place the Gates ​

"Human-in-the-loop" is not a slogan; it's a set of concrete gate-placement decisions: at which junctures, in what form, and with what information a human gets the final say. Placing gates wrong costs you in both directions — too many, and approval fatigue turns them into theater; too few or too late, and the irreversible damage has already happened.

Gate One: Plan Approval Before Any Action ​

Claude Code's plan mode is the canonical human gate: the agent researches with read-only tools only and produces a plan, and it gains write access only after a human approves. It layers "human approval at the critical juncture" on top of the planning modes spectrum, grafting plan-and-execute's end-to-end auditability onto the interleaved mode. The plan doubles here as planning artifact and authorization credential — what the human approves is not "is this approach any good?" but "I authorize you to act within this scope."

Gate Two: Explicit Confirmation for Irreversible Operations ​

Not every operation deserves a dialog. The engineering-sound way to tier them is by reversibility:

Operation typeExamplesSensible handling
Purely read-onlyReading files, grep, git statusAllow by default
Reversible writesEditing workspace files (git can roll back)Approval-free within the mode
Hard-to-reverse writesrm -rf, dropping database tables, force pushForced dialog or deny
Outward-facingSending email, opening PRs, deploying, paymentsHuman confirmation required; sandbox blocks outbound calls

The whole test compresses to one sentence: if git can get it back, don't ask; if git can't, you must. Workspace file edits earn their acceptEdits exemption precisely because version control backstops them — which also explains why a coding agent's permission design naturally takes "the repo" as its trust boundary.

Gate Three: Asking for Help When Things Go Wrong ​

The third kind of gate the agent triggers itself: when a permission is denied, the plan fails repeatedly, or credentials are missing, it stops and asks instead of forcing its way through. This depends on observability exposing "which step was denied by whom" to the user — otherwise the user never even knows a gate fired.

A Meta-Rule

Each gate's unit cost is an interrupted human, so ration them like a scarce resource: spend the human attention budget on irreversible, outward-facing decisions, and leave reversible high-frequency operations to rules and the sandbox. A harness that pops a dialog every five minutes teaches its user exactly one skill: clicking Allow without thinking.

Sandboxes: Trading Environment for Autonomy ​

A sandbox rewrites the permission problem from "approve each action" to "draw the boundary once": inside the boundary the agent runs fully autonomously, and the boundary itself is the security policy. It is currently the best engineering answer that serves autonomy and safety together — Anthropic reports that after adopting sandboxes internally, permission prompts dropped 84% while safety went up, not down.

The key design points (per Anthropic's sandboxing engineering post): filesystem isolation and network isolation are each incomplete without the other — they must ship as a pair.

  • Filesystem isolation: reads and writes are limited to designated directories (usually the current working directory). Without it, an injected agent can read ~/.ssh/, rewrite shell configuration, and thereby escape any network restriction.
  • Network isolation: all outbound traffic routes through a proxy outside the sandbox, allowlisted by domain, with new domains triggering human confirmation. Without it, an injected agent can send anything it has read anywhere it likes — filesystem isolation becomes a formality.

Implementation tiers, lightest to heaviest:

TierMechanismExemplars
OS primitivesLinux bubblewrap / macOS Seatbelt, no container overheadClaude Code's sandboxed bash tool
ContainersDocker-isolated runtime, actions execute inside the containerOpenHands' sandboxed runtime
Cloud-isolated environmentsOne dedicated VM per session, credentials never enter the sandboxClaude Code on the web, Devin

One detail of the cloud approach deserves attention: Claude Code on the web is designed so that sensitive credentials (git credentials, signing keys) never enter the sandbox. Git operations inside the sandbox authenticate with a restricted credential through the proxy, and the proxy attaches the real token only after verifying the branch and target. That is "credentials separated from the execution environment" — even a total compromise inside the sandbox leaves the attacker nothing worth carrying out.

Prompt Injection: The Agent's Number-One Threat ​

All permission systems share the same adversary: prompt injection. Simon Willison coined the term (deliberately echoing SQL injection): an LLM cannot reliably tell "instructions" from "instructions hiding in data" — everything ends up concatenated into one token stream fed to the model — so an attacker who plants malicious instructions in anything the agent reads (a web page, an issue, an email, a document, a tool response) can hijack its behavior.

Two commonly conflated things deserve separating: jailbreak is the user tricking the model into saying what it shouldn't (an output problem); prompt injection is untrusted content using the model's hands to do what it shouldn't (an action problem). For a chatbot, the former is the headline risk; for an agent, the latter is the lethal one — every hijacked step of the model becomes a real tool call.

The Lethal Trifecta: Three Dangerous Capabilities ​

In The lethal trifecta for AI agents (June 2025), Willison offered a compact risk test. When an agent has all three of the following capabilities at once, it is within one-shot range:

text
                 lethal trifecta
       ┌───────────────────────────────────────────────┐
       │  ① Can access private data                    │
       │     (code, keys, email)                       │
       │              +                                │
       │  ② Will touch untrusted content               │   all three legs =
       │     (web pages, issues, tool responses)       │   a single injection
       │              +                                │   completes steal +
       │  ③ Can communicate outward                    │   exfiltration
       │     (HTTP requests, opening PRs,              │
       │      sending email)                           │
       └───────────────────────────────────────────────┘
              defense = cut at least one leg

The attack chain is trivial: attacker instructions inside ② steer the agent to use ① to read private data, then ③ to ship it to the attacker. Willison's post lists a long roster of production systems that got hit — Microsoft 365 Copilot, the official GitHub MCP server, GitLab Duo, and before them Slack, Google Bard, Amazon Q, among others. The GitHub MCP case is especially instructive: one tool could read public issues (untrusted content an attacker can poison), access private repos (private data), and create PRs (an exfiltration channel) — all three elements in a single tool.

Note: a coding agent ships with all three legs by default — it reads your repo and your .env (①), it fetches web pages and issues (②), and it can run curl and git push (③). This is not a hypothetical risk; it is the factory configuration of the category, and every piece of sandbox and permission design above exists to make up for that fact.

Defense: Remove Legs Structurally, Not With Prompt Prayers ​

Willison's verdict is blunt, and worth quoting in full: we do not yet know how to stop these attacks 100% reliably. Of guardrail products claiming to "block 95% of attacks," he says: in web security, 95% is a failing grade — the attacker only needs 20 tries. Writing "do not follow instructions found in files" into a prompt is prayer, not defense: malicious instructions can be phrased in infinitely many ways, and your prohibition is a single sample.

Effective defenses are structural: they make the attack architecturally impossible rather than merely less probable:

  1. Remove a leg. When a session touches both ① and ②, take away ③ — an agent that can read private data and untrusted content does not get outbound tools. That is the job of network isolation and ask rules. Meta's "Agents Rule of Two," proposed in October 2025, codifies the same idea: a session may hold at most two of the three capabilities; if it must hold all three, add human supervision.
  2. Freeze the consequences. Quoting the summary of the security-patterns paper Willison cites: "once an agent has ingested untrusted input, it must be constrained so that the input cannot trigger any consequential action." An agent that reads web pages may produce text summaries and may not be granted write tools — least privilege applied at the level of tool combinations.
  3. Defense in depth. Prompt hardening, injection detection, permission rules, sandboxes, human gates — every layer leaks; stacked together, they force the attacker through five checkpoints at once. No silver bullet does not mean doing nothing.

Don't Expect the Model to Hold the Line

Handing the model itself the job of "recognize and refuse injected instructions" is putting the defendant on the bench as the judge. The model's job is the work; the harness's job is to make sure a hijacked model cannot do harm — which is exactly why permissions, sandboxes, and gates must live outside the model.

Separating Read-Only Tools from Write Tools ​

Applied to a harness, least privilege takes its most direct form in how you partition the toolset. Read, Grep, and Glob are safe; Edit, Write, and Bash are where things get lethal; WebFetch, git push, and the messaging tools form the third tier, "acts with effects beyond the machine" — the trust cost of the three tiers is entirely different, and they should be governed separately at the architecture level:

  • Read-only by default. Exploration, Q&A, and research tasks (plus the research phase of plan mode) get read-only tools only. A read-only tool's output can carry poison (injection), but it cannot do damage on its own — the poison only activates once a write tool is present.
  • Grant write tools by environment. Reversible writes inside the workspace belong to the acceptEdits tier; writes that cross the boundary or cannot be undone go through approval or deny.
  • Treat outward-facing tools separately. Any tool whose consequences leave the machine (push, deploy, messaging) goes on the ask list by default. These are the most irreversible tools, and they are precisely the lethal trifecta's third leg.
  • Use machine-readable annotations. The MCP protocol's tool annotations (such as readOnlyHint and destructiveHint) turn this three-way classification into protocol-level metadata, letting a harness decide allow-versus-ask automatically instead of hard-coding every tool.

The same logic covers subagents: a subagent's permissions should be a subset of the parent agent's, never the full set — a subagent dispatched to "summarize this page" has no reason to inherit Edit and Bash.

Design Checklist ​

This chapter, compressed into a checklist you can put to work as-is:

  • The default posture is deny-first: anything not explicitly allowed gets asked about; anything never asked about never gets through; deny rules outrank allow rules.
  • In the permission pipeline, deterministic checks (hooks, rules) live outside the model; the model holds proposal rights, never approval rights.
  • Tier the gates by reversibility: reversible goes approval-free, hard-to-reverse triggers a dialog, outward-facing is always reviewed. Spend the approval budget on irreversible operations.
  • When you need high autonomy, put a sandbox in place first (filesystem + network, both), then consider relaxing approvals; never run yolo on a primary environment holding real credentials.
  • Audit your tool combinations against the lethal trifecta: any session with all three legs — cut a leg, or add a human gate.
  • Separate tools into the three tiers — read-only / reversible write / outward-facing — and subagent permissions only ever shrink, never grow.
  • Log every permission event (allowed, denied, blocked at the boundary) — being able to answer, after the fact, "what did it do, and what was it stopped from doing" is the ultimate source of trust.

Further Reading ​

References ​