Coverage for src/ai_jury/privilege.py: 100%
643 statements
« prev ^ index » next coverage.py v7.16.1, created at 2026-09-30 06:29 +0000
« prev ^ index » next coverage.py v7.16.1, created at 2026-09-30 06:29 +0000
1"""Least-privilege auditing for review agents (OWASP LLM01 defense-in-depth).
3Reviewers process attacker-controlled content (the PR diff and, via ``--pr``, the
4PR title/body). If an agent CLI is invoked with write/tool/network powers, a
5successful prompt injection could escalate from "bad review text" to real
6side effects. The jury mitigates this by running agents read-only.
8This module both ENFORCES that restriction (:func:`enforce_read_only`, which
9every adapter that spawns a CLI routes its args through) and AUDITS it
10(:func:`audit_agent`), and it audits the argv a seat is actually spawned with
11rather than the ``extra_args`` written in the config (issue #750). The audit is
12advisory by default (a warning, surfaced via ``run_jury``); ``--strict`` promotes
13the warnings to a hard failure.
15Required read-only invocation per adapter (documented here and in docs/security.md):
17- ``claude`` : no tools at all. ``--tools ""`` (an empty allow-list of built-in
18 tools), ``--disallowed-tools`` naming every write, shell, read,
19 network and subagent tool, ``--strict-mcp-config`` with no
20 ``--mcp-config`` (so the operator's own MCP servers are not
21 loaded), ``--safe-mode``, ``--no-session-persistence`` and
22 ``--permission-mode dontAsk``. A permission mode, settings,
23 plugins, extra directories or agents the operator named are
24 kept and reported. Each is injected when absent and the deny list is
25 merged into a narrower one, unconditionally — config can add
26 denials, never remove them. A ``--tools`` list or ``--mcp-config``
27 the operator DID write is kept as written and flagged here,
28 louder when permission checks are skipped as well. The reviewer
29 only needs its prompt, which already carries the diff; a tool it
30 can call is a way for a prompt injection to read a file outside
31 the diff (``.env``) or send one to the network (``WebFetch``).
32- ``codex`` : ``-s read-only`` (the shipped default, issue #100), injected when
33 the config names no sandbox at all. A wider sandbox the operator
34 DID name (``workspace-write``/``danger-full-access``) is kept as
35 written and flagged here, so the opt-in is a knowing one.
36- ``agy``/gemini : ``--sandbox`` (the shipped default), injected unless the
37 argv already has that boolean flag — codex's ``-s read-only``,
38 which agy does not have, does not count (#902), and neither
39 does a ``--sandbox`` that is another option's value. Every
40 ``--sandbox=<value>`` is removed first, since agy reads
41 ``--sandbox=false`` as switching the flag off (#908 review).
42 ``--dangerously-skip-permissions`` / ``--yolo`` only skip an
43 approval prompt, so the sandbox beside them — whether the
44 operator wrote it or this module injected it — is what settles
45 the question as far as this module can settle it. What
46 ``--sandbox`` confines is agy's to decide: measured on agy 1.2.9,
47 the shipped argv still read and wrote files outside its
48 working directory and reached the network (docs/security.md).
49 It is the only restriction agy offers, not proof of one — so
50 agy is out of the default panel, and every seat on this
51 adapter is flagged by :func:`audit_agent` (``--strict`` fails).
52- ``cli``/``xai`` : the operator's own binary, for which this tool knows no
53 sandbox flag to add. Nothing is enforced, so for these the
54 declared ``extra_args`` really are the whole story and an
55 unsandboxed seat still warns.
56"""
58from __future__ import annotations
60from pathlib import PureWindowsPath
62from .config import GENERIC_CLI_VENDORS, normalise_vendor, spawns_process, spec_adapter
63from .redaction import redact
65# Flags that grant broad write/tool/network powers — dangerous for a reviewer.
66_DANGEROUS_FLAGS: tuple[str, ...] = (
67 "--dangerously-skip-permissions",
68 "--yolo",
69 "danger-full-access",
70 "--full-auto",
71 "workspace-write",
72)
74# Tool names that allow filesystem writes or shell execution.
75_WRITE_TOOLS: tuple[str, ...] = ("Edit", "Write", "NotebookEdit", "Bash")
77#: Claude tools that read the filesystem, reach the network, or hand the turn to a
78#: subagent that could do either. Denying only the write tools left all of these
79#: auto-approved under ``--dangerously-skip-permissions``: an injected instruction
80#: in the diff could ``Read`` the repository's ``.env`` and ``WebFetch`` it out.
81#:
82#: Only names the CLI still has. Claude Code 2.1.236 answers a deny rule for a
83#: tool it no longer ships (``LS``, ``NotebookRead``) with "Permission deny rule
84#: … matches no known tool" on stderr, on every run, and ``classify_stderr``
85#: reads that line as ``permission_prompt`` whenever the seat fails for another
86#: reason. ``--tools ""`` already leaves neither available.
87_READ_NETWORK_TOOLS: tuple[str, ...] = (
88 "Read",
89 "Grep",
90 "Glob",
91 "WebFetch",
92 "WebSearch",
93 "Task",
94 "Agent",
95)
97#: Every tool a claude reviewer is denied. The deny list is the second layer: the
98#: first is ``--tools ""``, which leaves no built-in tool available at all, and is
99#: an allow-list, so it also covers a tool a later Claude Code release adds. The
100#: shipped default in ``config.DEFAULT_CONFIG`` spells this same list out (config
101#: cannot import this module); ``tests/test_privilege.py`` keeps the two equal.
102_CLAUDE_DENIED_TOOLS: tuple[str, ...] = _WRITE_TOOLS + _READ_NETWORK_TOOLS
104#: claude's permission settings that approve a tool call without asking. With no
105#: tool available there is nothing for them to approve; with one available, they
106#: are what turns "the model asked for a file" into "the model read the file".
107_CLAUDE_APPROVING_MODES: tuple[str, ...] = ("bypassPermissions", "auto")
109#: The permission mode every read-only claude call runs in: a tool call that is
110#: not pre-approved is denied, never prompted for (so ``-p`` cannot hang) and
111#: never waved through. Injected when the argv names no ``--permission-mode``; a
112#: mode the operator DID name is kept and reported by :func:`audit_agent`. The
113#: write role swaps it for the bypass it had before.
114_CLAUDE_REVIEW_MODE = "dontAsk"
116#: claude options that point the CLI at configuration beyond its prompt:
117#: settings files and sources (hooks live there), plugins, extra directories
118#: (their CLAUDE.md), custom agents. ``--safe-mode`` keeps all of them from
119#: loading — measured on Claude Code 2.1.236, a ``--settings`` file's
120#: SessionStart/UserPromptSubmit hooks, ``--setting-sources project`` in a
121#: checkout with hooks and a CLAUDE.md, and an ``--add-dir`` holding a CLAUDE.md
122#: each loaded nothing under ``--safe-mode`` (and the ``--settings`` hooks ran
123#: without it). Kept as written, then, and reported: a reviewer needs none of
124#: them, and an operator should not be surprised to find one ignored.
125_CLAUDE_CONFIG_OPTIONS: tuple[str, ...] = (
126 "--settings",
127 "--setting-sources",
128 "--plugin-dir",
129 "--plugin-url",
130 "--add-dir",
131 "--agents",
132 "--agent",
133)
135# The subset of _DANGEROUS_FLAGS a sandbox does NOT settle, because they SELECT a
136# sandbox themselves rather than merely skipping an approval prompt (issue #750).
137#
138# `--dangerously-skip-permissions` and `--yolo` suppress a confirmation; the
139# sandbox beside them still confines the agent, which is why the shipped agy
140# default pairs the two (issue #100). codex's `--full-auto` is shorthand for a
141# *workspace-write* sandbox, and `_ensure_value_sandbox` looks only for an
142# `-s`/`--sandbox` token — so `enforce_read_only` passes `-s read-only` alongside
143# `--full-auto` without removing it, and which of the two the CLI honours is that
144# CLI's own argument-precedence rule, not something this module can assert. The
145# other two entries (`workspace-write`, `danger-full-access`) are sandbox VALUES,
146# which enforcement leaves exactly as written and the audit already reaches.
147#: codex flags that pick a *different* sandbox, and ones that remove the sandbox
148#: altogether. Two lists because the operator has to be told which: `--full-auto`
149#: selects workspace-write and leaves a sandbox, while `--yolo` is the documented
150#: alias of `--dangerously-bypass-approvals-and-sandbox` and leaves none at all.
151#: A name is not a meaning — agy's `--yolo` only skips approval prompts and stays
152#: inside its sandbox, so reading one CLI's spelling with another's dictionary is
153#: how the most dangerous flag on the codex path went unmentioned (#750).
154_CODEX_SANDBOX_SELECTORS: tuple[str, ...] = ("--full-auto",)
155_CODEX_SANDBOX_DISABLERS: tuple[str, ...] = (
156 "--yolo",
157 "--dangerously-bypass-approvals-and-sandbox",
158)
161#: The only argv values a warning repeats (#908 review, round 5): words from the
162#: CLIs' own vocabularies. A warning is printed to a terminal and a CI log, and
163#: an argv value can be anything an operator put in ``extra_args`` — a
164#: ``--settings`` JSON with an API key in it was echoed whole — so any other
165#: value is shown as ``<value>``, whatever it looks like.
166_SHOWN_VALUES: frozenset[str] = frozenset(
167 {
168 # claude permission modes, accepted and hidden
169 "acceptEdits",
170 "auto",
171 "bypassPermissions",
172 "default",
173 "dontAsk",
174 "manual",
175 "plan",
176 # codex sandbox values
177 "read-only",
178 "workspace-write",
179 "danger-full-access",
180 # Go's ParseBool spellings (agy's `--sandbox=<value>`)
181 "1",
182 "t",
183 "T",
184 "TRUE",
185 "true",
186 "True",
187 "0",
188 "f",
189 "F",
190 "FALSE",
191 "false",
192 "False",
193 }
194)
196#: The ``--tools`` names a warning may repeat: a fixed list, never a shape (#908
197#: review, round 6 — a letters-only rule printed ``CorrectHorseBatteryStaple``).
198#: Claude Code's built-in tools as this module already names them in the deny
199#: list (:data:`_CLAUDE_DENIED_TOOLS`, each checked against 2.1.236), plus
200#: ``TodoWrite``, the built-in that only edits the session's todo list, and the
201#: ``default`` keyword ``claude --help`` documents for ``--tools`` ("``default``
202#: to use all tools"). Anything else, an ``mcp__…`` tool included, is ``<tool>``.
203_SHOWN_TOOLS: frozenset[str] = frozenset({*_CLAUDE_DENIED_TOOLS, "TodoWrite", "default"})
206def _shown_value(value: str) -> str:
207 """*value* if it is in :data:`_SHOWN_VALUES`, else ``<value>``."""
208 return value if value in _SHOWN_VALUES else "<value>"
211def _shown_flag(token: str) -> str:
212 """A ``--flag`` or ``--flag=value`` token with its value passed through :func:`_shown_value`."""
213 if "=" not in token:
214 return token
215 flag, value = token.split("=", 1)
216 return f"{flag}={_shown_value(value)}"
219def _shown_positions(found: list[tuple[int, str]], offset: int) -> str:
220 """Unplaced risky tokens, by position in ``extra_args`` and the marker each carries.
222 The token itself is never shown (#908 review, round 5): it can hold a whole
223 ``--settings`` JSON. *offset* is how many tokens enforcement put in front of
224 the configured ones, which it only ever prepends.
225 """
226 return ", ".join(
227 f"item {i - offset + 1} of `extra_args` (it mentions `{mark}`)" for i, mark in found
228 )
231def _disallowed_tools_at(args: list[str], i: int) -> tuple[str, int] | None:
232 """Read a ``--disallowed-tools`` flag at *args[i]*, in either spelling.
234 Returns ``(value, span)`` — the comma-separated tool list and how many
235 tokens the flag occupies (2 for ``--disallowed-tools Edit,Write``, 1 for
236 ``--disallowed-tools=Edit,Write``) — or ``None`` when the token does not
237 start a readable flag, including a trailing space form with no value left.
239 Shared by :func:`_ensure_claude_disallowed` (which enforces the flag) and
240 :func:`_claude_is_locked_down` (which audits it) so the two spellings
241 cannot drift apart again (issue #717): the audit knew only the space form,
242 so a seat configured with ``--disallowed-tools=Edit,Write,NotebookEdit,Bash``
243 was enforced correctly, reported as *not* read-only, and aborted ``--strict``.
244 """
245 a = args[i]
246 if a == "--disallowed-tools":
247 return (args[i + 1], 2) if i + 1 < len(args) else None
248 if a.startswith("--disallowed-tools="):
249 return a.split("=", 1)[1], 1
250 return None
253def _args_str(extra_args: list[str]) -> str:
254 return " ".join(extra_args)
257# Codex sandbox VALUES that actually restrict the agent. A value sandbox like
258# ``workspace-write`` / ``danger-full-access`` does NOT (issue #292): the audit
259# must not treat the mere presence of a ``-s``/``--sandbox`` token as proof of a
260# read-only run when its value grants write/tool powers.
261_RESTRICTING_SANDBOX_VALUES: tuple[str, ...] = ("read-only",)
264def _is_sandboxed(extra_args: list[str], vendor: str = "") -> bool:
265 """True when a non-claude agent runs under a *restricting* sandbox.
267 Vendor-aware (issue #292) so a bare ``--sandbox`` token cannot give false
268 assurance: only the agy/gemini terminal sandbox is a genuine boolean
269 ``--sandbox``; for codex the sandbox takes a VALUE and only ``read-only``
270 restricts (``-s read-only`` / ``--sandbox read-only``). A bare ``--sandbox``
271 from any other vendor — e.g. ``["--sandbox", "--dangerously-skip-permissions",
272 "--yolo"]`` — is no longer accepted as a sandbox. When a restricting sandbox
273 is active, an otherwise-broad flag no longer grants real powers (issue #100).
275 The VALUE form is codex's, and only codex's (#901): it was accepted from every
276 vendor, so a bring-your-own ``aider`` or ``cursor-agent`` seat that wrote
277 ``--sandbox read-only`` — the very flag the audit's own warning names — was
278 audited as confined and passed ``--strict``, though neither CLI has such a
279 sandbox (cursor-agent's ``--sandbox`` takes ``enabled``/``disabled``, aider has
280 none). The codex identity is :func:`_is_codex`, the rule enforcement uses to
281 inject ``-s read-only``, so what counts as confined here is what enforcement
282 would have produced there.
283 """
284 vendor = normalise_vendor(vendor)
285 is_agy = vendor == "google"
286 is_codex = _is_codex(vendor)
287 args = _codex_options(extra_args) if is_codex else list(extra_args)
288 for i, a in enumerate(args):
289 # Equals form (issue #316/L-6): `-s=read-only` / `--sandbox=read-only`,
290 # which `enforce_read_only._ensure_value_sandbox` already recognizes — so
291 # the audit must too, or it false-positives a genuinely-safe config under
292 # `--strict`.
293 if a.startswith(("-s=", "--sandbox=")):
294 value = a.split("=", 1)[1]
295 if is_codex and value in _RESTRICTING_SANDBOX_VALUES:
296 return True
297 continue
298 if a in ("-s", "--sandbox"):
299 nxt = args[i + 1] if i + 1 < len(args) else ""
300 # Codex only: an explicit read-only sandbox value.
301 if is_codex and nxt in _RESTRICTING_SANDBOX_VALUES:
302 return True
303 # agy/gemini: its boolean --sandbox, read by the same rule enforcement injects
304 # it by (#902), so what counts as confined here is what enforcement produced.
305 return is_agy and _agy_bare_sandbox(args)
308#: agy options that take a value, from ``agy --help`` (1.2.12): each answered
309#: ``flag needs an argument`` when given alone (``agy --model``, checked by hand;
310#: no prompt, no model call). agy parses its flags with Go's ``flag`` rules, which
311#: hand such an option the next token *whatever it looks like*, so in
312#: ``--model --sandbox`` the ``--sandbox`` is a model name, not the sandbox (#908
313#: review) — the agy counterpart of :data:`_CLAUDE_VALUE_OPTIONS`. Names without
314#: dashes: Go accepts ``-x`` and ``--x`` alike.
315_AGY_VALUE_OPTIONS: frozenset[str] = frozenset(
316 {
317 "add-dir",
318 "agent",
319 "conversation",
320 "effort",
321 "i",
322 "input-format",
323 "json-schema",
324 "log-file",
325 "mode",
326 "model",
327 "output-format",
328 "p",
329 "print",
330 "print-timeout",
331 "project",
332 "prompt",
333 "prompt-interactive",
334 }
335)
337#: The values Go's ``strconv.ParseBool`` reads as true. agy 1.2.12 parses
338#: ``--sandbox=<value>`` with it (``--sandbox=`` answers ``invalid boolean value
339#: … strconv.ParseBool``), so ``--sandbox=false``/``0``/``f``/``FALSE`` turn the
340#: sandbox off, and — Go's flags being last-wins — override a ``--sandbox``
341#: before them. Anything outside these and the false spellings is rejected.
342_GO_TRUE_VALUES: tuple[str, ...] = ("1", "t", "T", "TRUE", "true", "True")
343_GO_FALSE_VALUES: tuple[str, ...] = ("0", "f", "F", "FALSE", "false", "False")
346def _agy_flag_name(token: str) -> str | None:
347 """The name *token* gives agy's Go flag parser (``-x``, ``--x``, ``--x=v``), or None."""
348 if not token.startswith("-") or token in ("-", "--"):
349 return None
350 body = token[2:] if token.startswith("--") else token[1:]
351 if not body or body.startswith(("-", "=")):
352 return None
353 return body.split("=", 1)[0]
356def _agy_flag_positions(extra_args: list[str]) -> list[int]:
357 """The indices of *extra_args* agy reads as flags, by Go's ``flag`` rules.
359 An option in :data:`_AGY_VALUE_OPTIONS` written without ``=`` takes the next
360 token as its value, and parsing stops at the first token that is not a flag
361 (a positional) or after ``--``: a ``--sandbox`` after either is not a flag.
362 The argv the adapter puts before these (``--input-format stream-json
363 --output-format stream-json`` and ``--model <id>``) is flags and values only,
364 so the configured args start where agy is still reading flags.
365 """
366 args = list(extra_args)
367 flags: list[int] = []
368 i = 0
369 while i < len(args):
370 name = _agy_flag_name(args[i])
371 if name is None:
372 break
373 flags.append(i)
374 i += 2 if name in _AGY_VALUE_OPTIONS and "=" not in args[i] else 1
375 return flags
378def _agy_bare_sandbox(extra_args: list[str]) -> bool:
379 """agy's boolean ``--sandbox`` at a flag position: followed by another flag or nothing.
381 The one sandbox agy has. ``agy --help`` (1.2.12) lists ``--sandbox`` with no
382 value and no ``-s`` at all, so codex's ``-s read-only`` is not a sandbox on an
383 agy seat, and neither is a ``--sandbox`` that a value follows (agy reads the
384 value as prompt text), one that is another option's value, or one after a
385 positional (Go stops reading flags there). ``--sandbox=<value>`` is not it
386 either: :func:`enforce_read_only` removes those. Shared by that function,
387 which injects ``--sandbox`` unless this holds, and :func:`_is_sandboxed`,
388 which audits it (#902): the enforcement used to accept any
389 ``-s``/``--sandbox=`` token, so an agy seat with ``-s read-only`` was spawned
390 without its real sandbox.
391 """
392 args = list(extra_args)
393 for i in _agy_flag_positions(args):
394 if not _is_agy_bare_sandbox_token(args[i]):
395 continue
396 nxt = args[i + 1] if i + 1 < len(args) else ""
397 if nxt == "" or nxt.startswith("-"):
398 return True
399 return False
402def _is_agy_bare_sandbox_token(token: str) -> bool:
403 """*token* names agy's sandbox with no inline value: ``--sandbox`` or ``-sandbox``.
405 Go's ``flag`` reads one dash and two alike, so a configured ``-sandbox`` is
406 the sandbox too (#910): counting only ``--sandbox`` injected a second one in
407 front of it.
408 """
409 return _agy_flag_name(token) == "sandbox" and "=" not in token
412def _agy_sandbox_values(extra_args: list[str]) -> list[int]:
413 """Indices of ``--sandbox=<value>`` / ``-sandbox=<value>`` at agy flag positions."""
414 args = list(extra_args)
415 return [
416 i
417 for i in _agy_flag_positions(args)
418 if _agy_flag_name(args[i]) == "sandbox" and "=" in args[i]
419 ]
422def _ensure_agy_sandbox(extra_args: list[str]) -> list[str]:
423 """agy's boolean ``--sandbox``, guaranteed; every ``--sandbox=<value>`` removed.
425 Removed first, fail-closed (#908 review): agy reads ``--sandbox=false`` as a
426 Go bool and its flags are last-wins, so injecting ``--sandbox`` in front of it
427 gave ``--sandbox --sandbox=false`` — a seat the audit accepted as sandboxed
428 and agy ran without one. A true value only repeats the flag and an
429 unparseable one stops agy starting, so removing every value is safe;
430 :func:`audit_agent` reports the ones that were not true.
431 """
432 args = list(extra_args)
433 drop = set(_agy_sandbox_values(args))
434 kept = [a for i, a in enumerate(args) if i not in drop]
435 if _agy_bare_sandbox(kept):
436 return kept
437 return ["--sandbox", *kept]
440def _agy_sandbox_switched_off(extra_args: list[str]) -> list[tuple[str, bool]]:
441 """The ``--sandbox=<value>`` tokens written to turn agy's sandbox off, as written.
443 A false value (``false``, ``0``, ``f``, ``FALSE`` …) or one Go's
444 ``ParseBool`` rejects; a true one only repeats the flag and is not reported.
445 Read from the *declared* args, since enforcement removes them. Each is
446 ``(shown, is_false)``; a value outside Go's spellings is shown as ``<value>``.
447 """
448 args = list(extra_args)
449 found: list[tuple[str, bool]] = []
450 for i in _agy_sandbox_values(args):
451 flag, value = args[i].split("=", 1)
452 if value not in _GO_TRUE_VALUES:
453 found.append((f"{flag}={_shown_value(value)}", value in _GO_FALSE_VALUES))
454 return found
457def _agy_foreign_sandbox_tokens(extra_args: list[str]) -> list[str]:
458 """Sandbox-looking tokens agy does not read as its sandbox, as written (#902).
460 codex's ``-s`` in either spelling (agy has no ``-s``) and a ``--sandbox`` with
461 a value after it, at flag positions. The boolean ``--sandbox`` is injected
462 beside them, so the seat is sandboxed as far as agy's flag goes; they are
463 reported because they do not do what they look like they do.
464 ``--sandbox=<value>`` is not here: enforcement removes it, and
465 :func:`_agy_sandbox_switched_off` reports it.
466 """
467 args = list(extra_args)
468 found: list[str] = []
469 for i in _agy_flag_positions(args):
470 a = args[i]
471 if a.startswith("-s="):
472 found.append(_shown_flag(a))
473 continue
474 if a != "-s" and not _is_agy_bare_sandbox_token(a):
475 continue
476 nxt = args[i + 1] if i + 1 < len(args) else ""
477 if nxt and not nxt.startswith("-"):
478 found.append(f"{a} {_shown_value(nxt)}")
479 elif a == "-s":
480 found.append(a)
481 return found
484def _agy_sandbox_not_read(extra_args: list[str]) -> list[str]:
485 """Sandbox tokens agy reads as something other than a flag, each with what it is (#910).
487 ``["--log-file", "--sandbox"]`` names a log file called ``--sandbox``: Go's
488 ``flag`` hands a value option the next token whatever it looks like. And
489 parsing stops at the first positional or ``--``, so a ``--sandbox`` after one
490 is prompt text. Enforcement injects the real flag when none is read; this
491 says why the configured one did not count. Named in agy's vocabulary only:
492 the token through :func:`_shown_flag`, the option by its name from
493 :data:`_AGY_VALUE_OPTIONS`.
494 """
495 args = list(extra_args)
496 flags = _agy_flag_positions(args)
497 found: list[str] = []
498 for i, a in enumerate(args):
499 if i in flags or _agy_flag_name(a) != "sandbox":
500 continue
501 owner = _agy_flag_name(args[i - 1]) if i - 1 in flags else None
502 if owner in _AGY_VALUE_OPTIONS and "=" not in args[i - 1]:
503 found.append(f"`{_shown_flag(a)}` as the value of `--{owner}`")
504 else:
505 found.append(f"`{_shown_flag(a)}` after agy has stopped reading flags")
506 return found
509def _agy_enforced_as_unknown(vendor: str) -> bool:
510 """True for a vendor :func:`enforce_read_only` gives agy's handling without being agy.
512 The vendor falls through to the last branch there: not claude, codex or
513 agy, and not a vendor that spawns no process or a bring-your-own CLI. With a
514 ``command`` it is spawned by the generic CLI adapter, without one as agy —
515 and either way every ``--sandbox=<value>`` is removed from its argv (#910).
516 """
517 vendor = normalise_vendor(vendor)
518 if vendor in _NO_SANDBOX_VENDORS or vendor.endswith("-api"):
519 return False
520 return vendor not in ("anthropic", "google") and not _is_codex(vendor)
523def _codex_options(extra_args: list[str]) -> list[str]:
524 """The part of a codex argv its parser reads options from: up to the first ``--``.
526 codex's parser (clap) reads what follows ``--`` as arguments, so a ``-s
527 read-only`` there is prompt text, not a sandbox (#908 review): it must not
528 stop the injection or pass the audit.
529 """
530 args = list(extra_args)
531 return args[: args.index("--")] if "--" in args else args
534#: Substrings that make a codex token able to widen or remove its sandbox,
535#: wherever it sits (#908 review): the wide sandbox values (also as a ``-c
536#: sandbox_mode=…`` override), the selector and both bypass spellings.
537_CODEX_RISK_MARKS: tuple[str, ...] = (
538 "danger-full-access",
539 "workspace-write",
540 "full-auto",
541 "yolo",
542 "dangerously-bypass-approvals-and-sandbox",
543)
546def _codex_unread_risks(extra_args: list[str]) -> list[tuple[int, str]]:
547 """codex tokens that could widen its sandbox and that no other warning names.
549 ``(index, marker)`` pairs, never the token (#908 review, round 5). The
550 fail-closed counterpart of :func:`_competing_sandboxes`, which reads only
551 options before ``--`` and only the selector, bypass and ``-s``/``--sandbox``
552 spellings: a marker after ``--``, inside a ``-c`` override, or anywhere else
553 is reported instead of trusted to be inert.
554 """
555 args = list(extra_args)
556 cut = len(_codex_options(args))
557 flags = (*_CODEX_SANDBOX_SELECTORS, *_CODEX_SANDBOX_DISABLERS)
558 covered: set[int] = set()
559 for i in range(cut):
560 a = args[i]
561 if any(a == f or a.startswith(f + "=") for f in flags) or a.startswith(
562 ("-s=", "--sandbox=")
563 ):
564 covered.add(i)
565 elif a in ("-s", "--sandbox") and i + 1 < cut:
566 covered.add(i + 1)
567 found: list[tuple[int, str]] = []
568 for i, a in enumerate(args):
569 mark = next((m for m in _CODEX_RISK_MARKS if m in a), None)
570 if mark is not None and i not in covered:
571 found.append((i, mark))
572 return found
575def _is_codex(vendor: str) -> bool:
576 """The identity rule :func:`enforce_read_only` uses for the codex branch.
578 Kept in one place so the audit cannot key off a different half of the seat
579 than enforcement did (#750). The rule is the **adapter key** and nothing
580 else (#758): reading the seat's name here made agy's ``--yolo`` — which only
581 skips approval prompts — get codex's meaning, where it is the alias of
582 ``--dangerously-bypass-approvals-and-sandbox``, so a correctly sandboxed
583 google seat that happened to be named ``codex-vs-gemini`` failed ``--strict``
584 on the identical argv a seat named ``agy`` passed with.
585 """
586 return normalise_vendor(vendor) == "openai"
589def _present(flag: str, args: list[str]) -> str | None:
590 """The token in *args* that spells *flag*, bare or with an ``=`` value.
592 This module treats the ``=`` spelling as first-class for ``-s`` (#316) and
593 ``--disallowed-tools`` (#717); a selector list matching only the bare token
594 let ``--full-auto=true`` through while ``--full-auto`` warned.
595 """
596 for a in args:
597 if a == flag or a.startswith(flag + "="):
598 return a
599 return None
602def _competing_sandboxes(extra_args: list[str], vendor: str = "") -> list[tuple[str, bool]]:
603 """Sandbox statements in the argv besides the enforced read-only one.
605 Each entry is ``(token, disables)`` — ``disables`` marks a flag that removes
606 the sandbox rather than picking a different one, because the operator needs
607 to be told which.
609 :func:`_is_sandboxed` asks "is a restricting sandbox named here?" and stops at
610 the first one. That is the right question for enforcement and the wrong one for
611 an audit: codex takes the sandbox as a *value*, so a second ``-s
612 workspace-write`` sits beside the enforced ``-s read-only`` — enforcement adds
613 nothing, a sandbox token already exists — and which one the CLI honours is its
614 own argument precedence, not something this module can read.
616 A ``-s``/``--sandbox`` whose next token is another flag, or absent, states
617 no value and so nothing here. Codex only: a value sandbox is codex's syntax,
618 and on an agy seat the same tokens are not a second sandbox but a flag agy
619 does not have — :func:`_agy_foreign_sandbox_tokens` reports those (#902).
620 """
621 args = _codex_options(extra_args)
622 found: list[tuple[str, bool]] = []
623 if not _is_codex(vendor):
624 return found
625 for flag in _CODEX_SANDBOX_SELECTORS:
626 token = _present(flag, args)
627 if token:
628 found.append((_shown_flag(token), False))
629 for flag in _CODEX_SANDBOX_DISABLERS:
630 token = _present(flag, args)
631 if token:
632 found.append((_shown_flag(token), True))
633 for i, a in enumerate(args):
634 if a.startswith(("-s=", "--sandbox=")):
635 value, shown = a.split("=", 1)[1], _shown_flag(a)
636 elif a in ("-s", "--sandbox"):
637 value = args[i + 1] if i + 1 < len(args) else ""
638 shown = f"{a} {_shown_value(value)}"
639 else:
640 continue
641 if not value or value.startswith("-"):
642 continue
643 if value not in _RESTRICTING_SANDBOX_VALUES:
644 found.append((shown, False))
645 return found
648#: claude options that take one required value. claude's parser (commander)
649#: hands such an option the next token *even when it starts with ``--``*, so in
650#: ``--append-system-prompt --tools`` the ``--tools`` is prompt text, not a flag.
651#: Every reader below skips value positions, or a configured value could pass
652#: for a restriction that is not there (and the real one would not be injected).
653#: From ``claude --help`` of Claude Code 2.1.236, plus the two ``-file`` forms
654#: its ``--bare`` text names.
655_CLAUDE_VALUE_OPTIONS: frozenset[str] = frozenset(
656 {
657 "--agent",
658 "--agents",
659 "--append-system-prompt",
660 "--append-system-prompt-file",
661 "--autocompact",
662 "--debug-file",
663 "--effort",
664 "--environment",
665 "--fallback-model",
666 "--input-format",
667 "--json-schema",
668 "--max-budget-usd",
669 "--model",
670 "-n",
671 "--name",
672 "--output-format",
673 "--permission-mode",
674 "--plugin-dir",
675 "--plugin-url",
676 "--remote-control-session-name-prefix",
677 "--session-id",
678 "--setting-sources",
679 "--settings",
680 "--system-prompt",
681 "--system-prompt-file",
682 }
683)
685#: claude options declared ``<values...>``: the first value is taken whatever it
686#: looks like, then every following token up to the next one starting with ``-``.
687_CLAUDE_VARIADIC_OPTIONS: frozenset[str] = frozenset(
688 {
689 "--add-dir",
690 "--allowed-tools",
691 "--allowedTools",
692 "--betas",
693 "--disallowed-tools",
694 "--disallowedTools",
695 "--file",
696 "--mcp-config",
697 "--tools",
698 }
699)
702#: claude options declared ``[value]``: commander gives one the next token only
703#: when that token does not start with ``-``. From ``claude --help`` of Claude
704#: Code 2.1.236, with the short aliases it lists (``-d``, ``-r``, ``-w``).
705_CLAUDE_OPTIONAL_VALUE_OPTIONS: frozenset[str] = frozenset(
706 {
707 "--cloud",
708 "--debug",
709 "--from-pr",
710 "--prompt-suggestions",
711 "--remote-control",
712 "--resume",
713 "--teleport",
714 "--worktree",
715 "-d",
716 "-r",
717 "-w",
718 }
719)
721#: Every short option ``claude --help`` (2.1.236) lists, and whether it takes a
722#: value: ``None`` for a switch, ``"required"`` for ``-n <name>``, ``"optional"``
723#: for ``-d [filter]``, ``-r [value]`` and ``-w [name]``. commander reads a
724#: combined token such as ``-pn`` as ``-p -n`` (#908 review): a switch hands the
725#: rest of the token on as another short option, and a value option takes the
726#: rest of the token as its value, or, when nothing is left, the next token — so
727#: in ``-pn -- --permission-mode=bypassPermissions`` the ``--`` is the session
728#: name, and the mode after it is read.
729_CLAUDE_SHORT_OPTIONS: dict[str, str | None] = {
730 "c": None,
731 "h": None,
732 "p": None,
733 "v": None,
734 "n": "required",
735 "d": "optional",
736 "r": "optional",
737 "w": "optional",
738}
741def _claude_short_cluster(token: str) -> str | None:
742 """What a single-dash *token* leaves for the next token: its value kind, or None.
744 ``"required"`` / ``"optional"`` when the token ends on a value option with no
745 inline value (``-n``, ``-pn``, ``-cd``), ``None`` otherwise — a run of
746 switches, a value written inline (``-nfoo``), or a letter claude does not
747 know (it rejects the argv then, so no value is taken).
748 """
749 if not token.startswith("-") or token.startswith("--") or len(token) < 2:
750 return None
751 for j, letter in enumerate(token[1:], start=1):
752 if letter not in _CLAUDE_SHORT_OPTIONS:
753 return None
754 kind = _CLAUDE_SHORT_OPTIONS[letter]
755 if kind is None:
756 continue
757 return kind if j == len(token) - 1 else None
758 return None
761def _claude_value_positions(args: list[str]) -> frozenset[int]:
762 """Indices of *args* that claude reads as an option's value, not as a flag."""
763 values: set[int] = set()
764 i = 0
765 while i < len(args):
766 a = args[i]
767 if a in _CLAUDE_VALUE_OPTIONS:
768 values.add(i + 1)
769 i += 2
770 continue
771 kind = "optional" if a in _CLAUDE_OPTIONAL_VALUE_OPTIONS else _claude_short_cluster(a)
772 if kind == "required":
773 values.add(i + 1)
774 i += 2
775 continue
776 if kind == "optional":
777 if i + 1 < len(args) and not args[i + 1].startswith("-"):
778 values.add(i + 1)
779 i += 2
780 continue
781 i += 1
782 continue
783 if a in _CLAUDE_VARIADIC_OPTIONS:
784 j = i + 1
785 if j < len(args):
786 values.add(j)
787 j += 1
788 while j < len(args) and not args[j].startswith("-"):
789 values.add(j)
790 j += 1
791 i = j
792 continue
793 i += 1
794 return frozenset(v for v in values if v < len(args))
797def _claude_terminator(args: list[str]) -> int:
798 """The index of the ``--`` that ends claude's options, or ``len(args)``.
800 claude's parser (commander) reads every token after ``--`` as an argument,
801 not an option: ``--permission-mode=bypassPermissions -- --permission-mode=dontAsk``
802 runs in bypass mode, and the last token is prompt text (#908 review, checked
803 against Claude Code 2.1.236 with help-only probes). A ``--`` that is another
804 option's value (``--model --``) is that value, not the terminator.
805 """
806 values = _claude_value_positions(args)
807 for i, a in enumerate(args):
808 if a == "--" and i not in values:
809 return i
810 return len(args)
813#: Substrings that make a claude token able to change the reviewer's permission
814#: mode or tools, wherever it sits (#908 review): a ``--permission-mode…``
815#: spelling, either skip-permissions flag (``--dangerously-skip-permissions`` and
816#: ``--allow-dangerously-skip-permissions``, both in ``claude --help`` 2.1.236),
817#: an approving mode name, a settings ``defaultMode``, and the options that hand
818#: tools back (``--tools``, ``--mcp-config``). The mode words ``auto``,
819#: ``manual``, ``default`` and ``plan`` are not here: each takes effect only as
820#: the value of a ``--permission-mode`` token, which is, and ``auto`` is also an
821#: ordinary ``--autocompact`` value.
822_CLAUDE_RISK_MARKS: tuple[str, ...] = (
823 "permission-mode",
824 "dangerously-skip-permissions",
825 "bypassPermissions",
826 "acceptEdits",
827 "defaultMode",
828 "--tools",
829 "--mcp-config",
830)
833def _claude_unread_risks(extra_args: list[str]) -> list[tuple[int, str]]:
834 """Risky tokens the precise readers here do not read as flags: ``(index, marker)``.
836 Never the token itself (#908 review, round 5): a ``--settings`` JSON that
837 names ``defaultMode`` can carry an API key beside it.
839 Fail-closed on purpose (#908 review). The readers above model claude's parser
840 — value positions, ``=`` forms, ``--``, combined short options — and each
841 round of review found a corner they missed, every time one where a token
842 they skipped was one Claude applied. So a token carrying one of
843 :data:`_CLAUDE_RISK_MARKS` that those readers do not take as a flag — after
844 ``--``, as another option's value, or anywhere else — is reported rather
845 than trusted to be inert. ``--permission-mode dontAsk`` (in either spelling)
846 and ``--tools ""`` grant nothing and are not reported. Text that only
847 mentions such a token is reported too; that over-warning is the price.
848 """
849 args = list(extra_args)
850 cut = _claude_terminator(args)
851 values = _claude_value_positions(args)
852 read = {i for i in range(cut) if i not in values}
853 # The value of a `--permission-mode` that is read is read with it.
854 read |= {i + 1 for i in read if args[i] == "--permission-mode"}
855 found: list[tuple[int, str]] = []
856 for i, a in enumerate(args):
857 mark = next((m for m in _CLAUDE_RISK_MARKS if m in a), None)
858 if i in read or mark is None:
859 continue
860 nxt = args[i + 1] if i + 1 < len(args) else None
861 if a == "--permission-mode=dontAsk" or (a == "--permission-mode" and nxt == "dontAsk"):
862 continue
863 # An empty `--tools` grants nothing only when no list value follows it:
864 # `--tools` is variadic, so `--tools '' Bash` would give Claude Bash (#910).
865 after = args[i + 2] if i + 2 < len(args) else "-"
866 if a == "--tools=" or (a == "--tools" and nxt == "" and after.startswith("-")):
867 continue
868 found.append((i, mark))
869 return found
872def _claude_options(extra_args: list[str]) -> list[str]:
873 """The part of a claude argv its parser reads options from: up to the terminator.
875 Every claude reader here scans this, not the whole argv, so text after ``--``
876 can neither hide a flag in front of it nor pass for one.
877 """
878 args = list(extra_args)
879 return args[: _claude_terminator(args)]
882def _claude_flag_present(flag: str, args: list[str]) -> bool:
883 """*flag* (bare or ``flag=…``) at a flag position of *args*, not as a value."""
884 args = _claude_options(args)
885 values = _claude_value_positions(args)
886 return any(
887 (a == flag or a.startswith(flag + "=")) and i not in values for i, a in enumerate(args)
888 )
891def _ensure_claude_disallowed(extra_args: list[str]) -> list[str]:
892 """Guarantee ``--disallowed-tools`` covers every denied tool (issue #288).
894 Merges the mandatory tools — write and shell, and since the read/network
895 hardening also read, network and subagent tools — into any existing
896 ``--disallowed-tools`` value (config may ADD denials, never REMOVE the
897 mandatory ones), or injects the flag when absent. Idempotent: the shipped
898 default already lists all of them, so it is returned unchanged.
899 """
901 def _merged(value: str) -> str:
902 existing = [t.strip() for t in value.split(",") if t.strip()]
903 for tool in _CLAUDE_DENIED_TOOLS:
904 if tool not in existing:
905 existing.append(tool)
906 return ",".join(existing)
908 # Options only (#908 review): a `--disallowed-tools` after `--` is prompt
909 # text, so it is neither merged nor taken as proof the flag is present.
910 everything = list(extra_args)
911 args = _claude_options(everything)
912 rest = everything[len(args) :]
913 values = _claude_value_positions(args)
914 out: list[str] = []
915 i = 0
916 found = False
917 while i < len(args):
918 a = args[i]
919 if i in values: # another option's value, whatever it spells
920 out.append(a)
921 i += 1
922 continue
923 # Both spellings, via the shared reader: the space form
924 # (--disallowed-tools Edit,Write) and the equals form
925 # (--disallowed-tools=Edit,Write — review of #288: the exact-match check
926 # missed it, so a narrower =-value could sit after the injected safe set
927 # and, if the CLI is last-wins, narrow the deny set). The value is
928 # rewritten in the spelling it was written in.
929 hit = _disallowed_tools_at(args, i)
930 if hit is not None:
931 value, span = hit
932 found = True
933 if span == 2:
934 out.extend([a, _merged(value)])
935 else:
936 out.append("--disallowed-tools=" + _merged(value))
937 i += span
938 continue
939 out.append(a)
940 i += 1
941 if not found:
942 out = ["--disallowed-tools", ",".join(_CLAUDE_DENIED_TOOLS), *out]
943 return [*out, *rest]
946def _claude_tools_at(args: list[str], i: int) -> tuple[list[str], int] | None:
947 """Read a ``--tools`` flag at *args[i]*: ``(tool names, span)`` or ``None``.
949 claude declares the option variadic (``--tools <tools...>``), so the space
950 form takes its first value whatever it looks like and then every following
951 token up to the next flag (see :data:`_CLAUDE_VARIADIC_OPTIONS`), and each
952 token may itself be a comma or space separated list. ``--tools ""`` is one
953 empty token and so names no tool at all — the documented way to disable
954 them. The equals form (``--tools=Read,Grep``) takes exactly its own value.
955 """
956 a = args[i]
957 if a.startswith("--tools="):
958 values, span = [a.split("=", 1)[1]], 1
959 elif a == "--tools":
960 j = min(i + 2, len(args))
961 while j < len(args) and not args[j].startswith("-"):
962 j += 1
963 values, span = args[i + 1 : j], j - i
964 else:
965 return None
966 names = [t for v in values for t in v.replace(",", " ").split() if t]
967 return names, span
970def _claude_tools(args: list[str]) -> list[str] | None:
971 """The built-in tools a ``--tools`` flag leaves available, or ``None`` if absent.
973 ``None`` means the flag is not there, which leaves the CLI's whole default
974 tool set available. Repeated flags accumulate, as claude's parser does for a
975 variadic option.
976 """
977 found: list[str] | None = None
978 args = _claude_options(args)
979 values = _claude_value_positions(args)
980 i = 0
981 while i < len(args):
982 hit = None if i in values else _claude_tools_at(args, i)
983 if hit is None:
984 i += 1
985 continue
986 names, span = hit
987 found = [*(found or []), *names]
988 i += span
989 return found
992#: Flags every read-only claude invocation carries, injected when absent.
993#: ``--strict-mcp-config`` (with no ``--mcp-config``) loads none of the user's MCP
994#: servers. ``--safe-mode`` starts claude with every customization off — CLAUDE.md
995#: and its imports, skills, plugins, hooks, MCP servers, custom agents — while
996#: login, model selection and permissions work as usual (``claude --help``,
997#: Claude Code 2.1.236): measured, a project ``.claude/settings.json`` hook ran
998#: under the no-tool argv without it and not with it, and ``~/.claude/CLAUDE.md``
999#: was in the reviewer's context without it and not with it.
1000#: ``--no-session-persistence`` keeps claude from writing a transcript — which
1001#: holds the untrusted diff — under ``~/.claude/projects/`` for every call.
1002_CLAUDE_LOCKDOWN_FLAGS: tuple[str, ...] = (
1003 "--strict-mcp-config",
1004 "--safe-mode",
1005 "--no-session-persistence",
1006)
1009def _ensure_claude_locked_down(extra_args: list[str]) -> list[str]:
1010 """The claude reviewer argv: no tools, no customizations, no transcript.
1012 ``--tools ""``, the full deny list and :data:`_CLAUDE_LOCKDOWN_FLAGS`.
1014 Injection, not override, like every other enforcement here: a ``--tools``
1015 list the operator wrote is kept (the deny list still removes every tool this
1016 module names from it), and so is an ``--mcp-config``; both are what
1017 :func:`audit_agent` then reports. The injected flags go in front, as one
1018 block, and the empty ``--tools`` value is followed by a flag — the injected
1019 ``--strict-mcp-config`` or ``--disallowed-tools``, or the configured argv's
1020 own first flag — never by a bare token its variadic parser would take as
1021 another tool name. (A configured argv that starts with a bare token was
1022 already broken: claude reads that token as the prompt.)
1023 """
1024 args = _ensure_claude_disallowed(extra_args)
1025 front: list[str] = []
1026 if _claude_tools(args) is None:
1027 front += ["--tools", ""]
1028 # Flag-aware: a token that is another option's value does not count.
1029 front += [f for f in _CLAUDE_LOCKDOWN_FLAGS if not _claude_flag_present(f, args)]
1030 # The reviewer's permission mode, unless the operator named one — which is
1031 # kept, and reported by `audit_agent`. With `--tools ""` in force a mode has
1032 # no tool to approve, so keeping it grants nothing; overriding it silently
1033 # would hide a configuration the operator believes is in effect.
1034 if not _claude_flag_present("--permission-mode", args):
1035 front += ["--permission-mode", _CLAUDE_REVIEW_MODE]
1036 return [*front, *args]
1039def _ensure_value_sandbox(extra_args: list[str], default: list[str]) -> list[str]:
1040 """Ensure SOME sandbox flag is present; inject ``default`` only when none is.
1042 If the operator already specified ``-s``/``--sandbox`` (even a wider value
1043 like codex ``workspace-write``), respect it — that is a documented, audited
1044 opt-in. We only inject the secure default when no sandbox flag exists at all,
1045 which is the actual hole (empty/misconfigured ``extra_args``, issue #288).
1046 """
1047 args = list(extra_args)
1048 # Recognize both the space form (-s read-only) and the equals form
1049 # (--sandbox=read-only) so an existing sandbox is never double-specified.
1050 # Options only (#908 review): one after `--` is prompt text, not a sandbox.
1051 if any(
1052 a in ("-s", "--sandbox") or a.startswith(("-s=", "--sandbox="))
1053 for a in _codex_options(args)
1054 ):
1055 return args
1056 return [*default, *args]
1059#: Vendors whose args this module leaves exactly as configured, in either
1060#: direction — :func:`enforce_read_only` and :func:`enable_write` both return
1061#: them unchanged. Two reasons land a vendor here. Most reach their model over
1062#: the network rather than by spawning a CLI, so they have no sandbox/tool
1063#: surface at all. The generic bring-your-own-CLI profiles (`cli`, and `xai`
1064#: for a Grok seat driven through Cursor's `cursor-agent`, issue #701) do spawn
1065#: a CLI, but it is the operator's CLI: there is no vendor-specific sandbox
1066#: flag this tool could add or remove, and injecting agy's `--sandbox` into an
1067#: unrelated binary breaks the seat rather than confining it.
1068_NO_SANDBOX_VENDORS: tuple[str, ...] = (
1069 "local",
1070 "anthropic-api",
1071 "openai-api",
1072 "google-api",
1073 "xai-api",
1074 "openai-compatible",
1075 *GENERIC_CLI_VENDORS,
1076)
1078#: Which seats run no subprocess at all is no longer a vendor list here: it is
1079#: ``config.spawns_process``, which answers for the adapter ``make_adapter`` builds
1080#: (#901 review), so the audit and the spawner cannot disagree about a seat.
1083def _drop_claude_disallowed(extra_args: list[str]) -> list[str]:
1084 """Drop every ``--disallowed-tools`` flag, in both its spelling forms.
1086 The inverse of :func:`_ensure_claude_disallowed`. Removing the flag entirely
1087 (rather than narrowing its value) restores the CLI's own default tool set,
1088 which is what an implementer role needs.
1089 """
1090 everything = list(extra_args)
1091 args = _claude_options(everything)
1092 rest = everything[len(args) :]
1093 values = _claude_value_positions(args)
1094 out: list[str] = []
1095 i = 0
1096 while i < len(args):
1097 a = args[i]
1098 if i in values:
1099 out.append(a)
1100 i += 1
1101 continue
1102 if a == "--disallowed-tools":
1103 i += 2 if i + 1 < len(args) else 1
1104 continue
1105 if a.startswith("--disallowed-tools="):
1106 i += 1
1107 continue
1108 out.append(a)
1109 i += 1
1110 return [*out, *rest]
1113def _permission_mode_at(args: list[str], i: int) -> tuple[str, int] | None:
1114 """Read a ``--permission-mode`` flag at *args[i]*: ``(mode, span)`` or ``None``."""
1115 a = args[i]
1116 if a == "--permission-mode":
1117 return (args[i + 1], 2) if i + 1 < len(args) else ("", 1)
1118 if a.startswith("--permission-mode="):
1119 return a.split("=", 1)[1], 1
1120 return None
1123def _claude_write_args(extra_args: list[str]) -> list[str]:
1124 """The claude implementer argv: the reviewer lockdown lifted (#661).
1126 Drops everything :func:`_ensure_claude_locked_down` adds — the deny list, the
1127 ``--tools`` allow-list and :data:`_CLAUDE_LOCKDOWN_FLAGS` — so the CLI's own
1128 default tool set, MCP servers, CLAUDE.md, hooks and session history are back:
1129 an implementer works in the operator's own worktree, as it did before. The reviewer's ``--permission-mode
1130 dontAsk`` would deny every edit the implementer exists to make, so it becomes
1131 ``--dangerously-skip-permissions``: the flag the shipped default carried for
1132 both roles until they were split, which keeps the write role's argv what it
1133 was.
1134 """
1135 everything = _drop_claude_disallowed(extra_args)
1136 args = _claude_options(everything)
1137 rest = everything[len(args) :]
1138 values = _claude_value_positions(args)
1139 skip_bypass = _claude_flag_present("--dangerously-skip-permissions", args)
1140 out: list[str] = []
1141 i = 0
1142 while i < len(args):
1143 a = args[i]
1144 if i in values:
1145 out.append(a)
1146 i += 1
1147 continue
1148 tools = _claude_tools_at(args, i)
1149 if tools is not None:
1150 i += tools[1]
1151 continue
1152 mode = _permission_mode_at(args, i)
1153 if mode is not None and mode[0] == _CLAUDE_REVIEW_MODE:
1154 # Once: a second `dontAsk` (#888) must not add the flag again.
1155 if not skip_bypass:
1156 out.append("--dangerously-skip-permissions")
1157 skip_bypass = True
1158 i += mode[1]
1159 continue
1160 if a not in _CLAUDE_LOCKDOWN_FLAGS:
1161 out.append(a)
1162 i += 1
1163 return [*out, *rest]
1166def _set_value_sandbox(extra_args: list[str], flag: str, value: str) -> list[str]:
1167 """Replace any existing value sandbox with ``flag value`` (codex form)."""
1168 everything = list(extra_args)
1169 args = _codex_options(everything)
1170 rest = everything[len(args) :]
1171 out: list[str] = []
1172 i = 0
1173 while i < len(args):
1174 a = args[i]
1175 if a in ("-s", "--sandbox"):
1176 # Drop the flag AND its value — unless the next token is another
1177 # flag, in which case there is no value to consume.
1178 nxt = args[i + 1] if i + 1 < len(args) else ""
1179 i += 2 if nxt and not nxt.startswith("-") else 1
1180 continue
1181 if a.startswith(("-s=", "--sandbox=")):
1182 i += 1
1183 continue
1184 out.append(a)
1185 i += 1
1186 return [flag, value, *out, *rest]
1189def _drop_bare_sandbox(extra_args: list[str]) -> list[str]:
1190 """Drop agy's boolean ``--sandbox`` (both spellings), leaving the rest."""
1191 return [a for a in extra_args if a != "--sandbox" and not a.startswith("--sandbox=")]
1194def enable_write(vendor: str, extra_args: list[str]) -> list[str]:
1195 """Return ``extra_args`` with the vendor's write/tool mode enabled (issue #661).
1197 *vendor* is the seat's ADAPTER key — the protocol whose flags this function
1198 speaks — which is the vendor itself unless the seat named an ``adapter``
1199 (issue #705). The flags belong to the CLI being spawned, not to whose model
1200 answers, so a GPT seat driven through ``cursor-agent`` is handled as ``cli``.
1201 The seat's *name* is deliberately not an argument here and must not become
1202 one (#758, #768): it is free text an operator picks to tell two seats apart
1203 in a report, and letting a substring of it override the declared adapter key
1204 spawned three configurations with another CLI's flags.
1206 The deliberate mirror image of :func:`enforce_read_only`, and the ONLY place
1207 the read-only guarantee is lifted. It exists for ``jury run-agent --role
1208 implement|fix --allow-write``: an orchestrator dispatching an *implementer*
1209 needs the agent to edit files, which is precisely what a reviewer must never
1210 do. Nothing on the panel path calls this — a review/debate/verify/synthesis
1211 invocation still goes through :func:`enforce_read_only`, so a prompt
1212 injection in an attacker-controlled diff cannot reach it.
1214 Per vendor: claude drops ``--disallowed-tools``, ``--tools`` and
1215 ``--strict-mcp-config`` (restoring its own default tool set and MCP servers)
1216 and trades the reviewer's ``--permission-mode dontAsk`` for
1217 ``--dangerously-skip-permissions``, codex moves to ``-s workspace-write``, and agy (plus any unknown
1218 vendor, which routes to the agy adapter) drops the boolean ``--sandbox``.
1219 Network vendors have no such surface and are returned unchanged.
1220 """
1221 vendor = normalise_vendor(vendor)
1222 args = list(extra_args or [])
1223 # The no-sandbox vendors first, for the same reason as in enforce_read_only:
1224 # there is no write mode of theirs to enable — a network vendor spawns no
1225 # process at all, and a bring-your-own-CLI seat runs a binary whose flags
1226 # this tool does not speak — so falling through would have the agy branch
1227 # below strip a `--sandbox` that belongs to somebody else's command line.
1228 if vendor in _NO_SANDBOX_VENDORS or vendor.endswith("-api"):
1229 return args
1230 if vendor == "anthropic":
1231 return _claude_write_args(args)
1232 if vendor == "openai":
1233 return _set_value_sandbox(args, "-s", "workspace-write")
1234 return _drop_bare_sandbox(args)
1237def enforce_read_only(vendor: str, extra_args: list[str]) -> list[str]:
1238 """Return ``extra_args`` with the mandatory read-only restriction guaranteed.
1240 *vendor* is the seat's ADAPTER key — the protocol whose flags this function
1241 speaks — which is the vendor itself unless the seat named an ``adapter``
1242 (issue #705). The flags belong to the CLI being spawned, not to whose model
1243 answers, so a GPT seat driven through ``cursor-agent`` is handled as ``cli``.
1244 The seat's *name* is deliberately not an argument here and must not become
1245 one (#758, #768): it is free text an operator picks to tell two seats apart
1246 in a report, and letting a substring of it override the declared adapter key
1247 spawned three configurations with another CLI's flags.
1249 The sandbox is enforced here (issue #288) rather than left to config, so on the
1250 adapters that have enforcement an **empty** ``extra_args`` cannot produce a
1251 write-capable reviewer of an attacker-controlled diff: the sandbox is injected
1252 when the config names none.
1253 It is injection, not override, and the difference is what the audit exists to
1254 cover (issue #750). Config that names a sandbox keeps it — ``-s
1255 workspace-write`` is passed through as written — and codex's bypass flags
1256 (``--yolo``, ``--dangerously-bypass-approvals-and-sandbox``) are passed through
1257 too, so enforcement alone does not guarantee a restriction survives. The
1258 ``cli`` and ``xai`` adapters have no enforcement at all. Each of those is a case
1259 :func:`audit_agent` reports, and ``--strict`` turns into a failure. A ``local`` (network) agent runs no
1260 subprocess and is returned unchanged; neither does a hosted-API agent
1261 (issue #430) — it makes one HTTP call with no tool/file/shell access at
1262 all, so there is no ``extra_args``/sandbox concept to enforce. An
1263 **unknown vendor** is treated like agy here and gets ``--sandbox`` injected
1264 (issue #310, completes #300) — fail-closed in the sense that matters, that
1265 the flag is added rather than omitted. What *runs* is usually
1266 ``GenericCLIAdapter``, not ``AgyAdapter``: ``make_adapter`` returns the
1267 generic adapter whenever the seat sets a ``command``, and ``AgyAdapter`` is
1268 only the no-command fallback. So ``--sandbox`` reaches an unknown binary as a
1269 passthrough token, which that binary may honour, ignore, or reject — it is not
1270 agy's confinement. That uncertainty is why :func:`audit_agent` still warns for
1271 an unknown vendor instead of accepting the injected flag as proof (#292).
1272 """
1273 # `normalise_vendor`, not `.lower()`: lowercasing alone left `" XAI "`
1274 # outside `GENERIC_CLI_VENDORS`, so the xai seat fell through to the agy
1275 # branch and had `--sandbox` injected into `cursor-agent` — the exact flag
1276 # the xai profile exists to keep off that CLI (issue #701, review round 3).
1277 # A spelling validation accepts must be the spelling the guard enforces.
1278 vendor = normalise_vendor(vendor)
1279 extra_args = list(extra_args or [])
1280 # The no-sandbox vendors are checked FIRST (review of #310): a network agent
1281 # runs no subprocess to confine, and a bring-your-own-CLI seat runs the
1282 # operator's own binary, for which this tool knows no sandbox flag. Without
1283 # this branch both would reach the unknown-vendor fallback at the end and
1284 # have agy's `--sandbox` injected into a command line that never asked for
1285 # it — the exact flag the xai profile exists to keep off `cursor-agent`.
1286 if vendor in _NO_SANDBOX_VENDORS or vendor.endswith("-api"):
1287 return extra_args
1288 if vendor == "anthropic":
1289 return _ensure_claude_locked_down(extra_args)
1290 if vendor == "openai":
1291 return _ensure_value_sandbox(extra_args, ["-s", "read-only"])
1292 # google / agy / gemini AND any unknown vendor (issue #310, completes #300):
1293 # an unknown vendor routes to the generic AgyAdapter (--print/--sandbox), so
1294 # inject --sandbox like agy. An agy-compatible CLI then runs sandboxed; an
1295 # incompatible one fails on the unknown flag rather than running UNSANDBOXED
1296 # — fail-closed either way, never fail-open. Only agy's own boolean
1297 # `--sandbox` makes the injection unnecessary (#902): codex's `-s read-only`
1298 # used to, and agy has no `-s`, so that seat ran without its sandbox. And a
1299 # `--sandbox=<value>` is removed, so `--sandbox=false` cannot switch the
1300 # injected flag off again (#908 review).
1301 return _ensure_agy_sandbox(extra_args)
1304def _claude_is_locked_down(extra_args: list[str]) -> bool:
1305 """True when claude's --disallowed-tools covers every denied tool.
1307 Every one of :data:`_CLAUDE_DENIED_TOOLS` — write and shell, read, network and
1308 subagent — not only the write tools: a deny list that stopped at those left
1309 ``Read`` and ``WebFetch`` open, and this check passed it.
1311 Reads the flag through :func:`_disallowed_tools_at`, the same reader
1312 :func:`_ensure_claude_disallowed` enforces it with (issue #717), so a seat
1313 whose args this module accepts as read-only is exactly a seat that module
1314 leaves unchanged — in either spelling of the flag.
1315 """
1316 disallowed: set[str] = set()
1317 args = _claude_options(extra_args)
1318 values = _claude_value_positions(args)
1319 i = 0
1320 while i < len(args):
1321 hit = None if i in values else _disallowed_tools_at(args, i)
1322 if hit is None:
1323 i += 1
1324 continue
1325 value, span = hit
1326 disallowed |= {t.strip() for t in value.split(",") if t.strip()}
1327 i += span
1328 return all(t in disallowed for t in _CLAUDE_DENIED_TOOLS)
1331def _claude_open_surface(extra_args: list[str]) -> list[str]:
1332 """What a claude reviewer argv still offers the model, in words; ``[]`` if nothing.
1334 Every tool a ``--tools`` list names counts, the denied ones included. The
1335 deny list is meant to remove those again, but it is the second layer, and a
1336 configuration that asks for ``Read`` or ``WebFetch`` on a reviewer is one the
1337 operator should hear about rather than one this module quietly trusts the
1338 CLI's precedence rules to undo. MCP servers count when ``--mcp-config`` loads
1339 some, or when ``--strict-mcp-config`` is missing and the operator's own
1340 Claude configuration would.
1341 """
1342 args = list(extra_args)
1343 surface: list[str] = []
1344 tools = _claude_tools(args)
1345 if tools is None: # pragma: no cover - enforcement injects `--tools ""`
1346 surface.append("the CLI's default tool set (no `--tools`)")
1347 elif tools:
1348 # Built-in names from a fixed list; others are not repeated.
1349 names = [t if t in _SHOWN_TOOLS else "<tool>" for t in dict.fromkeys(tools)]
1350 surface.append(f"`--tools {','.join(names)}`")
1351 if _claude_flag_present("--mcp-config", args):
1352 surface.append("MCP servers from `--mcp-config`")
1353 if not _claude_flag_present("--strict-mcp-config", args): # pragma: no cover - injected
1354 surface.append("the MCP servers in the user's Claude configuration")
1355 return surface
1358def _claude_rejected_mode(extra_args: list[str]) -> str | None:
1359 """A ``--permission-mode`` Claude Code rejects, as it should be shown; or None.
1361 The first decision on every audit path that talks about permissions: a
1362 rejected mode makes Claude Code exit before the model is reached, whatever
1363 else the argv holds — ``--dangerously-skip-permissions`` and ``--tools``
1364 included (measured on 2.1.236 with an invalid model name: ``bogus``, a
1365 missing value, an empty ``=`` and a value that is itself a flag, such as
1366 ``--permission-mode --dangerously-skip-permissions``, were each refused). So
1367 such a seat never runs, in bypass mode or any other. An empty value is shown
1368 as ``--permission-mode=`` and a missing one as ``--permission-mode``.
1369 """
1370 args = _claude_options(extra_args)
1371 values = _claude_value_positions(args)
1372 accepted = (*_CLAUDE_KNOWN_MODES, "default")
1373 for i in range(len(args)):
1374 mode = None if i in values else _permission_mode_at(args, i)
1375 if mode is None or mode[0] in accepted:
1376 continue
1377 if mode[0]:
1378 return f"--permission-mode {_shown_value(mode[0])}"
1379 return (
1380 "--permission-mode="
1381 if args[i].startswith("--permission-mode=")
1382 else ("--permission-mode")
1383 )
1384 return None
1387def _claude_effective_mode(extra_args: list[str]) -> str | None:
1388 """The ``--permission-mode`` Claude Code applies, or ``None`` when none is named.
1390 The **last** one at a flag position (#888). Measured on Claude Code 2.1.236:
1391 ``plan`` then ``dontAsk`` runs as ``dontAsk``, and ``dontAsk`` then ``plan``
1392 runs as ``plan``. Only meaningful once :func:`_claude_rejected_mode` has
1393 answered ``None`` — any rejected value stops the CLI, wherever it sits — and
1394 it does not account for ``--dangerously-skip-permissions``, which overrides
1395 whichever mode this returns, in either order.
1396 """
1397 args = _claude_options(extra_args)
1398 values = _claude_value_positions(args)
1399 named = [
1400 m[0] for i in range(len(args)) if i not in values and (m := _permission_mode_at(args, i))
1401 ]
1402 return named[-1] if named else None
1405def _claude_mode_override(extra_args: list[str]) -> str | None:
1406 """A permission setting other than the reviewer's ``dontAsk``, as written; or None.
1408 ``--dangerously-skip-permissions``, or a ``--permission-mode`` naming any other
1409 mode — one Claude Code accepts (``bypassPermissions``, ``auto``,
1410 ``acceptEdits``, ``manual`` or its hidden alias ``default``, ``plan``) or one
1411 it rejects (a value outside that set, an empty value, no value at all). Read
1412 at flag positions only.
1414 A rejected mode comes first, ahead of ``--dangerously-skip-permissions``:
1415 Claude Code 2.1.236 refuses to start on it whatever else is in the argv
1416 (measured with the skip flag beside ``bogus`` and beside a missing value), so
1417 the seat fails rather than running in bypass mode, and the warning must say so.
1418 An empty value is shown as ``--permission-mode=`` and a missing one as
1419 ``--permission-mode``, so neither renders as a flag with a blank after it.
1421 Of several accepted modes, the one reported is the one the seat runs in: the
1422 last (:func:`_claude_effective_mode`, #888). Reporting the first one that was
1423 not ``dontAsk`` told an operator whose ``plan`` was followed by ``dontAsk``
1424 that ``plan`` was in use, while the seat ran ``dontAsk``.
1425 """
1426 args = list(extra_args)
1427 rejected = _claude_rejected_mode(args)
1428 if rejected is not None:
1429 return rejected
1430 if _claude_flag_present("--dangerously-skip-permissions", args):
1431 return "--dangerously-skip-permissions"
1432 mode = _claude_effective_mode(args)
1433 if mode is not None and mode != _CLAUDE_REVIEW_MODE:
1434 return f"--permission-mode {mode}"
1435 return None
1438#: What each permission setting other than ``dontAsk`` does to a reviewer, as
1439#: ``(what it does, whether it approves tool calls unasked)``. Measured on Claude
1440#: Code 2.1.236: with ``--permission-mode dontAsk`` and ``Read`` available, a
1441#: ``Read`` outside the working directory was denied; adding
1442#: ``--dangerously-skip-permissions`` before or after it let the read through.
1443#: The flag also overrides a named ``plan`` or ``auto`` (the seat reports
1444#: ``bypassPermissions``). A named ``--permission-mode`` is otherwise kept as
1445#: written (nothing is injected beside it), so it is the mode the seat runs in —
1446#: the last one, when several are named (#888).
1447#: ``default`` is not in Claude Code's listed choices but is accepted as an alias
1448#: of ``manual``: its init event reports ``"permissionMode":"default"`` for both.
1449_CLAUDE_MODE_EFFECTS: dict[str, tuple[str, bool]] = {
1450 # The first half of this entry names the mode actually beside the flag; see
1451 # `_claude_skip_overrides`.
1452 "--dangerously-skip-permissions": (
1453 "so the seat runs in bypass mode, which approves every tool call without asking",
1454 True,
1455 ),
1456 "bypassPermissions": (
1457 "is used instead of the reviewer's `dontAsk`, and approves every tool call without asking",
1458 True,
1459 ),
1460 "auto": (
1461 "is used instead of the reviewer's `dontAsk`, and lets Claude Code approve "
1462 "tool calls without asking you",
1463 True,
1464 ),
1465 "acceptEdits": (
1466 "is used instead of the reviewer's `dontAsk`, and approves file edits without asking",
1467 True,
1468 ),
1469 "manual": ("is used instead of the reviewer's `dontAsk`", False),
1470 "default": ("is used instead of the reviewer's `dontAsk`", False),
1471 "plan": ("is used instead of the reviewer's `dontAsk`", False),
1472}
1474#: The ``--permission-mode`` choices Claude Code 2.1.236 lists; it also accepts
1475#: ``default`` (an unlisted alias of ``manual``). Anything else, an empty value
1476#: or no value makes it exit before the model is reached.
1477_CLAUDE_KNOWN_MODES: tuple[str, ...] = (
1478 "acceptEdits",
1479 "auto",
1480 "bypassPermissions",
1481 "manual",
1482 "dontAsk",
1483 "plan",
1484)
1487def _claude_skip_overrides(extra_args: list[str]) -> str:
1488 """What ``--dangerously-skip-permissions`` overrides in *extra_args*, in words.
1490 ``dontAsk`` is injected only when the argv names no ``--permission-mode``, so
1491 beside a named ``plan`` or ``auto`` there is no ``dontAsk`` to override — the
1492 sentence names the mode that is actually there.
1493 """
1494 args = _claude_options(extra_args)
1495 values = _claude_value_positions(args)
1496 modes = [
1497 m[0] for i in range(len(args)) if i not in values and (m := _permission_mode_at(args, i))
1498 ]
1499 measured = (
1500 "measured on Claude Code 2.1.236 for `dontAsk` in either order, and for `plan` and `auto`"
1501 )
1502 if not modes: # pragma: no cover - enforcement injects `--permission-mode dontAsk`
1503 return f"overrides any `--permission-mode` ({measured})"
1504 named = ", ".join(f"`--permission-mode {_shown_value(m)}`" for m in dict.fromkeys(modes))
1505 return f"overrides the {named} beside it ({measured})"
1508def _claude_mode_warning(label: str, override: str, extra_args: list[str]) -> str:
1509 """The audit's sentence for *override*, saying what that setting really does."""
1510 key = override.split(" ", 1)[1] if override.startswith("--permission-mode ") else override
1511 head = f"agent '{label}' (claude) is configured with `{override}`, which "
1512 # A rejected mode first: `--permission-mode --dangerously-skip-permissions`
1513 # names the skip flag as its VALUE, and must not be read as the flag itself.
1514 if override == _claude_rejected_mode(extra_args) or key not in _CLAUDE_MODE_EFFECTS:
1515 if override == "--permission-mode":
1516 head = f"agent '{label}' (claude) is configured with `--permission-mode` and no value, which "
1517 elif override == "--permission-mode=":
1518 head = (
1519 f"agent '{label}' (claude) is configured with an empty `--permission-mode=`, which "
1520 )
1521 return (
1522 f"{head}Claude Code 2.1.236 rejects (it accepts "
1523 f"{', '.join(_CLAUDE_KNOWN_MODES)}, and `default` as an alias of `manual`), "
1524 f"so the seat fails before it reviews anything. Drop it — a reviewer runs "
1525 f"in `dontAsk`."
1526 )
1527 effect, approves = _CLAUDE_MODE_EFFECTS[key]
1528 if key == "--dangerously-skip-permissions":
1529 effect = f"{_claude_skip_overrides(extra_args)}, {effect}"
1530 if _claude_tools(list(extra_args)):
1531 tail = "The seat is also given tools (see above), so this applies to them. Drop it."
1532 elif approves:
1533 tail = (
1534 'With `--tools ""` in force there is no tool for it to approve, so it grants '
1535 "nothing today, but it would approve whatever a later flag or Claude Code "
1536 "release made available. Drop it."
1537 )
1538 else:
1539 tail = (
1540 'With `--tools ""` in force there is no tool for it to act on, so it changes '
1541 "nothing today, but a reviewer runs in `dontAsk`. Drop it."
1542 )
1543 return f"{head}{effect}. {tail}"
1546def _claude_config_options(extra_args: list[str]) -> list[str]:
1547 """The :data:`_CLAUDE_CONFIG_OPTIONS` present at flag positions, in that order."""
1548 args = list(extra_args)
1549 return [f for f in _CLAUDE_CONFIG_OPTIONS if _claude_flag_present(f, args)]
1552def _claude_permission_bypass(extra_args: list[str]) -> str | None:
1553 """The token that makes claude approve tool calls unasked, or ``None``.
1555 ``None`` too when the argv carries a rejected ``--permission-mode``: Claude
1556 Code refuses to start on it, so there is no seat to approve anything, and
1557 :func:`_claude_mode_warning` reports the rejection instead.
1558 """
1559 args = list(extra_args)
1560 if _claude_rejected_mode(args) is not None:
1561 return None
1562 if _claude_flag_present("--dangerously-skip-permissions", args):
1563 return "--dangerously-skip-permissions"
1564 # The mode the seat runs in, not any mode written (#888): an approving mode
1565 # followed by `dontAsk` runs as `dontAsk`, and approves nothing.
1566 mode = _claude_effective_mode(args)
1567 if mode in _CLAUDE_APPROVING_MODES:
1568 return f"--permission-mode {mode}"
1569 return None
1572#: Bring-your-own CLIs that obey configuration in the directory they run in (#859).
1573#: A `cli` seat runs in the directory `jury` was started from — on a pull request,
1574#: the author's checkout — so that configuration is the author's, and no flag in
1575#: the seat turns it off. Keyed by the command's file name.
1576#:
1577#: - aider reads ``.aider.conf.yml`` from its working directory and git root and
1578#: ``.env`` from both (aider ``main.py``, ``main`` 5dc9490); those can set
1579#: ``test``/``lint`` with a command, or ``load``, which it runs before
1580#: ``--message``. ``--test``/``--lint`` have no ``--no-`` form, and ``--config``
1581#: / ``--env-file`` add a file instead of replacing the ones found
1582#: (ConfigArgParse ``_open_config_files`` always opens the defaults).
1583#: - cursor-agent: Cursor documents that project hooks (``.cursor/hooks.json``)
1584#: "run in any trusted workspace", and that its ``workspaceOpen`` hook "Runs in
1585#: the Cursor desktop app and CLI". A headless run starts only in a trusted
1586#: directory (``--trust``, or ``--force``/``--yolo``).
1587_CHECKOUT_CONFIG_RISKS: dict[str, str] = {
1588 "aider": (
1589 "reads .aider.conf.yml and .env from the directory it runs in, which can "
1590 "turn on its test, lint or load commands and run them before the review, "
1591 "whatever its flags say"
1592 ),
1593 "cursor-agent": (
1594 "runs the hooks in .cursor/hooks.json of the directory it runs in once "
1595 "that workspace is trusted (--trust), whatever its flags say"
1596 ),
1597}
1600def _checkout_config_warning(label: str, command: str) -> str:
1601 """The sentence for a CLI that obeys its working directory's config (#859).
1603 Empty for any other command. Keyed by the binary, not the adapter: no argv
1604 sandbox stops either CLI reading its checkout's config, so the risk is
1605 reported whatever the seat's sandbox status (#901). Appended to the seat's
1606 one warning where it has one, so a seat still gets one warning.
1607 """
1608 # The stem, case-insensitively (#901 review): `cursor-agent.cmd`, `.bat`,
1609 # `.ps1` and `AIDER.EXE` are the same CLI as the bare name.
1610 # `PureWindowsPath` splits on both slash and backslash, so one reading serves both.
1611 name = PureWindowsPath(command).stem.lower()
1612 risk = _CHECKOUT_CONFIG_RISKS.get(name)
1613 if risk is None:
1614 return ""
1615 return (
1616 f" Also, {name} {risk}. It runs in the directory jury was started from — on "
1617 f"a pull request, the author's checkout — so seat '{label}' only on "
1618 f"checkouts you trust."
1619 )
1622def _checkout_only_warning(label: str, checkout: str) -> str:
1623 """The checkout-config warning on its own, for a seat with no other one (#901).
1625 Used where the audit would otherwise accept the seat — a sandbox flag it
1626 recognizes, or the claude lockdown — because neither stops aider or
1627 cursor-agent obeying the config in the checkout it runs in. Reporting it
1628 is what makes ``--strict`` fail such a seat.
1629 """
1630 return (
1631 f"agent '{label}' is confined only in what its flags control; its CLI's "
1632 f"checkout config is not.{checkout}"
1633 )
1636def audit_agent(spec) -> list[str]:
1637 """Return least-privilege warnings for a single agent spec, redacted.
1639 Each sentence is built to repeat no argv value outside :data:`_SHOWN_VALUES`
1640 (#908 review, round 5); passing it through :func:`redaction.redact` as well
1641 catches any secret-shaped text a future sentence lets through.
1642 """
1643 return [redact(w)[0] for w in _audit_agent(spec)]
1646def _audit_agent(spec) -> list[str]:
1647 """The warnings :func:`audit_agent` returns, before redaction.
1649 The subject is the **effective** argv — what ``adapters._read_only_extra_args``
1650 will spawn this seat with — not the ``extra_args`` written in the config
1651 (issue #750). Every panel invocation is routed through
1652 :func:`enforce_read_only`, so the declared list is only half the command
1653 line: a bare ``claude`` seat with no ``extra_args`` at all — the documented,
1654 recommended configuration — is spawned with the full write-tool denylist and
1655 was nevertheless reported as write-capable, failing ``--strict`` on the one
1656 configuration the docs tell operators to write. An audit that cries wolf on
1657 the recommended setup is one operators learn to pass ``--no-strict`` around.
1659 Auditing the enforced argv is also what keeps the check honest where
1660 enforcement cannot help, with no special case for it: :func:`enforce_read_only`
1661 is a **no-op** for the bring-your-own-CLI vendors (``cli``/``xai``), so for
1662 those the effective argv *is* the declared one and every warning that fired
1663 before still fires. The same holds for a sandbox the operator widened on
1664 purpose, which enforcement keeps as written rather than narrowing.
1666 What this can and cannot promise: the enforcement is only as good as the
1667 adapter that applies it. Every adapter here that spawns a subprocess builds
1668 its argv through ``adapters._read_only_extra_args`` — claude, codex, agy, and
1669 the generic CLI adapter an unknown vendor falls through to — so for those the
1670 audited argv is the real one. A custom adapter registered through
1671 ``adapters.register_adapter`` is operator-supplied code that may build its
1672 argv however it likes; for one of those this reports what enforcement *would*
1673 produce, which is also why a bare ``--sandbox`` is still not trusted from a
1674 vendor whose CLI is unknown (issue #292).
1675 """
1676 warnings: list[str] = []
1677 # The ADAPTER, not the vendor (issue #705). Every question this function
1678 # asks — is there a subprocess, which sandbox flag would confine it, is one
1679 # present — is about the CLI that gets spawned. A seat with
1680 # `vendor = "openai", adapter = "cli"` spawns the operator's own binary, for
1681 # which this tool knows no sandbox flag; auditing it as codex would demand a
1682 # `-s read-only` that `cursor-agent` does not have.
1683 vendor = spec_adapter(spec)
1684 label = getattr(spec, "name", "agent")
1686 # Local/HTTP agents (issue #43) and hosted-API agents (issue #430) run no
1687 # subprocess to sandbox — there is no write/tool/network surface to flag
1688 # (a hosted-API call has strictly less access than even a sandboxed CLI:
1689 # no filesystem, no shell, nothing to disallow), so they are out of scope
1690 # for this audit. Which seats those are is asked of the adapter the spawner
1691 # builds, not of the seat's keys (#901 review): the presence of `endpoint`
1692 # skipped the audit, while `make_adapter` ignores `endpoint` on every
1693 # registered CLI adapter — a `cli` aider seat, a claude seat with
1694 # `--dangerously-skip-permissions`, a codex seat with `danger-full-access`,
1695 # each with a stray endpoint, ran as CLIs and were audited clean.
1696 if not spawns_process(spec):
1697 return warnings
1699 # The argv this seat is spawned with, byte for byte what
1700 # `adapters._read_only_extra_args(spec)` returns: same adapter key, same
1701 # declared args, same function. The seat's name is an input to neither
1702 # (#758, #768) — which is what stops the auditor and the spawner disagreeing
1703 # about a seat whose name says one CLI and whose adapter key says another.
1704 declared = list(getattr(spec, "extra_args", []) or [])
1705 extra_args = enforce_read_only(vendor, declared)
1706 args_text = _args_str(extra_args)
1707 # A CLI that obeys its working directory's config (aider, cursor-agent) is
1708 # told so on every path below, sandboxed or not (#859, #901): no flag in the
1709 # argv stops it reading the checkout's config.
1710 checkout = _checkout_config_warning(label, str(getattr(spec, "command", "") or ""))
1712 is_claude = vendor == "anthropic"
1714 if is_claude:
1715 # A tripwire, not a live check: `_ensure_claude_disallowed` merges the denied
1716 # tools into the argv this function just enforced, so a locked-down result is
1717 # guaranteed and the body below cannot run. It is kept because the guarantee
1718 # lives in another function — if enforcement ever stops injecting, this is
1719 # what says so instead of the audit silently passing a writable seat. Before
1720 # #758 removed the name-based identity it was reachable, via a `cli`-adapter
1721 # seat whose *name* contained "claude".
1722 if not _claude_is_locked_down(extra_args): # pragma: no cover - see above
1723 warnings.append(
1724 f"agent '{label}' (claude) is not restricted to read-only: add "
1725 f"`--disallowed-tools {','.join(_CLAUDE_DENIED_TOOLS)}` so a prompt "
1726 f"injection in the diff cannot edit files, run commands, read "
1727 f"files outside the diff or reach the network."
1728 )
1729 # What enforcement keeps as written: a `--tools` list, or MCP servers the
1730 # operator loaded. The shipped default has neither, so it raises nothing.
1731 surface = _claude_open_surface(extra_args)
1732 bypass = _claude_permission_bypass(extra_args)
1733 if surface:
1734 named = ", ".join(surface)
1735 if bypass:
1736 warnings.append(
1737 f"agent '{label}' (claude) is given {named} and skips permission "
1738 f"checks (`{bypass}`), so a prompt injection in the diff can use "
1739 f"them unasked — to read files outside the diff or reach the "
1740 f"network. Drop them — a reviewer only reads its prompt."
1741 )
1742 elif _claude_rejected_mode(extra_args) is not None:
1743 # The seat never starts: Claude Code rejects its `--permission-mode`
1744 # (reported below). Said conditionally, then (#888 follow-up), so
1745 # this does not read as a live exposure beside that sentence.
1746 warnings.append(
1747 f"agent '{label}' (claude) is given {named} while reviewing "
1748 f"untrusted content; once the seat can start, a prompt injection "
1749 f"in the diff could ask for them, and whatever your Claude "
1750 f"settings pre-approve would run unasked. Drop them — a reviewer "
1751 f"only reads its prompt."
1752 )
1753 else:
1754 warnings.append(
1755 f"agent '{label}' (claude) is given {named} while reviewing "
1756 f"untrusted content; a prompt injection in the diff can ask for "
1757 f"them, and whatever your Claude settings pre-approve runs "
1758 f"unasked. Drop them — a reviewer only reads its prompt."
1759 )
1760 # A permission setting other than dontAsk, kept as written. Said once:
1761 # the warning above already named an approving one next to its tools.
1762 override = _claude_mode_override(extra_args)
1763 if override and not (surface and bypass):
1764 warnings.append(_claude_mode_warning(label, override, extra_args))
1765 # Configuration beyond the prompt. `--safe-mode` keeps it from loading
1766 # (measured), so this is a surprise to prevent, not a hole to close.
1767 # Fail-closed (#908 review): a risky token the readers above did not read
1768 # as a flag may still be one Claude applies.
1769 unread = _claude_unread_risks(extra_args)
1770 if unread:
1771 named = _shown_positions(unread, len(extra_args) - len(declared))
1772 one = len(unread) == 1
1773 warnings.append(
1774 f"agent '{label}' (claude) has {named} where jury cannot tell whether "
1775 f"Claude reads {'it' if one else 'them'} as {'a flag' if one else 'flags'} "
1776 f"(after `--`, as another option's value, or in a spelling jury does not "
1777 f"model). If Claude does, the reviewer's permission mode or tools "
1778 f"change, and jury has not checked how. Drop "
1779 f"{'it' if one else 'them'} — a reviewer only reads its prompt."
1780 )
1781 loaded = _claude_config_options(extra_args)
1782 if loaded:
1783 named = ", ".join(f"`{f}`" for f in loaded)
1784 warnings.append(
1785 f"agent '{label}' (claude) is configured with {named}; `--safe-mode` "
1786 f"keeps them from loading hooks, CLAUDE.md, plugins or agents into a "
1787 f"reviewer, so they have no effect there. A reviewer needs none of "
1788 f"them — drop them."
1789 )
1790 if checkout:
1791 warnings.append(_checkout_only_warning(label, checkout))
1792 return warnings
1794 # agy is spawned with `--sandbox` (enforced above), and that is still not
1795 # confinement: measured on agy 1.2.9 the shipped argv read and wrote files
1796 # outside its working directory and reached the network, and agy has no
1797 # flag that removes its tools. It is out of the default panel for that
1798 # reason, so a seat on this adapter is one the operator configured — and is
1799 # told, whatever flags it carries. The implementer role never reaches this:
1800 # the audit covers panel seats only.
1801 if vendor == "google":
1802 warnings.append(
1803 f"agent '{label}' (agy) cannot be confined: even with --sandbox it reads "
1804 f"and writes files and reaches the network; do not use it on untrusted "
1805 f"diffs."
1806 )
1807 # A `--sandbox=<value>` that would switch the sandbox off (#908 review).
1808 # Enforcement removed it, so it is read from the declared args.
1809 switched = _agy_sandbox_switched_off(declared)
1810 off = [shown for shown, is_false in switched if is_false]
1811 unread = [shown for shown, is_false in switched if not is_false]
1812 for tokens, effect in (
1813 (
1814 off,
1815 "would turn agy's sandbox off: agy reads a `--sandbox=` value as true "
1816 "or false, and the last `--sandbox` it reads wins",
1817 ),
1818 (unread, "is not a value agy reads as true or false, so agy refuses to start"),
1819 ):
1820 if tokens:
1821 named = ", ".join(f"`{t}`" for t in tokens)
1822 one = len(tokens) == 1
1823 warnings.append(
1824 f"agent '{label}' (agy) is configured with {named}, which {effect}. "
1825 f"jury removes {'it' if one else 'them'} and spawns the seat with "
1826 f"`--sandbox`. Drop {'it' if one else 'them'}."
1827 )
1828 # codex's sandbox spelling on an agy seat (#902). Enforcement added agy's
1829 # own `--sandbox` beside it; what the operator wrote does nothing agy
1830 # documents, and may keep the seat from starting.
1831 foreign = _agy_foreign_sandbox_tokens(extra_args)
1832 if foreign:
1833 named = ", ".join(f"`{t}`" for t in foreign)
1834 warnings.append(
1835 f"agent '{label}' (agy) is configured with {named}, which is not agy's "
1836 f"sandbox: agy 1.2.12 has only a boolean `--sandbox` (jury adds it) "
1837 f"and no `-s`, so agy may refuse the argv or read a value as prompt "
1838 f"text. Drop {'it' if len(foreign) == 1 else 'them'}."
1839 )
1840 # A sandbox token agy reads as an option's value or as prompt text (#910).
1841 # Enforcement added the real flag unless one was already read.
1842 not_read = _agy_sandbox_not_read(declared)
1843 if not_read:
1844 one = len(not_read) == 1
1845 warnings.append(
1846 f"agent '{label}' (agy) is configured with {', '.join(not_read)}, so "
1847 f"agy does not read {'it' if one else 'them'} as its sandbox flag. The "
1848 f"seat's sandbox is a `--sandbox` agy does read, which jury adds when "
1849 f"there is none. Drop {'it' if one else 'them'}."
1850 )
1851 elif _agy_enforced_as_unknown(vendor):
1852 # An unknown vendor gets agy's `--sandbox` handling, command or not (#910):
1853 # every `--sandbox=<value>` is removed, whatever it means to this CLI.
1854 removed = [_shown_flag(declared[i]) for i in _agy_sandbox_values(declared)]
1855 if removed:
1856 named = ", ".join(f"`{t}`" for t in removed)
1857 one = len(removed) == 1
1858 warnings.append(
1859 f"agent '{label}' is configured with {named}, which jury removes: a "
1860 f"vendor jury does not know gets agy's sandbox handling, which drops "
1861 f"every `--sandbox=<value>` and adds a bare `--sandbox`, whatever the "
1862 f"value means to this CLI. Drop {'it' if one else 'them'}, or set "
1863 f'`adapter = "cli"` to pass the argv through as written.'
1864 )
1866 # Non-claude agents must run under a restricting sandbox (issue #100) — one
1867 # the config named, or one enforcement injected above.
1868 if _is_sandboxed(extra_args, vendor=vendor):
1869 # …unless a sandbox-SELECTING flag is sitting in the same argv, in which
1870 # case two sandboxes are specified and the CLI, not this module, decides
1871 # which one wins (issue #750). Said plainly rather than through the
1872 # generic message below, which would recommend the `-s read-only` that
1873 # is already there.
1874 competing = _competing_sandboxes(extra_args, vendor=vendor)
1875 if _is_codex(vendor):
1876 unread = _codex_unread_risks(extra_args)
1877 if unread:
1878 named = _shown_positions(unread, len(extra_args) - len(declared))
1879 warnings.append(
1880 f"agent '{label}' has {named} where jury cannot tell whether codex "
1881 f"reads it as a sandbox setting (after `--`, inside a `-c` "
1882 f"override, or in a spelling jury does not model). If codex does, "
1883 f"the enforced read-only sandbox may be widened or removed. Drop it."
1884 )
1885 if competing:
1886 named = ", ".join(f"`{token}`" for token, _ in competing)
1887 one = len(competing) == 1
1888 # Say which failure mode it is. `--full-auto` picks a write-capable
1889 # sandbox; `--yolo` removes the sandbox altogether. Telling an operator
1890 # that a bypass merely "selects a sandbox of its own" describes the
1891 # milder of the two and understates what they configured.
1892 # `any`, not `all`: a seat carrying `--full-auto` *and* `--yolo` must be
1893 # described by the worse of the two, not the milder.
1894 if any(disables for _, disables in competing):
1895 effect = (
1896 "disables the sandbox entirely"
1897 if one
1898 else "include a flag that disables the sandbox entirely"
1899 )
1900 tail = (
1901 f"the enforced read-only sandbox may not apply at all. "
1902 f"Drop {'it' if one else 'them'} — a reviewer only reads its prompt."
1903 )
1904 else:
1905 effect = (
1906 "selects a sandbox of its own" if one else "each state a sandbox of their own"
1907 )
1908 tail = (
1909 f"{'it is' if one else 'they are'} passed alongside the enforced "
1910 f"read-only sandbox, so which {'of the two applies' if one else 'applies'} "
1911 f"is up to the CLI. Drop {'it' if one else 'them'} — a reviewer only "
1912 f"reads its prompt."
1913 )
1914 warnings.append(f"agent '{label}' is configured with {named}, which {effect}; {tail}")
1915 if checkout:
1916 warnings.append(_checkout_only_warning(label, checkout))
1917 return warnings
1918 # Not sandboxed. A broad-powers flag gets a specific message…
1919 for flag in _DANGEROUS_FLAGS:
1920 if flag in extra_args or flag in args_text:
1921 warnings.append(
1922 f"agent '{label}' is configured with `{flag}`, granting "
1923 f"write/tool/network powers while reviewing untrusted content; "
1924 f"prefer a read-only sandbox (e.g. codex `-s read-only` or agy "
1925 f"`--sandbox`).{checkout}"
1926 )
1927 return warnings
1928 # …otherwise warn that it simply isn't sandboxed. This closes the audit
1929 # blind spot (issue #300): an unknown-vendor or no-flag agent previously
1930 # produced ZERO warnings and ran via the generic adapter without the
1931 # read-only guarantee — and so `--strict` could not fail it on this basis.
1932 # A bring-your-own CLI (#901 item 1): jury adds and checks no sandbox for it,
1933 # so telling its operator there is "no `--sandbox`" is wrong for a cursor-agent
1934 # seat that passes its own `--sandbox enabled`, and "add a sandbox" names a fix
1935 # jury could not verify either way.
1936 if vendor in GENERIC_CLI_VENDORS:
1937 warnings.append(
1938 f"agent '{label}' is not under a sandbox jury can verify: jury adds or "
1939 f"checks none for a bring-your-own CLI, so its own flags are the whole "
1940 f"restriction, and a prompt injection in the diff could reach "
1941 f"write/tool/network if they allow it. Run with `--strict` to fail.{checkout}"
1942 )
1943 return warnings
1944 warnings.append(
1945 f"agent '{label}' is not running under a recognized read-only sandbox "
1946 f"({_missing_sandbox(vendor)}); a prompt injection in the diff could "
1947 f"reach write/tool/network. Add a sandbox, or run with `--strict` to fail.{checkout}"
1948 )
1949 return warnings
1952def _missing_sandbox(vendor: str) -> str:
1953 """What the generic warning says is missing, in the spawned CLI's own terms (#902).
1955 One sentence for every CLI named both codex's and agy's flag, which read
1956 wrongly on an agy seat carrying codex's ``-s read-only``, and told a seat
1957 whose ``--sandbox`` jury had injected that it had "no ``--sandbox``".
1958 """
1959 if _is_codex(vendor):
1960 return "codex needs `-s read-only`"
1961 if normalise_vendor(vendor) == "google":
1962 return "agy needs its boolean `--sandbox`, followed by another flag or nothing"
1963 return (
1964 "jury knows no sandbox flag for this CLI; the `--sandbox` it adds is agy's, "
1965 "which this CLI may ignore"
1966 )
1969def audit_privilege(specs) -> list[str]:
1970 """Return all least-privilege warnings across the configured agents."""
1971 warnings: list[str] = []
1972 for spec in specs:
1973 warnings.extend(audit_agent(spec))
1974 return warnings