Files
claude-projects/agent-builder/constrained_autonomy_design_decisions.md
T
Backtalk6858 6cffad860d docs: CA-P1b — widen fast-path on logged evidence (hook v2.1)
Records why the v2.0 classifier taxed the user 22 prompts and what changed.

Root cause: v2.0 conflated "contains a redirect character" with "writes
something". `2>&1` is file-descriptor plumbing and `>/dev/null` is a discard —
neither writes. That single mistake caused 16 of the 22 prompts. It was found
by reading ca_decisions.jsonl, not by inspection: the bug was the classifier's
reasoning, so re-reading the code only reproduced it.

v2.1 changes (all evidence-driven, none speculative):
- _SAFE_REDIRECT_RE strips fd plumbing + /dev/null discards before the
  redirect check; a redirect to any REAL path still disqualifies, anchored so
  `>/dev/nullx` cannot ride the prefix.
- `cd` added to READONLY_ALLOW (no filesystem effect; every other segment must
  independently qualify anyway).
- `ssh <host> '<cmd>'` classifies the inner command under identical rules,
  depth-limited to one hop.

Verified: 42/42 self-test, e2e ALL PASS, classify() median 0.011ms.
Tier-1 writes deliberately NOT shipped — widen on evidence, not guesses.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-14 19:33:39 -05:00

16 KiB
Raw Blame History

Constrained Autonomy for Claude Code — Design Decisions

Status: DESIGN LOCKED (grill-me 2026-07-14). Not yet built. Build = phased background agents (CA-P0…P5), mirroring the Agent-Sudo build model. Do NOT implement ad hoc — follow this record.

Driving insight (user): a permission prompt the human rubber-stamps — especially a truncated command they can't read or don't understand — adds no safety. It is the same theater as the NTFY approval gate we removed from Agent-Sudo. Kill the theater; replace it with an ENFORCED code constraint layer. Governing memory: feedback_autonomous_security_constrain_not_gate (constrain capability, don't gate access), feedback_evolution_by_default, feedback_always_include_training_loop, feedback_secrets_via_proxy_only.

The core problem

Claude Code's Bash tool runs as administrator (docker group + sudo). Today, consequential commands stop for a human permission prompt. --dangerously-skip-permissions removes that gate but adds NO constraint → strictly worse. We want: remove the theater prompt, but move the real gate into deterministic enforced code that reuses the Agent-Sudo constraint+recovery stack.

D1 — Enforcement architecture: HYBRID (hook classifies+routes; daemon executes-with-constraint)

  • The hook = security-enforcement.py (existing PreToolUse Bash hook, pure-regex, <50ms, no-I/O) extended into the decision point. It runs regardless of permission mode and is code, not a rubber stamp. It emits a permissionDecision:
    • tier-0/1 (read/reversible)allow LOCALLY (no prompt, no daemon round-trip → kills the theater, stays <50ms).
    • tier-2/3 (mutating/destructive) + UNKNOWNdeny the raw command, message routes it to the Agent-Sudo daemon /exec (which already does sandbox-test-first + rollback + audit). This generalizes what the hook ALREADY does for sudo today (block → route to bridge) into "anything consequential → Agent-Sudo".
    • tier-4 (catastrophic)deny, no override.
  • The daemon = Agent-Sudo /exec (LIVE: server-01 8082, primary 8084). All heavy lifting (Incus sandbox mirror test, Timeshift/rollback, circuit-breaker, command_audit logging) stays where it's built and tested. We reuse it; we do not reimplement it in a 50ms regex hook.
  • Rejected: pure-hook (can't sandbox/rollback in 50ms), route-EVERYTHING-through-daemon (latency + hard dependency + wrong root context for harmless reads).

D2 — Classification: STRUCTURAL effect-classifier, unknown→sandbox, learned allowlist on top

  • SUDO.md only covers the sudo surface; ~90% of Claude Code commands are non-sudo (git/docker/python/curl/file ops). So the classifier judges a command by its effect class, not exact-match:
    • reads/inspection → tier-0/1 allow
    • workspace-local / reversible writes (git add/commit, edits under project dirs, scratch scripts) → tier-1 allow
    • blast-radius mutations (docker rm/restart/compose down, writes outside workspace, DB writes, net-config, mv/rm on real paths) → tier-2/3 → daemon sandbox
    • catastrophic verbs (rm -rf / patterns, mkfs, dd of=/dev/…, fork bombs, disk/partition ops) → tier-4 deny
    • sudo-prefixed → daemon's SUDO.md governs (unchanged)
    • genuinely unrecognizable structure → default to daemon sandbox (treat as tier-2) — unknown = "prove it in the mirror first," NEVER "ask the tired human." (User-confirmed.)
  • Learned allowlist layered ON TOP: commands that ran clean N times (from command_audit + human_verdict gold labels) get promoted to the local fast-path → system gets faster with use (feedback_evolution_by_default). Starts from structural analysis so day-one has zero rubber-stamp gaps.
  • Policy source of truth = a signed CLAUDE_CMDS.md tier-map file (the effect-class rules), analogous to SUDO.md.

D3 — Secrets: via secrets-proxy env_secrets (dormant until #128)

  • Any command that touches a secret runs via secrets-proxy /shell + env_secrets (vault://path#field / bitwarden://item) so values are sourced at runtime, never in stdout/history/argv. The hook's existing Category-2/3 blocks (secret exposure, hardcoded secrets) stay as the backstop that FORCES commands onto the proxy path.
  • Dependency: secrets-proxy = task #128 (pending/down). Secret-handling path is built but dormant until #128; interim fallback = feedback_secret_to_script_via_file. Decoupled from the classifier build.

D4 — Tamper-evidence: Vault transit signature — UNBLOCKED 2026-07-14 (#145 resolved)

Transit is LIVE. Vault admin WAS recovered (user was right; my first probe was wrong — it checked only the least-privilege AppRole, which is denied transit by design, and never followed the Bitwarden path). Chain: vault approle → secret/data/bitwarden-bridge#BRIDGE_API_KEY → bridge :8083 GET /secret?item=Hashicorp Vault&field=notes → root token. GOTCHA: the token LABELED "Root token:" is REVOKED; a second UNLABELED token at the bottom of the note is the live one. Search the whole note — do not trust the label. transit enabled + ed25519 key agent-sudo-sudomd created + SUDO.md SIGNED (sha256 38eca778…, transit/verify PASSED) + tamper test PASSED (fake widened rule → gate REFUSED). CLAUDE_CMDS.md (CA-P4) and proxy.md (#150) can now be signed the same way.

D4 (original) — Vault transit signature (was dormant until #145)

  • CLAUDE_CMDS.md gets the IDENTICAL mechanism as sudo_sign.py: Vault transit ed25519 key, sign file sha256, verify_gate on load via scoped AppRole, ..._VERIFY_ENFORCE flag. Tampered policy → refuse to honor.
  • Dependency: live signing blocked on #145 (OpenBao admin lockout; transit engine not enableable yet). SUDO.md itself runs unsigned (enforce=false) for this reason. CLAUDE_CMDS.md inherits the posture: built to verify, ships enforce=false, flips to true when #145 lands — one activation signs SUDO.md + proxy.md + CLAUDE_CMDS.md together. (proxy.md signature status UNVERIFIED — do not assert it is signed.)

D5 — Fail mode: FAIL-CLOSED

  • If the hook crashes or exceeds its 5s timeout → deny everything until healthy. A broken security control must not silently become no control. (User-confirmed.)
  • Mitigations: (a) hook stays dead-simple pure-regex (minimal crash surface); (b) an ESCAPE_HATCH env flag the USER sets to drop to prompt-mode if the hook ever bricks all Bash; (c) SessionStart self-test so a broken deploy is caught before it blocks work.

D6 — Recovery ladder (two phases)

  • Phase 1 (now, background-agent automation): fail-closed → user reachable to grant permission / co-diagnose → fix → resume autonomy.
  • Phase 2 (full autonomy): hook fail → ALL WORK PAUSES → Hermes (already monitoring) attempts diagnose+fix → success → resume; impossible → ALL WORK STOPS + high-priority NTFY ("catastrophic failure, all work halted, interactive session needed"). In that session Claude/Hermes still do the heavy lifting; user only grants permission.
  • Depends on Hermes (#138) for Phase 2.

D7 — Circuit-breaker + self-disarm (inherited from Agent-Sudo)

  • Same anomaly trip: N tier-3 failures in a window OR any tier-4 attempt → daemon self-disarms to read-only + NTFY. No new mechanism.

D8 — Training capture (inherited; zero new schema)

  • Every classify+outcome → command_audit (projects DB): assigned_tier, decision_type, matched_rule, evidence, exit_code, rollback_taken, verify_passed, human_verdict (gold label), training_signal. This is the local-model training export.

Phased build (background agents, like Agent-Sudo P0P5)

  • CA-P0 — write CLAUDE_CMDS.md tier-map + the structural classifier module (pure-regex effect-classes). No activation.
  • CA-P1 — wire the hook to emit permissionDecision allow/deny/ask; tier-0/1 local allow; tier-4 deny; tier-2/3+unknown → route to daemon /exec. Fail-closed + ESCAPE_HATCH + SessionStart self-test. Test in the mirror first.
  • CA-P2 — daemon side: accept routed Claude-Code commands, sandbox-test unknowns in Incus mirror, capture rollback, log command_audit. (Mostly exists — extend/verify.)
  • CA-P3 — secrets-proxy env_secrets path for secret-touching commands. GATED on #128.
  • CA-P4 — sign CLAUDE_CMDS.md (transit) + learned-allowlist promotion from command_audit. GATED on #145.
  • CA-P5 — Hermes recovery-ladder (Phase 2) integration. GATED on #138.

EXECUTION ORDER (user-approved 2026-07-14, revised)

User's stated order was: (1) finish Agent-Sudo P3 subsystems [command testing + recovery] + sign SUDO.md; (2) finish/test/deploy secrets-proxy + sign proxy.md; (3) design/test/deploy constrained autonomy + sign CLAUDE_CMDS.md. User's concern: "I'll have to be here to give permission for everything."

REVISED — insert item 0 first. The ordering paradox: item 3 is what frees the user, but sat last, behind the two most babysitting-heavy items. Item 3's fast-path has NO dependency on items 1/2 (only CA-P2 needs item 1's sandbox; only CA-P3 needs item 2's proxy).

  • CA-P1a — CONSERVATIVE FAST-PATH (do FIRST; zero dependencies). Hook classifies: a TIGHT, EXPLICITLY-ENUMERATED read-only set (git status/diff/log, ls, grep, find, docker ps/inspect, cat non-secret, scratch scripts) → allow locally, no prompt. Catastrophic verbs → deny (protection that does NOT exist today). Everything else falls through to the normal prompt exactly as today. Purely additive: net SAFER (adds tier-4 deny) AND less annoying. Removes ~8090% of build-time prompts since building is overwhelmingly reads. Do NOT start with the full structural classifier — that's where a mis-classification could auto-allow something real. Start tight, widen with command_audit evidence.
  • Then 1 → 2 → 3 in the user's order, each one widening the fast-path: item 1's sandbox unlocks CA-P2 routing (destructive minority stops prompting); item 2 unlocks CA-P3 secrets path; item 3 completes + signs.
  • Net: user is present only for the SHRINKING MINORITY of commands, immediately — instead of all of them until the very end.

#145 (Vault admin) — user reports RECOVERED but task not marked complete. Per feedback_verify_before_persist: VERIFY transit is actually enableable BEFORE marking done or signing. Next-session first action: verify → sign SUDO.md → close #145 → CA-P4's signing gate also falls.

CA-P1a — BUILT + LIVE (2026-07-14, task #148)

Shipped in /opt/appdata/docker/.claude/hooks/security-enforcement.py v2.0 as "Category 4", layered under the existing Category 13 blocks (which run FIRST and still win — a fast-path candidate that trips any security rule is BLOCKED, never allowed).

  • allow — every segment matches a tight enumerated read-only set (git read-verbs, ls/cat/grep/ find, docker ps/inspect/logs, systemctl status, journalctl, ip show, sysinfo). No prompt.
  • deny — catastrophic verbs (rm -rf / + system dirs, mkfs, dd→/dev, wipefs, fork bomb, destructive partition ops). New protection that did not exist before. No override.
  • prompt — everything else falls through exactly as before. Purely additive.

Safety rests on two independent conditions, both required: (1) transparent structure — ANY $, backtick, redirection, or backgrounding disqualifies; (2) every pipeline segment individually enumerated. One unknown segment disqualifies the whole command.

A real hole was caught by the tests, not by review: echo $VAULT_TOKEN initially classified allowecho is enumerated and $VAR is not $(, and Categories 13 don't catch it (not docker exec env, not a cat of a known secret path). It would have printed a live secret with no prompt. Fix: any $ disqualifies the fast-path — the hook cannot know what a variable holds, so it cannot certify the command as a read. Locked in as a regression case. Lesson: enumerating safe verbs is not enough; the structure must also be transparent, and only an adversarial test suite finds the gap.

Verified: 23/23 self-test cases; e2e over the real stdin/stdout protocol ALL PASS; classify() median 0.009ms (budget 50ms); malformed stdin → exit 0; non-Bash ignored. D5 fail mode: in P1a the hook is NOT the only gate — the prompt still backstops everything not fast-pathed, so "closed" = fall back to the prompt, never auto-allow. Any exception → no decision → prompt (strictly no worse than pre-v2.0). Deny-everything fail-closed arrives with CA-P1, when the hook becomes the sole gate. Escape hatch: CLAUDE_CA_ESCAPE_HATCH=1. D5 self-test wired into session-start.sh (step 9) — reports at every SessionStart, non-fatal. D8 training loop: every decision (including prompt) appends to /opt/appdata/docker/.claude/hooks/ca_decisions.jsonl — best-effort, never breaks the hook. The prompt rows are the D2 promotion candidates (approved-every-time ⇒ widen the fast-path).

CA-P1b — WIDENED ON EVIDENCE (2026-07-14, hook v2.1, task #147)

The first production data made the case, not intuition. After CA-P1a shipped, the user observed he was still approving nearly everything. ca_decisions.jsonl answered why: 29 decisions, 22 prompts, and 16 were "opaque structure" — not mutations, not danger.

Root cause: 2>&1 contains a >. v2.0's rule was "any redirection is a WRITE ⇒ disqualify." But 2>&1 is file-descriptor plumbing and >/dev/null is a discard — neither can write anything. The rule was rejecting the exact idiom ordinary diagnostic reads are written in (ls -la 2>&1 | head). The classifier wasn't being cautious; it was being wrong.

Three changes, all still tier-0. None widens WHAT may run — they let the hook recognise reads it was already supposed to allow:

  1. Safe redirects (_SAFE_REDIRECT_RE) stripped before the redirect check: 2>&1, >&2, 2>/dev/null, &>/dev/null. A redirect to any REAL path still disqualifies; /dev/null is anchored so >/dev/nullx and >/dev/null/../../etc/passwd cannot ride the prefix (tested).
  2. cd enumerated — no filesystem effect, and every other segment is still checked independently (cd /etc && rm -rf x still prompts on the rm).
  3. ssh <host> '<cmd>' — classify the INNER command under the identical rules, depth-limited to one hop. A read is a read regardless of which host runs it. Strict shape only: no options (ssh -o ProxyCommand=… prompts), no unquoted form. This is what stopped server-01 work from taxing the user on every single ls.

Security argument for the ssh hop, and its regression test: Categories 13 scan the FULL raw text (including the inner) before the fast-path is consulted, so ssh server-01 'cat …/agent-sudo/.env' is blockedcat is structurally a read verb, so the secret-path check is the ONLY thing standing between the ssh fast-path and an exfil channel. That case is a locked regression test; if it ever goes green-to-allow, the hop must be withdrawn.

Verified: 42/42 self-test (up from 23); e2e ALL PASS, no regression; classify() median 0.011ms. Live-proved: ssh server-01 'systemctl is-active …' → allow, no prompt. Known edge (accepted): 2>&1 immediately followed by a quote isn't stripped (lookahead wants whitespace/;/|/EOL), so quoted compounds stay opaque → prompt. Safe direction, low value to fix.

Tier-1 (reversible writes) NOT shipped — deliberately. D2 authorises it, but the log says Bash-tier-1 is a small slice of real traffic: file edits go through the Write/Edit tools (not this hook), scratchpad writes are already pre-authorised, and git add/commit collides with playbook_git_criteria_universal (commits are checklist-triggered — auto-allowing removes the last friction on a rule enforced only by judgement). Correct next move per feedback_evolution_by_default: run v2.1, let ca_decisions.jsonl name the next real tax, and widen on evidence. The log found this one; it can find the next one.

Verified live state (2026-07-14)

  • Hook: /opt/appdata/docker/.claude/hooks/security-enforcement.py, registered PreToolUse matcher=Bash timeout=5 in ~/.claude/settings.json. Currently block-only (exit 0/2), no classifier, no permissions allowlist set.
  • Agent-Sudo daemon /health+/exec+/allowlist: server-01 8082 (LIVE, server_id=server-01), primary 8084 (server_id=primary; primary cutover swap still pending #146).
  • command_audit (projects DB): 15 cols as above, present.