# Constrained Autonomy for Claude Code — Design Decisions **Status:** DESIGN LOCKED (grill-me 2026-07-14). Not yet built. Build = phased background agents (CA-P0…P5), mirroring the Agent-Sudo build model. Do NOT implement ad hoc — follow this record. **Driving insight (user):** a permission prompt the human rubber-stamps — especially a truncated command they can't read or don't understand — adds *no* safety. It is the same theater as the NTFY approval gate we removed from Agent-Sudo. Kill the theater; replace it with an ENFORCED code constraint layer. Governing memory: `feedback_autonomous_security_constrain_not_gate` (constrain capability, don't gate access), `feedback_evolution_by_default`, `feedback_always_include_training_loop`, `feedback_secrets_via_proxy_only`. ## The core problem Claude Code's Bash tool runs as `administrator` (docker group + sudo). Today, consequential commands stop for a human permission prompt. `--dangerously-skip-permissions` removes that gate but adds NO constraint → strictly worse. We want: **remove the theater prompt, but move the real gate into deterministic enforced code** that reuses the Agent-Sudo constraint+recovery stack. ## D1 — Enforcement architecture: HYBRID (hook classifies+routes; daemon executes-with-constraint) - **The hook** = `security-enforcement.py` (existing PreToolUse Bash hook, pure-regex, <50ms, no-I/O) extended into the **decision point**. It runs regardless of permission mode and is code, not a rubber stamp. It emits a `permissionDecision`: - **tier-0/1 (read/reversible)** → `allow` LOCALLY (no prompt, no daemon round-trip → kills the theater, stays <50ms). - **tier-2/3 (mutating/destructive) + UNKNOWN** → `deny` the raw command, message routes it to the **Agent-Sudo daemon `/exec`** (which already does sandbox-test-first + rollback + audit). This generalizes what the hook ALREADY does for `sudo` today (block → route to bridge) into "anything consequential → Agent-Sudo". - **tier-4 (catastrophic)** → `deny`, no override. - **The daemon** = Agent-Sudo `/exec` (LIVE: server-01 8082, primary 8084). All heavy lifting (Incus sandbox mirror test, Timeshift/rollback, circuit-breaker, `command_audit` logging) stays where it's built and tested. We reuse it; we do not reimplement it in a 50ms regex hook. - Rejected: pure-hook (can't sandbox/rollback in 50ms), route-EVERYTHING-through-daemon (latency + hard dependency + wrong root context for harmless reads). ## D2 — Classification: STRUCTURAL effect-classifier, unknown→sandbox, learned allowlist on top - SUDO.md only covers the `sudo` surface; ~90% of Claude Code commands are non-sudo (git/docker/python/curl/file ops). So the classifier judges a command by its **effect class**, not exact-match: - reads/inspection → tier-0/1 `allow` - workspace-local / reversible writes (git add/commit, edits under project dirs, scratch scripts) → tier-1 `allow` - blast-radius mutations (docker rm/restart/compose down, writes outside workspace, DB writes, net-config, mv/rm on real paths) → tier-2/3 → daemon sandbox - catastrophic verbs (rm -rf / patterns, mkfs, dd of=/dev/…, fork bombs, disk/partition ops) → tier-4 `deny` - `sudo`-prefixed → daemon's SUDO.md governs (unchanged) - **genuinely unrecognizable structure → default to daemon sandbox (treat as tier-2)** — unknown = "prove it in the mirror first," NEVER "ask the tired human." (User-confirmed.) - **Learned allowlist** layered ON TOP: commands that ran clean N times (from `command_audit` + `human_verdict` gold labels) get promoted to the local fast-path → system gets faster with use (`feedback_evolution_by_default`). Starts from structural analysis so day-one has zero rubber-stamp gaps. - Policy source of truth = a signed **`CLAUDE_CMDS.md`** tier-map file (the effect-class rules), analogous to SUDO.md. ## D3 — Secrets: via secrets-proxy env_secrets (dormant until #128) - Any command that touches a secret runs via **secrets-proxy `/shell` + `env_secrets`** (`vault://path#field` / `bitwarden://item`) so values are sourced at runtime, never in stdout/history/argv. The hook's existing Category-2/3 blocks (secret exposure, hardcoded secrets) stay as the backstop that FORCES commands onto the proxy path. - **Dependency:** secrets-proxy = task **#128** (pending/down). Secret-handling path is built but dormant until #128; interim fallback = `feedback_secret_to_script_via_file`. Decoupled from the classifier build. ## D4 — Tamper-evidence: Vault transit signature — **UNBLOCKED 2026-07-14 (#145 resolved)** > **Transit is LIVE.** Vault admin WAS recovered (user was right; my first probe was wrong — > it checked only the least-privilege AppRole, which is *denied* transit by design, and never > followed the Bitwarden path). Chain: vault approle → `secret/data/bitwarden-bridge#BRIDGE_API_KEY` > → bridge :8083 `GET /secret?item=Hashicorp Vault&field=notes` → root token. > **GOTCHA: the token LABELED "Root token:" is REVOKED; a second UNLABELED token at the bottom of > the note is the live one. Search the whole note — do not trust the label.** > transit enabled + ed25519 key `agent-sudo-sudomd` created + **SUDO.md SIGNED** (sha256 > `38eca778…`, transit/verify PASSED) + **tamper test PASSED** (fake widened rule → gate REFUSED). > `CLAUDE_CMDS.md` (CA-P4) and proxy.md (#150) can now be signed the same way. ## D4 (original) — Vault transit signature (was dormant until #145) - `CLAUDE_CMDS.md` gets the IDENTICAL mechanism as `sudo_sign.py`: Vault transit ed25519 key, sign file sha256, `verify_gate` on load via scoped AppRole, `..._VERIFY_ENFORCE` flag. Tampered policy → refuse to honor. - **Dependency:** live signing blocked on **#145** (OpenBao admin lockout; transit engine not enableable yet). SUDO.md itself runs unsigned (enforce=false) for this reason. `CLAUDE_CMDS.md` inherits the posture: **built to verify, ships `enforce=false`, flips to `true` when #145 lands** — one activation signs SUDO.md + proxy.md + CLAUDE_CMDS.md together. (proxy.md signature status UNVERIFIED — do not assert it is signed.) ## D5 — Fail mode: FAIL-CLOSED - If the hook crashes or exceeds its 5s timeout → `deny` everything until healthy. A broken security control must not silently become no control. (User-confirmed.) - Mitigations: (a) hook stays dead-simple pure-regex (minimal crash surface); (b) an `ESCAPE_HATCH` env flag the USER sets to drop to prompt-mode if the hook ever bricks all Bash; (c) SessionStart self-test so a broken deploy is caught before it blocks work. ## D6 — Recovery ladder (two phases) - **Phase 1 (now, background-agent automation):** fail-closed → user reachable to grant permission / co-diagnose → fix → resume autonomy. - **Phase 2 (full autonomy):** hook fail → ALL WORK PAUSES → **Hermes** (already monitoring) attempts diagnose+fix → success → resume; impossible → ALL WORK STOPS + **high-priority NTFY** ("catastrophic failure, all work halted, interactive session needed"). In that session Claude/Hermes still do the heavy lifting; user only grants permission. - Depends on Hermes (#138) for Phase 2. ## D7 — Circuit-breaker + self-disarm (inherited from Agent-Sudo) - Same anomaly trip: N tier-3 failures in a window OR any tier-4 attempt → daemon self-disarms to read-only + NTFY. No new mechanism. ## D8 — Training capture (inherited; zero new schema) - Every classify+outcome → `command_audit` (projects DB): `assigned_tier`, `decision_type`, `matched_rule`, `evidence`, `exit_code`, `rollback_taken`, `verify_passed`, `human_verdict` (gold label), `training_signal`. This is the local-model training export. ## Phased build (background agents, like Agent-Sudo P0–P5) - **CA-P0** — write `CLAUDE_CMDS.md` tier-map + the structural classifier module (pure-regex effect-classes). No activation. - **CA-P1** — wire the hook to emit `permissionDecision` allow/deny/ask; tier-0/1 local allow; tier-4 deny; tier-2/3+unknown → route to daemon `/exec`. Fail-closed + `ESCAPE_HATCH` + SessionStart self-test. Test in the mirror first. - **CA-P2** — daemon side: accept routed Claude-Code commands, sandbox-test unknowns in Incus mirror, capture rollback, log `command_audit`. (Mostly exists — extend/verify.) - **CA-P3** — secrets-proxy env_secrets path for secret-touching commands. GATED on #128. - **CA-P4** — sign `CLAUDE_CMDS.md` (transit) + learned-allowlist promotion from `command_audit`. GATED on #145. - **CA-P5** — Hermes recovery-ladder (Phase 2) integration. GATED on #138. ## EXECUTION ORDER (user-approved 2026-07-14, revised) User's stated order was: (1) finish Agent-Sudo P3 subsystems [command testing + recovery] + sign SUDO.md; (2) finish/test/deploy secrets-proxy + sign proxy.md; (3) design/test/deploy constrained autonomy + sign CLAUDE_CMDS.md. User's concern: "I'll have to be here to give permission for everything." **REVISED — insert item 0 first.** The ordering paradox: item 3 is what frees the user, but sat last, behind the two most babysitting-heavy items. Item 3's fast-path has NO dependency on items 1/2 (only CA-P2 needs item 1's sandbox; only CA-P3 needs item 2's proxy). - **CA-P1a — CONSERVATIVE FAST-PATH (do FIRST; zero dependencies).** Hook classifies: a TIGHT, EXPLICITLY-ENUMERATED read-only set (git status/diff/log, ls, grep, find, docker ps/inspect, cat non-secret, scratch scripts) → `allow` locally, no prompt. Catastrophic verbs → `deny` (protection that does NOT exist today). **Everything else falls through to the normal prompt exactly as today.** Purely additive: net SAFER (adds tier-4 deny) AND less annoying. Removes ~80–90% of build-time prompts since building is overwhelmingly reads. Do NOT start with the full structural classifier — that's where a mis-classification could auto-allow something real. Start tight, widen with `command_audit` evidence. - Then **1 → 2 → 3** in the user's order, each one widening the fast-path: item 1's sandbox unlocks CA-P2 routing (destructive minority stops prompting); item 2 unlocks CA-P3 secrets path; item 3 completes + signs. - Net: user is present only for the SHRINKING MINORITY of commands, immediately — instead of all of them until the very end. **#145 (Vault admin) — user reports RECOVERED but task not marked complete. Per `feedback_verify_before_persist`: VERIFY transit is actually enableable BEFORE marking done or signing.** Next-session first action: verify → sign SUDO.md → close #145 → CA-P4's signing gate also falls. ## CA-P1a — BUILT + LIVE (2026-07-14, task #148) Shipped in `/opt/appdata/docker/.claude/hooks/security-enforcement.py` **v2.0** as "Category 4", layered under the existing Category 1–3 blocks (which run FIRST and still win — a fast-path candidate that trips any security rule is BLOCKED, never allowed). - **allow** — every segment matches a tight enumerated read-only set (git read-verbs, ls/cat/grep/ find, docker ps/inspect/logs, systemctl status, journalctl, ip show, sysinfo). No prompt. - **deny** — catastrophic verbs (`rm -rf /` + system dirs, mkfs, dd→/dev, wipefs, fork bomb, destructive partition ops). **New protection that did not exist before.** No override. - **prompt** — everything else falls through exactly as before. Purely additive. **Safety rests on two independent conditions, both required:** (1) transparent structure — **ANY** `$`, backtick, redirection, or backgrounding disqualifies; (2) every pipeline segment individually enumerated. One unknown segment disqualifies the whole command. **A real hole was caught by the tests, not by review:** `echo $VAULT_TOKEN` initially classified **allow** — `echo` is enumerated and `$VAR` is not `$(`, and Categories 1–3 don't catch it (not `docker exec env`, not a cat of a known secret path). It would have printed a live secret with no prompt. Fix: **any `$` disqualifies the fast-path** — the hook cannot know what a variable holds, so it cannot certify the command as a read. Locked in as a regression case. Lesson: enumerating safe *verbs* is not enough; the *structure* must also be transparent, and only an adversarial test suite finds the gap. **Verified:** 23/23 self-test cases; e2e over the real stdin/stdout protocol ALL PASS; `classify()` median **0.009ms** (budget 50ms); malformed stdin → exit 0; non-Bash ignored. **D5 fail mode:** in P1a the hook is NOT the only gate — the prompt still backstops everything not fast-pathed, so "closed" = *fall back to the prompt*, never auto-allow. Any exception → no decision → prompt (strictly no worse than pre-v2.0). Deny-everything fail-closed arrives with CA-P1, when the hook becomes the sole gate. Escape hatch: `CLAUDE_CA_ESCAPE_HATCH=1`. **D5 self-test** wired into `session-start.sh` (step 9) — reports at every SessionStart, non-fatal. **D8 training loop:** every decision (including `prompt`) appends to `/opt/appdata/docker/.claude/hooks/ca_decisions.jsonl` — best-effort, never breaks the hook. The `prompt` rows are the D2 promotion candidates (approved-every-time ⇒ widen the fast-path). ## CA-P1b — WIDENED ON EVIDENCE (2026-07-14, hook v2.1, task #147) **The first production data made the case, not intuition.** After CA-P1a shipped, the user observed he was *still* approving nearly everything. `ca_decisions.jsonl` answered why: 29 decisions, **22 prompts**, and **16 were "opaque structure"** — not mutations, not danger. **Root cause: `2>&1` contains a `>`.** v2.0's rule was "any redirection is a WRITE ⇒ disqualify." But `2>&1` is file-descriptor plumbing and `>/dev/null` is a discard — neither can write anything. The rule was rejecting the exact idiom ordinary diagnostic reads are written in (`ls -la 2>&1 | head`). The classifier wasn't being cautious; it was being wrong. Three changes, all still tier-0. **None widens WHAT may run** — they let the hook recognise reads it was already supposed to allow: 1. **Safe redirects** (`_SAFE_REDIRECT_RE`) stripped before the redirect check: `2>&1`, `>&2`, `2>/dev/null`, `&>/dev/null`. A redirect to any REAL path still disqualifies; `/dev/null` is anchored so `>/dev/nullx` and `>/dev/null/../../etc/passwd` cannot ride the prefix (tested). 2. **`cd`** enumerated — no filesystem effect, and every other segment is still checked independently (`cd /etc && rm -rf x` still prompts on the `rm`). 3. **`ssh ''`** — classify the INNER command under the identical rules, depth-limited to one hop. A read is a read regardless of which host runs it. Strict shape only: no options (`ssh -o ProxyCommand=… ` prompts), no unquoted form. This is what stopped server-01 work from taxing the user on every single `ls`. **Security argument for the ssh hop, and its regression test:** Categories 1–3 scan the FULL raw text (including the inner) *before* the fast-path is consulted, so `ssh server-01 'cat …/agent-sudo/.env'` is **blocked** — `cat` is structurally a read verb, so the secret-path check is the ONLY thing standing between the ssh fast-path and an exfil channel. That case is a locked regression test; if it ever goes green-to-allow, the hop must be withdrawn. **Verified:** 42/42 self-test (up from 23); e2e ALL PASS, no regression; `classify()` median **0.011ms**. Live-proved: `ssh server-01 'systemctl is-active …'` → allow, no prompt. **Known edge (accepted):** `2>&1` immediately followed by a quote isn't stripped (lookahead wants whitespace/`;`/`|`/EOL), so quoted compounds stay opaque → prompt. Safe direction, low value to fix. **Tier-1 (reversible writes) NOT shipped — deliberately.** D2 authorises it, but the log says Bash-tier-1 is a *small* slice of real traffic: file edits go through the Write/Edit tools (not this hook), scratchpad writes are already pre-authorised, and `git add/commit` collides with `playbook_git_criteria_universal` (commits are checklist-triggered — auto-allowing removes the last friction on a rule enforced only by judgement). Correct next move per `feedback_evolution_by_default`: run v2.1, let `ca_decisions.jsonl` name the *next* real tax, and widen on evidence. The log found this one; it can find the next one. ## Verified live state (2026-07-14) - Hook: `/opt/appdata/docker/.claude/hooks/security-enforcement.py`, registered PreToolUse matcher=Bash timeout=5 in `~/.claude/settings.json`. Currently block-only (exit 0/2), no classifier, no permissions allowlist set. - Agent-Sudo daemon `/health`+`/exec`+`/allowlist`: server-01 8082 (LIVE, server_id=server-01), primary 8084 (server_id=primary; primary cutover swap still pending #146). - `command_audit` (projects DB): 15 cols as above, present.