# Agent-Sudo (projects DB id=176) — Design Grill-Me: Locked Decisions & Open Gaps **Status:** ✅ GRILL-ME COMPLETE (2026-07-08, Opus 4.8/high). All decisions D1–D10 + gap-#5 cross-cutting LOCKED. Phased build prompts P0–P5 written to `agent_prompts.md`. NEXT = user review of prompts, then dispatch (not yet launched). **Started:** 2026-07-08 (Opus 4.8 / high). Design doc for the pre-existing id=176 design; this file records the sub-design decisions the DB row left unspecified. ## Verified current state (checked live 2026-07-08) - ❌ No Incus/LXC/LXD installed on server-01 (192.168.1.90) — P1 installs it - ❌ No `command_audit` table in `projects` DB (table `infrastructure_projects` holds id=176) — P0 creates it - ✅ `agent_test_results` table EXISTS in `api_business` DB (verified 2026-07-08); `automation_ideas` + `behavior_changes` also in `api_business` (NOT the empty `automation_ideas` DB) - ✅ timeshift installed on server-01 (`/usr/bin/timeshift`) - ✅ **Coolify RETIRED** (control plane stopped 2026-06-25) — deploy Agent-Sudo via plain docker-compose + Jenkins/Gitea, never Coolify; old bridge decommission (P5) = plain stop+rm, nothing to deregister. `coolify-proxy` = standalone Traefik (name kept for cloudflared). See [[project-coolify-traefik-retirement]]. - ✅ **N8N is NOT a caller of the bridge** — all 12 production workflows checked 2026-07-08, zero reference sudo-bridge (they use the now-dead Coolify API + Vault + bitwarden-bridge). So the D8 cutover shadow-diff has ZERO N8N surface; real callers = claude-code / background-agents / (future) Hermes via PROXY_CALLERS. - ⚠️ **PORT CORRECTION to D8:** Agent-Sudo parallel port = **8084**, NOT 8083 (8083 is bitwarden-bridge). 8084/8085/8086/8087 free on primary as of 2026-07-08. - Existing sudo-bridge: primary `sudo-bridge-drbjegv07256ki2lpfyr00n8` (port 8082 @192.168.1.88); server-01 bridge (Coolify UUID o2kz1puml1mmneyiqd96mouj, port 8082 @192.168.1.90); shared image `gitea.local/backtalk6858/sudo-bridge:latest`; API keys in Vault `secret/data/sudo-bridge` and `secret/data/sudo-bridge-server01`. ## LOCKED DECISIONS ### D1 — LXC mirror fidelity = COMMAND-SURFACE mirror Same OS + package versions, same directory structure/paths, same users/sudoers/groups. **No** live data, **no** running containers. Static, script-rebuildable, zero production-data exposure in sandbox. (Rejected: state-replica mirror — constant drift, high upkeep, leaks prod data.) ### D2 — Tier-2 semantics = RE-TIER BY REPRODUCIBILITY The current 4-level testing methodology tests *built automations*; sandbox-testing *arbitrary sudo commands* universally is impractical (would need the rejected faithful replica). So: - **Reproducible ops** (file edits, package installs, permission/config changes) → run the **sandbox-execute** path in the command-surface mirror, using **static/dry-run validation** (parse, path resolution, forbidden-path check, `--dry-run`/`shellcheck`). Does not claim to prove live runtime behavior. - **Live-service ops** (`docker`/`systemctl` restarts, network changes, live DB) → **skip the sandbox** (it structurally can't test them) → **Tier 3**: Timeshift snapshot → execute on target → full audit trail. If it fails, restore from snapshot and post-mortem from the audit record. - Net: the LXC mirror earns its keep only for reproducible ops; live-service safety = snapshot + audit, not a theatrical sandbox run. ### D3 — LOCKED (confirmed 2026-07-08): rollback is scoped, not Timeshift; primary is stricter Emerged from the "primary has no failover" question. Confirmed by user. **Tier-4-on-primary = HARD REFUSE, interactive-session-only (option a).** No remote/NTFY approval path for primary Tier-4 — a command dangerous enough to be Tier-4 on the most critical host requires a human at a real terminal. (server-01 may still allow refuse-and-notify at its own lower tiers.) **Long-term (Horizon B / redundancy vision):** once a redundant failover setup exists (Proxmox mirror + server redundancy/k3s — see [[project_server_redundancy_vision]], [[project_proxmox_test_server]]), revisit allowing full Timeshift restore — with failover, actual downtime → ~0%, so the "full restore = total downtime" objection dissolves and the primary posture can relax. Until then: conservative-by-default. - **Default Tier-3 rollback = scoped command-specific undo**, captured BEFORE execution (file backup+restore of only touched files; package record+remove; container recovery via **Jenkins redeploy**, not system restore). Blast radius = only what the command touched. - **Timeshift full-restore = emergency last resort ONLY** (OS-level corruption: dpkg/systemd/libs). User-in-the-loop, NTFY-urgent, accept full downtime. Agent-Sudo may TAKE the snapshot but must NEVER auto-restore primary. Rationale: full restore = total downtime for ALL services (incl. Vault/Authelia/secrets-proxy/media) + reboot, AND reverting /var/lib/docker risks stateful DATA LOSS (Postgres). - **Host-aware tiering:** primary classifier stricter than server-01. Same command may be Tier-3 on server-01 but **Tier-4 (refuse → require interactive session)** on primary — especially anything touching critical security infra or stateful DB volumes. - Long-term real fix = Proxmox test-mirror + server redundancy/k3s failover (already in vision). Until then primary posture = conservative-by-default. - Feeds SUDO.md design (gap #1): the classifier must be **host-aware** (per-host tier overrides). ### D4 — LOCKED (2026-07-08): tier model has NO human-approval gate The current sudo-bridge tier-2 "NTFY approve/deny then execute" gate is being **eliminated**, not preserved. Rationale (why — this is the whole point of Agent-Sudo): - It didn't work: user misses NTFY notifications → command times out (900s APPROVAL_TIMEOUT), compounded by up to 5 parallel background agents competing for attention. - It was theatrical: user was asked to approve commands they couldn't evaluate for safety → rubber-stamping, not a real gate. - Real safety = testing/verifying the command (sandbox for reproducible ops; snapshot+audit+auto-verify for live ops), NOT a human who lacks context. - Only human touchpoint remaining = Tier-4 hard-refuse → interactive session. Classifier must be tuned so Tier-4 is a near-zero-frequency event. **Canonical tier model (safe→dangerous ascending; inverts the live bridge's 0=refuse):** | Tier | Name | Behavior | |---|---|---| | 0 | auto-safe (read-only) | auto-execute, no snapshot | | 1 | auto-reversible | auto-execute + capture scoped-undo | | 2 | sandbox-first (reproducible ops) | dry-run/validate in LXC mirror → execute → audit | | 3 | snapshot+execute (live-service ops) | scoped-undo (+Timeshift snapshot) → execute → full audit → auto-verify | | 4 | hard-refuse | refuse → require interactive session (host-aware: stricter on primary) | Migration note (P0/P4): existing allowlist.json uses live-bridge numbering (0=refuse,1=auto,2=NTFY-approve). Remap on cutover: live tier-1 → new tier-0/1 (by read vs write); live tier-2 → new tier-2 or 3 (by reproducible vs live-service); danger vetoes → new tier-4. ### D5 — LOCKED (2026-07-08): matching engine + seed corpus + audit-driven self-completing loop **Engine = option (a):** evolve the existing model — ordered rule list, most-specific-first (first-match-wins), `fnmatch` globs (existing allowlist.json migrates ~1:1), regex danger-veto = auto-Tier-4. Each rule gains optional per-host tier override + verb class. Verb-heuristic is a *fallback default only*, not the primary engine. **Miss behavior (host-aware, per D3):** allowlist miss on **primary → hard-refuse (Tier-4)** (primary never guesses); on **server-01 → verb-heuristic default** allowed (sandbox-side may guess). Ambiguous even under heuristic → Tier-4. **Seed corpus:** pre-populate SUDO.md via internet research on the common command set we actually run + what each does → assign each a tier up front. Reduces cold-start Tier-4 misses. **Audit-driven classification loop (the self-completing part — this is the training-data mechanism):** 1. Unknown command arrives. Reproducible-verb unknowns route to the **sandbox** (Tier-2 treatment) — NOT the live target. Live-service unknowns on primary → Tier-4 until classified (no blind execution on primary). 2. Run in sandbox → analyze the audit trail (exit code, output, stderr). 3. If it executes cleanly → evidence supports a tier; propose the classification. 4. If it fails because the sandbox is missing something (path/package/config/user) → analyze the audit trail to identify the gap → add it to the command-surface mirror **rebuild script** → rebuild → re-run. This makes the LXC mirror (D1) self-completing over time. 5. The AI evaluator (P3) produces a **proposed** tier + audit evidence and writes the rule into SUDO.md in a **`proposed` (inactive) state on BOTH hosts** — a proposed rule cannot be used to execute until reviewed. It fires a **low-priority NTFY** ("new proposed Agent-Sudo rule — load conversation in an interactive session to review"). Rule becomes `active` only after the review discussion. Logged as training data throughout. - **Why this does NOT contradict D4:** D4 killed the blocking *execution*-approval gate (timed out, on live commands the user couldn't evaluate). This is a **non-blocking async review of a config/rule change** — nothing is stalled waiting on it (no timeout failure mode), it's low-frequency (meta-level), and it happens in a session where the user CAN evaluate it with Claude's help. Different mechanism, different risk profile. - Implementation detail (defer): NTFY must carry a stable conversation pointer for the review session (candidate: a dedicated `agent-sudo-rules` conversation, or the agent-builder conversation). ### D6 — LOCKED (2026-07-08): operating principle = minimize human involvement, not eliminate it Until human-level AGI exists, reducing the user's involvement to exactly zero is not achievable. Design target for every autonomous system = **minimize** required human touchpoints, not **eliminate** them. Legitimate residual touchpoints (Tier-4 interactive execution; proposed-rule review) are kept deliberately, made non-blocking and low-frequency, and each is justified. (Candidate universal feedback memory — applies beyond Agent-Sudo.) ## OPEN GAPS (resume grill-me here) ### D7 — LOCKED (2026-07-08): LXC architecture on server-01 Verified live: both hosts = LMDE 7 "gigi" (Debian 13 trixie base), kernel 6.12.90; server-01 = 31GB RAM/27GB free, 173GB free disk, btrfs tooling present (no zfs), no LXC/LXD/Incus yet. Primary confirmed CLEAN of all testing infra (no LXC/sandbox/test containers; only prod postgres + vault; timeshift present = intended D3 snapshot tool). - **Container tech = Incus** (Debian-trixie-native; profiles/limits/instant snapshots for cheap reset-to-baseline). - **Base image = `images:debian/13` (trixie).** LMDE≈Debian under the hood; LMDE's deltas are desktop/theming, irrelevant to sudo command behavior. Confirmed acceptable by user. - **Storage = btrfs pool** (loop-file or subvolume) for near-instant snapshot/rollback; fallback `dir` backend if root isn't btrfs. - **Network = NAT via default `incusbr0`** — outbound works (nala/apt during classification tests), ZERO inbound LAN exposure (sandbox must not be reachable on 192.168.1.x). - **Limits = 2 vCPU / 2GB RAM / 10GB disk each** (ephemeral scaffolds, no services). - **Rebuild = one parameterized `build-sandbox.sh `** (D1 script-rebuildable; D5 loop appends missing deps). Differs per host only in package set + paths + users/sudoers. - **P1 prerequisite = read-only host inventory** (the mirror spec): `dpkg --get-selections`, users/uids+groups+sudoers STRUCTURE (never secrets), relevant dir tree (`/opt/appdata`, `/data/coolify`, `/etc/...`). Both mirrors (`sandbox-primary`, `sandbox-server01`) run ON server-01; `sandbox-primary` mirrors production's command surface but never runs on production. ### D9 — LOCKED (2026-07-08): privilege + security model = FULLY AUTONOMOUS, no human arming Hard constraint (user): must be fully autonomous — if the user is unavailable, automation must not stall. So NO human-arming gate on the execution path. Once full autonomy is required, "can the machine escalate on its own?" = yes (settled); security moves OFF access-gating (theater) ONTO the constraint layer. **The reframe (the crux):** separate two powers — - **EXECUTE an already-vetted command** → fully autonomous, no human ever, always available (Tiers 0–3). - **EXPAND what's allowed** (new command / raise tier / widen allowlist) → the ONLY human touchpoint; rare, async, non-blocking (D5 proposed-rule review). The autonomous system CANNOT widen its own allowlist. Attacker win = "run arbitrary dangerous command" ⇒ requires widening the allowlist ⇒ impossible autonomously; executor enforces danger-veto regardless. Full compromise of the smart brain ⇒ only inherits "run already-safe/reversible/audited commands" = bounded blast radius. **Unavailability is NOT a single point of failure:** all Tier 0–3 automation runs with user asleep/away; only (a) activating a brand-new capability and (b) catastrophic Tier-4 wait for a human — both SHOULD wait. **Privilege mechanism:** dumb root daemon everywhere (auth-agnostic — survives user's fingerprint→privacyIDEA MFA→YubiKey roadmap; never calls sudo/PAM). Both hosts = always-on SCOPED short-TTL AppRole (reuse existing rotate-and-revoke machinery; policy reads ONLY the one Agent-Sudo secret path). Bitwarden + OS-sudo-password both DROPPED from design (Bitwarden was only for the abandoned OS-password pull). **Autonomous-security stack (all human-free):** least-privilege allowlist · dumb minimal root executor · independent re-validation (already in code) · signed/read-only SUDO.md (Vault transit — allowlist can't be silently widened) · scoped short-TTL AppRole · executor sandboxing (AppArmor/seccomp/systemd) · egress lockdown (executor has no outbound net) · append-only audit shipped off-box · **anomaly circuit-breaker that self-disarms** (watches own command rate/mix/failure spike → auto-stops Tier-3 + alerts — replaces the human panic button) · **dual-service concurrence on Tier-3** (independent second classifier must agree). Primary stays stricter (D3): larger Tier-4 set, catastrophic → refuse-to-human. **New tool? NO.** All pieces already exist (allowlist+executor, Vault AppRole + transit signing, AppArmor/seccomp/egress free on host). Falco (runtime anomaly detection) / OPA (policy-as-code) = documented FUTURE-ENHANCEMENT candidates only — new tools add attack surface + maintenance for a solo operator. ### D10 — LOCKED (2026-07-08): Hermes as the always-on independent monitor (future layer) Hermes (id=138, monitoring brain, local model on server-01, no Claude-API cost) is the natural home for the monitoring layers of D9 — layered on AFTER Agent-Sudo v1 ships. **Security gains:** (1) Hermes IS the intelligent anomaly circuit-breaker/watchdog — reads the append-only audit stream, trips the breaker on anomaly, NTFYs user; being a SEPARATE process from Agent-Sudo = separation of duties (compromising Agent-Sudo ≠ compromising the watchdog). (2) The independent second opinion for dual-service Tier-3 concurrence. (3) Audit drift-detection + integrity monitoring (SUDO.md signature valid? AppArmor loaded? egress rules present? → tamper alert). **Usability gains:** (1) triages/batches the two human touchpoints (proposed-rule activation, Tier-4) into digestible evidence-backed decisions. (2) Drives the D5 sandbox-test-and-analyze loop autonomously so unknown-command classification is one-click-confirm by review time. (3) Self-healing: detects failed side effects → triggers scoped-undo or Jenkins redeploy (per project_jenkins_hermes_boundary). **HARD BOUNDARY:** Hermes is on the MONITORING side only — observe/analyze/alert/trip-breaker/trigger-pre-vetted-recovery. Hermes NEVER gets execution privilege, NEVER widens the allowlist, NEVER runs arbitrary Tier-3+. If Hermes could command the executor freely it becomes the soft underbelly (compromise Hermes = escalate). Keeps automated dual-control intact: Agent-Sudo (has root, dumb about judgment) + Hermes (no root, smart about judgment) check each other. **Build order:** Agent-Sudo v1 ships with a simple self-contained rule-based circuit-breaker; Hermes later augments/replaces it as the intelligent watchdog. Ties to project_hermes_exploration_day (capability-map after id=138). ## OPEN GAPS (resume grill-me here) 1. ✅ RESOLVED — SUDO.md classifier (see D4 tiers, D5 engine/seed/loop, D6 principle). 2. ✅ RESOLVED — LXC/Incus specifics (see D7). 3. ✅ RESOLVED — Cutover & rollback (D8). 4. ✅ DROPPED — Bitwarden ephemeral pull no longer in design (D9: Vault AppRole replaces it; no OS-password, no Bitwarden). [Security model fully resolved: D9 + D10.] ### D8 — LOCKED (2026-07-08): cutover sequence & rollback Drop-in replacement: response contract stays identical (`POST /exec` Bearer → `{exit_code, output, ...}`; T4 refuse = 403 like today's allowlist-miss). Callers need zero changes; T2/T3 just add latency (sandbox/snapshot) transparently. D4 removes the tier-2 NTFY-approval block entirely. - **Deploy order = server-01 FIRST, primary LAST** (user confirmed). server-01 = testing host, low blast radius, and hosts the sandbox LXCs (T2 local). Primary only after a clean 4-level pass on server-01. - **Parallel run via ports:** Agent-Sudo stands up on **new port 8084** (⚠️ corrected 2026-07-08 from 8083 — 8083 is bitwarden-bridge) while old bridge keeps serving **8082** untouched. Shadow-validate (4-level gate on 8083 + replay known-safe commands through both, diff results). Callers stay on 8082 during validation → zero disruption. - **Atomic swap when green:** flip Agent-Sudo to 8082, demote old bridge to 8083 hot-standby (kept running, not deleted). - **Rollback = swap ports back** (old bridge one flip from resuming 8082). Retain old bridge container + image **14 days** before decommission (P5). - **Primary cutover = MANUAL port-swap by user** in an interactive session (user confirmed). Agent-Sudo is built + validated autonomously everywhere; the final "make it live on production" flip on primary is user's hands on keyboard (per D3/D6). server-01 cutover = autonomous. 4. **Bitwarden ephemeral sudo password pull** — confirm path via bitwarden-bridge; never stored at rest. 5. ✅ RESOLVED (2026-07-08) — **Cross-cutting:** - **Per-phase rollback/abort** — matrix baked into the P0–P5 prompts (P0 drop-table / P1 incus-delete / P2 delete-repo / P3 revoke-policy / P4 port-flip-back / P5 re-pull-image). No phase is destructive to the OLD bridge until P5. - **Vault/N8N/Coolify dry-run paths** — Vault = additive `agent-sudo` policy + revoke-self smoke test (least-priv proof); N8N = ZERO surface (not a caller); Coolify = retired, Agent-Sudo deploys OUTSIDE it via Jenkins, P5 = plain stop+rm. - **Training-data logging** — 6 decision points → `command_audit` rows; GOLD label = AI-eval-proposal → human review verdict (`human_verdict`). P0 schema carries the fields. - **Blast radius per tier** — 0 minimal→3 medium/large→4 zero; capped on primary by shifting the dangerous middle up to refuse; **Tier-3-on-primary = highest-attention path** (most tests + dual-concurrence + tightest breaker). Full-compromise ⇒ inherits only already-vetted T0–3, cannot widen allowlist or reach T4. - **Blocks/unblocks** — id=176 has NO blocker (buildable now); no DB row lists it as a blocker (link is architectural, not recorded). Unblocks: Hermes self-healing (id=138), the sandbox-testing program (it's the first test of the harness), and any Horizon-B autonomous automation needing unattended privileged exec. Does NOT unblock the 12 broken N8N workflows (those need the Jenkins/Coolify-API migration). ## AFTER GRILL-ME COMPLETES Write the **phased** background-agent prompt SET (NOT one monolith — see reasoning in session): - P0: DB foundation (`command_audit` in projects + ensure `agent_test_results` in api_business) — design-complete, no server risk. - P1: LXC/Incus setup + command-surface mirror build on server-01. - P2: Python service = drop-in sudo-bridge replacement (same API), classifier reads SUDO.md, audit → command_audit. - P3: Bitwarden ephemeral pull + Timeshift automation + AI evaluator (tier-2/3 flow). - P4: deploy both servers behind old bridge, run 4-level gate. - P5: decommission old sudo-bridge. Opus+high prompt-craft upgrades over the Sonnet-era prompts (agent_prompts.md): bake in the 4-level test gate + agent_test_results logging (Sonnet's omitted it); split locked-contract from reference-implementation instead of embedding brittle literal scripts; right-size per-phase turn budgets; explicit "design is locked — do not redesign" block; per-phase rollback/abort criteria + idempotent pre-flight checks. --- ## CIRCUIT-BREAKER — REMAINING DECISIONS LOCKED (2026-07-15, grill-me Q3/Q4/Q5) These close the last three open questions on task #149. D-CB1..D-CB4 (2026-07-14) fixed the breaker's *concept*; these fix its *numbers, storage, and self-protection*. **#149 is now design-complete — an implementing agent must NOT redesign any of this.** ### D-CB5 — LOCKED: benign-trip threshold = 3 tier-3 failures / 15-min SLIDING window / PER-HOST - **Sliding, not fixed.** A fixed window lets failures straddle the boundary (2 at the end of window 1 + 2 at the start of window 2 = 4 failures, no trip). Sliding closes that. - **Per-host, not global.** The hosts have different roles (primary = production, server-01 = sandbox); a sandbox failure burst must not disarm production tier-3. Decisive reason: a global counter needs cross-host state, which recreates the **#173 failure class** (server-01's safety depending on primary ⇒ drop primary to disarm server-01). Per-host adds no new cross-host edge. - **Known, accepted cost:** an attacker probing BOTH hosts gets 2× budget before a benign trip. Accepted because probing is a **tier-4 signal**, which latches hard via D-CB1 regardless of this counter. The benign counter only ever governs broken scripts — and broken scripts are per-host. - Threshold governs the BENIGN path only. Blast radius + re-arm still per D-CB1/D-CB2: benign ⇒ stop **tier-3 only**, auto re-arm after cooldown w/ exponential backoff on repeat; attack (any tier-4 attempt, or repeated trips inside one window) ⇒ **full read-only, human-only reset, NTFY urgent.** ### D-CB6 — LOCKED: breaker state = LOCAL APPEND-ONLY EVENT LOG, state DERIVED by replay `/var/lib/agent-sudo/breaker.log` (0600 root, append-only), one JSON event per line: `{ts, host, event: tier3_fail|TRIP|REARM|RESET, cause: benign|attack, scope: tier3|readonly, cmd_id}` - **State is never stored — always `replay(log)`.** Consequence that matters: a daemon restart replays the log and the breaker is STILL TRIPPED. In-memory state would make **restart = free reset = bypass** (the trap Q4 existed to catch). - **Local only. No Vault, no Postgres.** Both live on primary ⇒ server-01 would depend on a remote host for its own local safety state, and Vault-sealed/pg-down would mean *no breaker at all*. Same #173 class as D-CB5's global-counter rejection. State must survive restart AND not add a cross-host edge; local storage is the only option meeting both. - **Append-only over a mutable `breaker.json` — because the daemon is root.** Root can rewrite any local file, so "tamper-proof" is theater ([[feedback-autonomous-security-constrain-not-gate]]). A mutable state file can be silently set back to `armed` leaving no trace. An append-only log cannot be *edited* into a lie, only **truncated** — and truncation is **evident** (log shrinks, audit trail has a hole). This converts *impossible-to-tamper* (false) into *tampering-is-visible* (true and enforceable). Constrain, don't gate. - Free wins: trip/reset events are labelled **training data**; and the replay boundary is the **D10 seam** letting Hermes assume trip-classification/reset later without a rewrite. ### D-CB7 — LOCKED: self-cannibalization ⇒ host_override TIER-4 in CODE, BOTH hosts, whole Set B surface Found live during #145: `systemctl restart agent-sudo-daemon` routed THROUGH Agent-Sudo has the root daemon kill its own executor mid-request — post-exec verify/audit can never complete. **The command destroys its own witness.** Structural; `SNAPSHOT_ENABLED` does not touch it. Refuse in code (a `host_override`, NOT merely a SUDO.md rule) on **both** hosts: - `systemctl restart|stop|disable|mask agent-sudo-daemon` - edits to `/etc/systemd/system/agent-sudo-daemon.service` - writes to `agent-sudo/bridge/SUDO.md` - writes to `/var/lib/agent-sudo/breaker.log` ← **new Set B member created by D-CB6** - **Why the whole surface, not just `restart`:** the reason restart is forbidden is D-CB4 **Set B** — the unit env is where `SUDO_MD_VERIFY_ENFORCE` lives, so self-restart *is* the gate-disarm primitive. But `stop`/`disable`/`mask`/unit-file-edit all reach the same place, and D-CB6's log makes "reset the breaker" a one-line write. Blocking only `restart` blocks one door in an open frame. - **Why code-level, not a SUDO.md rule:** *a rule that lives in the file it protects is not a constraint, it is a suggestion.* (SUDO.md is signed, but host_override must not depend on the gate it exists to protect.) - **Why both hosts (rejecting tier-4-primary/tier-3-server-01):** server-01 is where Agent-Sudo changes get PROVEN. A gate behaving differently in test than in prod means the sandbox stops predicting production — the one host that proves the gate would be the one host not running it. - **Known, accepted cost:** nobody can restart the daemon through Agent-Sudo, including for legitimate deploys — a human needs a shell. That is **Set B working as designed**: human-only *even though fully recoverable*, because automating it removes the ceiling. Per D6 (minimize, not eliminate) this is a *near-zero-frequency* touchpoint, not a stall. **XREF:** #149 (impl), #146 (primary cutover), #173 (same cross-host class — still OPEN, and D-CB5/D-CB6 both deliberately avoid adding to it), D-CB1/D-CB2 (benign-vs-attack drives both blast radius and re-arm), D-CB4 Set B, D9, D10, CA-D7. ### D-CB8 — LOCKED (2026-07-15): WHICH tier-4 sources are attack signals (closes D-CB1's open sub-decision) D-CB1 says "any tier-4 attempt => attack latch" and explicitly left open: *"each trip signal must be classified benign-noise vs attack at design time."* THIS IS THAT DECISION. It is load-bearing: `app.py` reaches tier-4 from THREE different sources and **one of them is routine operation.** | tier-4 source (`d.source`) | breaker signal | why | |---|---|---| | `danger_veto` | **ATTACK LATCH** (read-only) | matched a DANGER_PATTERN — attack-shaped | | host_override (D-CB4 **Set B**) | **ATTACK LATCH** (read-only) | an attempt to disarm the gate itself | | `miss_primary_refuse` | **NONE** (audit only) | **just an unlisted command** | **Why `miss_primary_refuse` must NOT latch:** on primary it is the *default outcome for anything not on the allowlist* — it is the system working, not an attack. Taken literally, D-CB1 would make every typo and every new tool take the host to FULL READ-ONLY pending a human reset. **Verified against reality:** during the 2026-07-15 session Agent-Sudo correctly refused every privileged command with `source=miss_primary_refuse`; under the literal rule the breaker would have attack-latched on the FIRST one and stayed latched all day. That is a false-positive machine and directly contradicts D9 ("if the user is unavailable, automation must not stall"). **Rejected — "repeated misses trip BENIGN" (3-in-15min via the D-CB5 counter):** attractive because rapid-fire misses resemble enumeration, but normal exploratory work produces exactly that shape, so it re-imports the false-positive problem one layer down. A miss is not evidence of intent. **Rejected — "only host_override latches, danger_veto trips benign":** a real destructive attempt would then keep write access. Danger patterns catching fat-fingering as often as malice is fine — the response (stop writing) is correct for BOTH. **IMPLEMENTATION NOTE — the seam holds, no breaker change needed.** `circuit_breaker.record_tier4_ attempt()` stays correct as a primitive that ALWAYS attack-latches. D-CB8 governs only its CALLERS in `app.py`: call it for `danger_veto` and host_override; for `miss_primary_refuse` emit the existing `tier4_refuse` audit row and NO breaker event. Classification lives behind `classify_trip()` (the D10 seam) so Hermes can refine this later without a rewrite. **[[feedback-instrument-the-gate]] applies:** log every breaker decision WITH its reason and read the log on a schedule. If `miss_primary_refuse` bursts turn out to correlate with real probing, D-CB8 can be revisited **on logged evidence** — not on imagination. Widen/tighten on evidence. ### D-CB9 — LOCKED (2026-07-15): escalation = 4+ benign trips / 6h sliding (fixes an incoherent constant pair) **The bug this fixes was in the SPEC, not the code**, and it was caught by the implementing agent rather than by me. As first specified: `BASE_COOLDOWN_SECONDS`=10min but `TRIP_ESCALATION_WINDOW_SECONDS`=1h with an implicit 2-trip threshold. Since re-arm happens at 10min, **any repeat benign trip is NECESSARILY inside the 1h window** ⇒ every repeat escalates to an attack latch ⇒ **the exponential backoff (20/40/80min) is unreachable dead code.** Two explicitly-required behaviours (backoff; escalate-on-repeat) were mutually unsatisfiable at those constants. **LOCKED VALUES:** - `BASE_COOLDOWN_SECONDS = 10 * 60` - `MAX_COOLDOWN_SECONDS = 4 * 60 * 60` - `TRIP_ESCALATION_WINDOW_SECONDS = 6 * 60 * 60` (was 1h) - `TRIP_ESCALATION_THRESHOLD = 4` (was implicitly 2) Benign trips 1→2→3 stay benign with doubling cooldown (10 → 20 → 40 min); the **4th within a 6h sliding window** attack-latches. Two counters, deliberately different: **escalation** counts benign trips inside the 6h window; **backoff** counts benign trips since the last `RESET`. **WHY 4/6h rather than latching on trip #2 — the key insight:** *exponential backoff IS the anti-"attacker waits out the cooldown" mechanism.* 10→20→40→80min rate-limits a patient attacker automatically, with no latch required. Meanwhile a broken cron job emits an IDENTICAL signal to a patient attacker, and the breaker cannot tell them apart (same ambiguity as D-CB2). So latching on the 2nd trip mostly punishes broken scripts and re-imports the false-positive problem D-CB8 just removed. **A latch should mean "this is not resolving", not "you failed twice."** - Rejected — *raise BASE_COOLDOWN above the 1h window*: makes backoff reachable, but one broken script then kills tier-3 for 90+ min. Slow recovery fights D9 (automation must not stall). - Rejected — *drop escalation, backoff only*: simplest and most autonomous, but a permanently broken system would quietly self-throttle to the 4h cap forever and never ask for a human. A latch is how the system says "I cannot fix this myself." **PROCESS NOTE — worth keeping:** the constants were set casually mid-session to avoid stalling on a question, and were not checked against each other. The implementing agent flagged the contradiction instead of silently picking a reading that made its tests pass. That is the behaviour we want from build agents, and it is why the "if you think a decision is wrong, implement it anyway but say so" instruction is load-bearing — a compliant-but-silent agent would have shipped unreachable backoff and nobody would have known until an incident. XREF [[feedback-instrument-the-gate]].