Files
claude-projects/agent-builder/agent_sudo_design_decisions.md
T
Backtalk6858 c50e72e17b docs: agent-sudo grill-me complete — D1-D10, P0-P5 build prompts, comparison
- Close Agent-Sudo (infra id=176) design grill-me; gap #5 cross-cutting locked
- Add P0-P5 phased background-agent build prompts (Opus 4.8 upgrades applied)
- Add agent_sudo_vs_sudo_bridge.md comparison doc
- Update context.md: sprint day 11 session 4, Coolify/N8N/port-8084 corrections

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-09 14:28:39 -05:00

21 KiB
Raw Blame History

Agent-Sudo (projects DB id=176) — Design Grill-Me: Locked Decisions & Open Gaps

Status: GRILL-ME COMPLETE (2026-07-08, Opus 4.8/high). All decisions D1D10 + gap-#5 cross-cutting LOCKED. Phased build prompts P0P5 written to agent_prompts.md. NEXT = user review of prompts, then dispatch (not yet launched). Started: 2026-07-08 (Opus 4.8 / high). Design doc for the pre-existing id=176 design; this file records the sub-design decisions the DB row left unspecified.

Verified current state (checked live 2026-07-08)

  • No Incus/LXC/LXD installed on server-01 (192.168.1.90) — P1 installs it
  • No command_audit table in projects DB (table infrastructure_projects holds id=176) — P0 creates it
  • agent_test_results table EXISTS in api_business DB (verified 2026-07-08); automation_ideas + behavior_changes also in api_business (NOT the empty automation_ideas DB)
  • timeshift installed on server-01 (/usr/bin/timeshift)
  • Coolify RETIRED (control plane stopped 2026-06-25) — deploy Agent-Sudo via plain docker-compose + Jenkins/Gitea, never Coolify; old bridge decommission (P5) = plain stop+rm, nothing to deregister. coolify-proxy = standalone Traefik (name kept for cloudflared). See project-coolify-traefik-retirement.
  • N8N is NOT a caller of the bridge — all 12 production workflows checked 2026-07-08, zero reference sudo-bridge (they use the now-dead Coolify API + Vault + bitwarden-bridge). So the D8 cutover shadow-diff has ZERO N8N surface; real callers = claude-code / background-agents / (future) Hermes via PROXY_CALLERS.
  • ⚠️ PORT CORRECTION to D8: Agent-Sudo parallel port = 8084, NOT 8083 (8083 is bitwarden-bridge). 8084/8085/8086/8087 free on primary as of 2026-07-08.
  • Existing sudo-bridge: primary sudo-bridge-drbjegv07256ki2lpfyr00n8 (port 8082 @192.168.1.88); server-01 bridge (Coolify UUID o2kz1puml1mmneyiqd96mouj, port 8082 @192.168.1.90); shared image gitea.local/backtalk6858/sudo-bridge:latest; API keys in Vault secret/data/sudo-bridge and secret/data/sudo-bridge-server01.

LOCKED DECISIONS

D1 — LXC mirror fidelity = COMMAND-SURFACE mirror

Same OS + package versions, same directory structure/paths, same users/sudoers/groups. No live data, no running containers. Static, script-rebuildable, zero production-data exposure in sandbox. (Rejected: state-replica mirror — constant drift, high upkeep, leaks prod data.)

D2 — Tier-2 semantics = RE-TIER BY REPRODUCIBILITY

The current 4-level testing methodology tests built automations; sandbox-testing arbitrary sudo commands universally is impractical (would need the rejected faithful replica). So:

  • Reproducible ops (file edits, package installs, permission/config changes) → run the sandbox-execute path in the command-surface mirror, using static/dry-run validation (parse, path resolution, forbidden-path check, --dry-run/shellcheck). Does not claim to prove live runtime behavior.
  • Live-service ops (docker/systemctl restarts, network changes, live DB) → skip the sandbox (it structurally can't test them) → Tier 3: Timeshift snapshot → execute on target → full audit trail. If it fails, restore from snapshot and post-mortem from the audit record.
  • Net: the LXC mirror earns its keep only for reproducible ops; live-service safety = snapshot + audit, not a theatrical sandbox run.

D3 — LOCKED (confirmed 2026-07-08): rollback is scoped, not Timeshift; primary is stricter

Emerged from the "primary has no failover" question. Confirmed by user. Tier-4-on-primary = HARD REFUSE, interactive-session-only (option a). No remote/NTFY approval path for primary Tier-4 — a command dangerous enough to be Tier-4 on the most critical host requires a human at a real terminal. (server-01 may still allow refuse-and-notify at its own lower tiers.) Long-term (Horizon B / redundancy vision): once a redundant failover setup exists (Proxmox mirror + server redundancy/k3s — see project_server_redundancy_vision, project_proxmox_test_server), revisit allowing full Timeshift restore — with failover, actual downtime → ~0%, so the "full restore = total downtime" objection dissolves and the primary posture can relax. Until then: conservative-by-default.

  • Default Tier-3 rollback = scoped command-specific undo, captured BEFORE execution (file backup+restore of only touched files; package record+remove; container recovery via Jenkins redeploy, not system restore). Blast radius = only what the command touched.
  • Timeshift full-restore = emergency last resort ONLY (OS-level corruption: dpkg/systemd/libs). User-in-the-loop, NTFY-urgent, accept full downtime. Agent-Sudo may TAKE the snapshot but must NEVER auto-restore primary. Rationale: full restore = total downtime for ALL services (incl. Vault/Authelia/secrets-proxy/media) + reboot, AND reverting /var/lib/docker risks stateful DATA LOSS (Postgres).
  • Host-aware tiering: primary classifier stricter than server-01. Same command may be Tier-3 on server-01 but Tier-4 (refuse → require interactive session) on primary — especially anything touching critical security infra or stateful DB volumes.
  • Long-term real fix = Proxmox test-mirror + server redundancy/k3s failover (already in vision). Until then primary posture = conservative-by-default.
  • Feeds SUDO.md design (gap #1): the classifier must be host-aware (per-host tier overrides).

D4 — LOCKED (2026-07-08): tier model has NO human-approval gate

The current sudo-bridge tier-2 "NTFY approve/deny then execute" gate is being eliminated, not preserved. Rationale (why — this is the whole point of Agent-Sudo):

  • It didn't work: user misses NTFY notifications → command times out (900s APPROVAL_TIMEOUT), compounded by up to 5 parallel background agents competing for attention.
  • It was theatrical: user was asked to approve commands they couldn't evaluate for safety → rubber-stamping, not a real gate.
  • Real safety = testing/verifying the command (sandbox for reproducible ops; snapshot+audit+auto-verify for live ops), NOT a human who lacks context.
  • Only human touchpoint remaining = Tier-4 hard-refuse → interactive session. Classifier must be tuned so Tier-4 is a near-zero-frequency event.

Canonical tier model (safe→dangerous ascending; inverts the live bridge's 0=refuse):

Tier Name Behavior
0 auto-safe (read-only) auto-execute, no snapshot
1 auto-reversible auto-execute + capture scoped-undo
2 sandbox-first (reproducible ops) dry-run/validate in LXC mirror → execute → audit
3 snapshot+execute (live-service ops) scoped-undo (+Timeshift snapshot) → execute → full audit → auto-verify
4 hard-refuse refuse → require interactive session (host-aware: stricter on primary)

Migration note (P0/P4): existing allowlist.json uses live-bridge numbering (0=refuse,1=auto,2=NTFY-approve). Remap on cutover: live tier-1 → new tier-0/1 (by read vs write); live tier-2 → new tier-2 or 3 (by reproducible vs live-service); danger vetoes → new tier-4.

D5 — LOCKED (2026-07-08): matching engine + seed corpus + audit-driven self-completing loop

Engine = option (a): evolve the existing model — ordered rule list, most-specific-first (first-match-wins), fnmatch globs (existing allowlist.json migrates ~1:1), regex danger-veto = auto-Tier-4. Each rule gains optional per-host tier override + verb class. Verb-heuristic is a fallback default only, not the primary engine. Miss behavior (host-aware, per D3): allowlist miss on primary → hard-refuse (Tier-4) (primary never guesses); on server-01 → verb-heuristic default allowed (sandbox-side may guess). Ambiguous even under heuristic → Tier-4. Seed corpus: pre-populate SUDO.md via internet research on the common command set we actually run + what each does → assign each a tier up front. Reduces cold-start Tier-4 misses. Audit-driven classification loop (the self-completing part — this is the training-data mechanism):

  1. Unknown command arrives. Reproducible-verb unknowns route to the sandbox (Tier-2 treatment) — NOT the live target. Live-service unknowns on primary → Tier-4 until classified (no blind execution on primary).
  2. Run in sandbox → analyze the audit trail (exit code, output, stderr).
  3. If it executes cleanly → evidence supports a tier; propose the classification.
  4. If it fails because the sandbox is missing something (path/package/config/user) → analyze the audit trail to identify the gap → add it to the command-surface mirror rebuild script → rebuild → re-run. This makes the LXC mirror (D1) self-completing over time.
  5. The AI evaluator (P3) produces a proposed tier + audit evidence and writes the rule into SUDO.md in a proposed (inactive) state on BOTH hosts — a proposed rule cannot be used to execute until reviewed. It fires a low-priority NTFY ("new proposed Agent-Sudo rule — load conversation in an interactive session to review"). Rule becomes active only after the review discussion. Logged as training data throughout.
    • Why this does NOT contradict D4: D4 killed the blocking execution-approval gate (timed out, on live commands the user couldn't evaluate). This is a non-blocking async review of a config/rule change — nothing is stalled waiting on it (no timeout failure mode), it's low-frequency (meta-level), and it happens in a session where the user CAN evaluate it with Claude's help. Different mechanism, different risk profile.
    • Implementation detail (defer): NTFY must carry a stable conversation pointer for the review session (candidate: a dedicated agent-sudo-rules conversation, or the agent-builder conversation).

D6 — LOCKED (2026-07-08): operating principle = minimize human involvement, not eliminate it

Until human-level AGI exists, reducing the user's involvement to exactly zero is not achievable. Design target for every autonomous system = minimize required human touchpoints, not eliminate them. Legitimate residual touchpoints (Tier-4 interactive execution; proposed-rule review) are kept deliberately, made non-blocking and low-frequency, and each is justified. (Candidate universal feedback memory — applies beyond Agent-Sudo.)

OPEN GAPS (resume grill-me here)

D7 — LOCKED (2026-07-08): LXC architecture on server-01

Verified live: both hosts = LMDE 7 "gigi" (Debian 13 trixie base), kernel 6.12.90; server-01 = 31GB RAM/27GB free, 173GB free disk, btrfs tooling present (no zfs), no LXC/LXD/Incus yet. Primary confirmed CLEAN of all testing infra (no LXC/sandbox/test containers; only prod postgres + vault; timeshift present = intended D3 snapshot tool).

  • Container tech = Incus (Debian-trixie-native; profiles/limits/instant snapshots for cheap reset-to-baseline).
  • Base image = images:debian/13 (trixie). LMDE≈Debian under the hood; LMDE's deltas are desktop/theming, irrelevant to sudo command behavior. Confirmed acceptable by user.
  • Storage = btrfs pool (loop-file or subvolume) for near-instant snapshot/rollback; fallback dir backend if root isn't btrfs.
  • Network = NAT via default incusbr0 — outbound works (nala/apt during classification tests), ZERO inbound LAN exposure (sandbox must not be reachable on 192.168.1.x).
  • Limits = 2 vCPU / 2GB RAM / 10GB disk each (ephemeral scaffolds, no services).
  • Rebuild = one parameterized build-sandbox.sh <primary|server-01> (D1 script-rebuildable; D5 loop appends missing deps). Differs per host only in package set + paths + users/sudoers.
  • P1 prerequisite = read-only host inventory (the mirror spec): dpkg --get-selections, users/uids+groups+sudoers STRUCTURE (never secrets), relevant dir tree (/opt/appdata, /data/coolify, /etc/...). Both mirrors (sandbox-primary, sandbox-server01) run ON server-01; sandbox-primary mirrors production's command surface but never runs on production.

D9 — LOCKED (2026-07-08): privilege + security model = FULLY AUTONOMOUS, no human arming

Hard constraint (user): must be fully autonomous — if the user is unavailable, automation must not stall. So NO human-arming gate on the execution path. Once full autonomy is required, "can the machine escalate on its own?" = yes (settled); security moves OFF access-gating (theater) ONTO the constraint layer. The reframe (the crux): separate two powers —

  • EXECUTE an already-vetted command → fully autonomous, no human ever, always available (Tiers 03).
  • EXPAND what's allowed (new command / raise tier / widen allowlist) → the ONLY human touchpoint; rare, async, non-blocking (D5 proposed-rule review). The autonomous system CANNOT widen its own allowlist. Attacker win = "run arbitrary dangerous command" ⇒ requires widening the allowlist ⇒ impossible autonomously; executor enforces danger-veto regardless. Full compromise of the smart brain ⇒ only inherits "run already-safe/reversible/audited commands" = bounded blast radius. Unavailability is NOT a single point of failure: all Tier 03 automation runs with user asleep/away; only (a) activating a brand-new capability and (b) catastrophic Tier-4 wait for a human — both SHOULD wait. Privilege mechanism: dumb root daemon everywhere (auth-agnostic — survives user's fingerprint→privacyIDEA MFA→YubiKey roadmap; never calls sudo/PAM). Both hosts = always-on SCOPED short-TTL AppRole (reuse existing rotate-and-revoke machinery; policy reads ONLY the one Agent-Sudo secret path). Bitwarden + OS-sudo-password both DROPPED from design (Bitwarden was only for the abandoned OS-password pull). Autonomous-security stack (all human-free): least-privilege allowlist · dumb minimal root executor · independent re-validation (already in code) · signed/read-only SUDO.md (Vault transit — allowlist can't be silently widened) · scoped short-TTL AppRole · executor sandboxing (AppArmor/seccomp/systemd) · egress lockdown (executor has no outbound net) · append-only audit shipped off-box · anomaly circuit-breaker that self-disarms (watches own command rate/mix/failure spike → auto-stops Tier-3 + alerts — replaces the human panic button) · dual-service concurrence on Tier-3 (independent second classifier must agree). Primary stays stricter (D3): larger Tier-4 set, catastrophic → refuse-to-human. New tool? NO. All pieces already exist (allowlist+executor, Vault AppRole + transit signing, AppArmor/seccomp/egress free on host). Falco (runtime anomaly detection) / OPA (policy-as-code) = documented FUTURE-ENHANCEMENT candidates only — new tools add attack surface + maintenance for a solo operator.

D10 — LOCKED (2026-07-08): Hermes as the always-on independent monitor (future layer)

Hermes (id=138, monitoring brain, local model on server-01, no Claude-API cost) is the natural home for the monitoring layers of D9 — layered on AFTER Agent-Sudo v1 ships. Security gains: (1) Hermes IS the intelligent anomaly circuit-breaker/watchdog — reads the append-only audit stream, trips the breaker on anomaly, NTFYs user; being a SEPARATE process from Agent-Sudo = separation of duties (compromising Agent-Sudo ≠ compromising the watchdog). (2) The independent second opinion for dual-service Tier-3 concurrence. (3) Audit drift-detection + integrity monitoring (SUDO.md signature valid? AppArmor loaded? egress rules present? → tamper alert). Usability gains: (1) triages/batches the two human touchpoints (proposed-rule activation, Tier-4) into digestible evidence-backed decisions. (2) Drives the D5 sandbox-test-and-analyze loop autonomously so unknown-command classification is one-click-confirm by review time. (3) Self-healing: detects failed side effects → triggers scoped-undo or Jenkins redeploy (per project_jenkins_hermes_boundary). HARD BOUNDARY: Hermes is on the MONITORING side only — observe/analyze/alert/trip-breaker/trigger-pre-vetted-recovery. Hermes NEVER gets execution privilege, NEVER widens the allowlist, NEVER runs arbitrary Tier-3+. If Hermes could command the executor freely it becomes the soft underbelly (compromise Hermes = escalate). Keeps automated dual-control intact: Agent-Sudo (has root, dumb about judgment) + Hermes (no root, smart about judgment) check each other. Build order: Agent-Sudo v1 ships with a simple self-contained rule-based circuit-breaker; Hermes later augments/replaces it as the intelligent watchdog. Ties to project_hermes_exploration_day (capability-map after id=138).

OPEN GAPS (resume grill-me here)

  1. RESOLVED — SUDO.md classifier (see D4 tiers, D5 engine/seed/loop, D6 principle).
  2. RESOLVED — LXC/Incus specifics (see D7).
  3. RESOLVED — Cutover & rollback (D8).
  4. DROPPED — Bitwarden ephemeral pull no longer in design (D9: Vault AppRole replaces it; no OS-password, no Bitwarden). [Security model fully resolved: D9 + D10.]

D8 — LOCKED (2026-07-08): cutover sequence & rollback

Drop-in replacement: response contract stays identical (POST /exec Bearer → {exit_code, output, ...}; T4 refuse = 403 like today's allowlist-miss). Callers need zero changes; T2/T3 just add latency (sandbox/snapshot) transparently. D4 removes the tier-2 NTFY-approval block entirely.

  • Deploy order = server-01 FIRST, primary LAST (user confirmed). server-01 = testing host, low blast radius, and hosts the sandbox LXCs (T2 local). Primary only after a clean 4-level pass on server-01.
  • Parallel run via ports: Agent-Sudo stands up on new port 8084 (⚠️ corrected 2026-07-08 from 8083 — 8083 is bitwarden-bridge) while old bridge keeps serving 8082 untouched. Shadow-validate (4-level gate on 8083 + replay known-safe commands through both, diff results). Callers stay on 8082 during validation → zero disruption.
  • Atomic swap when green: flip Agent-Sudo to 8082, demote old bridge to 8083 hot-standby (kept running, not deleted).
  • Rollback = swap ports back (old bridge one flip from resuming 8082). Retain old bridge container + image 14 days before decommission (P5).
  • Primary cutover = MANUAL port-swap by user in an interactive session (user confirmed). Agent-Sudo is built + validated autonomously everywhere; the final "make it live on production" flip on primary is user's hands on keyboard (per D3/D6). server-01 cutover = autonomous.
  1. Bitwarden ephemeral sudo password pull — confirm path via bitwarden-bridge; never stored at rest.
  2. RESOLVED (2026-07-08) — Cross-cutting:
    • Per-phase rollback/abort — matrix baked into the P0P5 prompts (P0 drop-table / P1 incus-delete / P2 delete-repo / P3 revoke-policy / P4 port-flip-back / P5 re-pull-image). No phase is destructive to the OLD bridge until P5.
    • Vault/N8N/Coolify dry-run paths — Vault = additive agent-sudo policy + revoke-self smoke test (least-priv proof); N8N = ZERO surface (not a caller); Coolify = retired, Agent-Sudo deploys OUTSIDE it via Jenkins, P5 = plain stop+rm.
    • Training-data logging — 6 decision points → command_audit rows; GOLD label = AI-eval-proposal → human review verdict (human_verdict). P0 schema carries the fields.
    • Blast radius per tier — 0 minimal→3 medium/large→4 zero; capped on primary by shifting the dangerous middle up to refuse; Tier-3-on-primary = highest-attention path (most tests + dual-concurrence + tightest breaker). Full-compromise ⇒ inherits only already-vetted T03, cannot widen allowlist or reach T4.
    • Blocks/unblocks — id=176 has NO blocker (buildable now); no DB row lists it as a blocker (link is architectural, not recorded). Unblocks: Hermes self-healing (id=138), the sandbox-testing program (it's the first test of the harness), and any Horizon-B autonomous automation needing unattended privileged exec. Does NOT unblock the 12 broken N8N workflows (those need the Jenkins/Coolify-API migration).

AFTER GRILL-ME COMPLETES

Write the phased background-agent prompt SET (NOT one monolith — see reasoning in session):

  • P0: DB foundation (command_audit in projects + ensure agent_test_results in api_business) — design-complete, no server risk.
  • P1: LXC/Incus setup + command-surface mirror build on server-01.
  • P2: Python service = drop-in sudo-bridge replacement (same API), classifier reads SUDO.md, audit → command_audit.
  • P3: Bitwarden ephemeral pull + Timeshift automation + AI evaluator (tier-2/3 flow).
  • P4: deploy both servers behind old bridge, run 4-level gate.
  • P5: decommission old sudo-bridge.

Opus+high prompt-craft upgrades over the Sonnet-era prompts (agent_prompts.md): bake in the 4-level test gate + agent_test_results logging (Sonnet's omitted it); split locked-contract from reference-implementation instead of embedding brittle literal scripts; right-size per-phase turn budgets; explicit "design is locked — do not redesign" block; per-phase rollback/abort criteria + idempotent pre-flight checks.