docs(prompts): ASH agent-sudo hardening prep + W4 WireGuard build prep (spawned 2026-09-30)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
Backtalk6858
2026-09-30 23:57:35 -05:00
parent 724e51e3db
commit b89f54558f
2 changed files with 176 additions and 0 deletions
@@ -0,0 +1,89 @@
# ASH — agent-sudo hardening pass (build prep, files only) — #176 / #187
--max-turns 15. You are one agent of a batch; the main session reports all results together after the whole batch
finishes — do not address the owner directly in your final message; write for the main session.
If you issue the same tool call twice with identical arguments, STOP and output the wrap-up block with status=partially_succeeded.
## Goal
Stage the one-pass "hardening" so the owner can apply it on BOTH hosts in one sitting with copy-paste commands:
root-owned daemon code (no more root executing administrator-writable files), a copied (not symlinked) host-neutral
systemd unit with Vault-wait + backoff, the daemon state dir, and the promotion + signing of the already-decided
`bridge/SUDO.md.proposed`. Plus one small code fix (Part B). FILES ONLY: you install nothing, restart nothing.
## Background (verified 2026-09-30 by the main session)
- Repo: `/opt/appdata/docker/docker-compose/agent-sudo/` (same path on server-01 192.168.1.90, but server-01's copy is
OLDER: it lacks the uncommitted AS1/AS2 code). Tests: `python3 -m pytest -q -p no:cacheprovider` → 178 passed.
- Daemon = `sudo_bridge_daemon.py`, imports `sudo_rules`, `daemon_ops` (AS1), and from `security/`: `sudo_sign`,
`scoped_undo`, `timeshift` (check `sys.path` handling in the file — the installed layout must keep those imports
working). `daemon_ops.UNDO_DIR` defaults to `/var/lib/agent-sudo-daemon/undo` (env `AGENT_SUDO_UNDO_DIR`).
`/var/lib/agent-sudo` (0700 root) is a DIFFERENT dir: the container's breaker log bind — do not reuse it.
- Live unit on BOTH hosts: `/etc/systemd/system/agent-sudo-daemon.service` is a SYMLINK (root-owned link) to the repo
file `agent-sudo-daemon.service` (administrator-owned). Primary's repo copy: `SERVER_ID=primary`; server-01's:
`SERVER_ID=server-01` + `VAULT_ADDR=http://192.168.1.88:8200`. Both: `SUDO_MD_VERIFY_ENFORCE=true`,
`SUDO_MD_PATH=<repo>/bridge/SUDO.md`, `SOCKET_PATH=<repo>/bridge/sudo-bridge.sock` (the container mounts
`<repo>/bridge` — keep the socket there), `ExecStart=/usr/bin/python3 <repo>/sudo_bridge_daemon.py`,
`After=network.target`, `Restart=always`, `RestartSec=5`.
- **DO NOT edit the live repo file `agent-sudo-daemon.service`** — it is the symlink target; any edit is live at the
next restart. Stage everything under `deploy/hardening/`.
- Daemon is fail-closed: `_verify_sudo_md_or_die()` exits 3 when the Vault signature check fails; today that busy-loops
(primary ≈53k journal lines since 09-01; server-01 restarts while primary's Vault is down).
- Decided design (readiness report `/opt/appdata/docker/research/agent_sudo_readiness_2026-09-28.md`, "Monday build
order" step 1): host-neutral `EnvironmentFile=/etc/agent-sudo/daemon.env`; `After=network-online.target
docker.service` + `Wants=network-online.target`; `RestartSec=30`; `StartLimitIntervalSec=0`; `ExecStartPre`
Vault-wait (poll `<VAULT_ADDR>/v1/sys/health` until unsealed, 5-minute cap, then exit non-zero); code installed
root-owned at `/usr/local/lib/agent-sudo/` (dirs 0755, files 0644, root:root) and the unit COPIED to
`/etc/systemd/system/`. Add `StateDirectory=agent-sudo-daemon` + `StateDirectoryMode=0700`. Do NOT add
sandboxing directives (ProtectSystem, NoNewPrivileges, …): this daemon's job is root mutations.
- AS2 owner decisions DONE 2026-09-30: `bridge/SUDO.md.proposed` is final (resolver start = `primary=4`). Promotion =
copy over `bridge/SUDO.md`, then sign: `VAULT_TOKEN=<privileged> python3 security/sudo_sign.py sign bridge/SUDO.md`
and `... digest ...` to check. The privileged token = Vault root token, which the owner retrieves from Bitwarden
(never written to disk, never echoed; runbook uses `read -rs VAULT_TOKEN` then `export`, then `unset`). Signing must
happen BEFORE the restart or the enforced gate refuses to start.
- Rollout order matters (AS1): new daemon first, container image later (AS3). A new container against an old daemon
returns 503 for tiers 1–3.
- Open finding to fix (Part B): on server-01 the `verb_heuristic` in `sudo_rules.py` gives tier 0 to read verbs
(cat/head/tail/grep/less/…) for ANY path, e.g. `head /etc/shadow`. Fix: a read verb gets tier 0 only when no
operand matches a sensitive-path list (at least `/etc/shadow`, `/etc/gshadow`, `/etc/sudoers*`, `/root/`, `*/.ssh/`,
`/etc/ssl/private/`, `/etc/wireguard/`, `*agent-sudo*`, `*approle*`, `*.env`); otherwise tier 4. Read the function
first and keep its structure; add tests in a NEW file `test_ash_heuristic.py`.
## Scope allowlist
- CREATE `deploy/hardening/` in the repo: `agent-sudo-daemon.service`, `daemon.env.primary`, `daemon.env.server-01`,
`vault-wait.sh`, `install.sh` (idempotent; copies code root-owned, writes /etc/agent-sudo/daemon.env from the right
host file, replaces the symlink with a copy, daemon-reload; does NOT restart), `verify.sh` (read-only post-checks),
`rollback.sh` (restore the symlink unit + restart instructions), `HARDENING_RUNBOOK.md`.
- EDIT `sudo_rules.py` (verb_heuristic only). CREATE `test_ash_heuristic.py`. APPEND `.claude/context.md`.
- Nothing else. No sudo, no systemctl mutations, no docker restarts, no Vault writes, no SUDO.md edits, no commits.
## Steps
1. Read: the daemon + imports, `daemon_ops.py`, `sudo_rules.verb_heuristic`, both hosts' current unit
(`systemctl cat agent-sudo-daemon`; server-01 over `ssh administrator@192.168.1.90`), `AS1_CHANGES.md`,
`AS2_CHANGES.md`, `DEPLOY_RUNBOOK.md`.
2. Part B code fix + tests; run the FULL suite; it must stay green (≥ 178 + your new tests).
3. Stage the files. `install.sh` takes the host name as `$1` (primary|server-01), refuses anything else, uses
`install -o root -g root`, and lists exactly which files it copies. On server-01 the source is a bundle the main
session rsyncs from primary to `/tmp/agent-sudo-hardening/` (write that rsync command in the runbook as a
main-session step, with a sha256 manifest check).
4. VALIDATE with the real tools: `systemd-analyze verify deploy/hardening/agent-sudo-daemon.service` (paths that
don't exist yet may warn — record the output and explain each line); `bash -n` and `shellcheck` (if installed) on
every script; run `vault-wait.sh` once against primary's real Vault read-only (`/v1/sys/health`) and record the
result; import-test the installed layout by copying the code to a scratch dir with the same structure and running
`python3 -c "import sys; sys.path[:0]=['<dir>','<dir>/security']; import sudo_bridge_daemon"`-style checks WITHOUT
starting the daemon (read its `__main__` guard first).
5. HARDENING_RUNBOOK.md — commands IN FULL, no `…`, one command per block, why + expected output for each:
(0) preflight (current state, take a copy of the live unit target); (1) promote SUDO.md.proposed → SUDO.md and
sign + digest-check (root token via `read -rs`); (2) primary: install.sh primary → verify.sh → ONE restart →
status + journal check (expect "SUDO.md signature OK"-style line — read the daemon for the real log text) → a
tier-0 `/exec` smoke test through the container is NOT possible yet (old image) — say what IS testable; (3)
server-01: main-session rsync bundle → install.sh server-01 → verify → restart → check; (4) rollback per host.
6. Wrap-up.
## Wrap-up (always)
Append the dated handoff to the repo's `.claude/context.md` (What was done / Decisions / Current state / Next step).
Run `python3 /opt/appdata/docker/.claude/scripts/embed_memory_dir.py --only-recent 3`. Never print secrets or
`.Config.Env`; mask git remote URLs with `sed -E 's#//[^@]*@#//<cred>@#'` if you ever look at them. Final message = JSON:
```
{"status":"succeeded|partially_succeeded|failed","project":"agent-sudo hardening prep (ASH)","actions_taken":[],
"actions_failed":[],"files_touched":[],"containers_restarted":[],"tests":{"before":178,"after":0,"failed":0},
"validation":{},"owner_steps_count":0,"risks":[],"unverified":[],"next_step":"","notes":""}
```
@@ -0,0 +1,87 @@
# W4 — per-client ACLs + DOCKER-USER backstop + health/ntfy (build prep, files only) — #274
--max-turns 15. You are one agent of a batch; the main session reports all results together after the whole batch
finishes — do not address the owner directly in your final message; write for the main session.
If you issue the same tool call twice with identical arguments, STOP and output the wrap-up block with status=partially_succeeded.
## Goal
Stage W4 so the owner can apply it in one sitting AFTER W3: each tunnel client reaches only what it needs, a host
firewall backstop limits the whole VPN even if wg-easy's per-client filter is wrong, and a health timer alerts via
ntfy. FILES ONLY. Design (LOCKED): `/opt/appdata/docker/research/diy_wireguard_design.md` — read §Q6 (access
control), §Q8 (failure modes + health), §5 W4 row, §6 risk 4. W3 is STAGED, NOT applied: read
`/opt/appdata/docker/docker-compose/wireguard/W3_RUNBOOK.md` and `deploy/w3/` — W4 assumes W3's end state.
## Background (verified 2026-09-29/30 by the main session)
- wg-easy 15.4.0 on primary, container `wg-easy`, docker network `wg` 10.42.42.0/24, gateway 10.42.42.1 (host),
wg-easy 10.42.42.42 (NATs all clients → the host sees them as 10.42.42.42), tunnel 10.8.0.0/24, UDP 45791.
Clients: laptop 10.8.0.2, phone 10.8.0.3, tablet 10.8.0.4. Client AllowedIPs `10.8.0.0/24,10.42.42.0/24`.
The admin UI answers at http://10.8.0.1:51821 from the tunnel (container `HOST=0.0.0.0`, `INSECURE=true`).
- After W3: dnsmasq `wg-dnsmasq` at 10.42.42.53:53; Traefik VPN listener 10.42.42.1:443 (`https-vpn` entrypoint).
- wg-easy v15.3+ has per-client server-side firewall filtering (iptables in the container namespace). Find the exact
field/format by reading the container's code (`docker exec wg-easy sh -c 'grep -rn -i firewall /app/server | head'`,
read-only) and WebFetch the wg-easy docs/release notes for it; cite URLs. The admin login is Bitwarden + TOTP (Ente):
you CANNOT log in to the UI/API — write per-client filters as exact owner UI steps.
- **ACL table — pre-decided (design §Q6 + owner directive 2026-09-30 #288: the PHONE must be able to SSH via Termius
to primary and server-01 for owner escalations):**
| client | allowed |
|---|---|
| phone | 10.42.42.1:443/tcp, 10.42.42.53:53/udp+tcp, 192.168.1.88:22/tcp, 192.168.1.90:22/tcp |
| tablet | 10.42.42.1:443/tcp, 10.42.42.53:53/udp+tcp |
| laptop | phone set + wg-easy UI 51821 |
| anything else | deny |
Everything else on 192.168.1.0/24 and 172.16.0.0/12 (Postgres 5432, Vault 8200, Samba, RustDesk, sudo-bridge) is
unreachable from the VPN. Work out — and state in the runbook — which path each destination takes on the HOST
(192.168.1.88/10.42.42.1 = INPUT chain, not DOCKER-USER; 192.168.1.90 = FORWARD → DOCKER-USER) and design the
backstop accordingly (DOCKER-USER for forwarded traffic; for host-local ports decide INPUT rules scoped to source
10.42.42.0/24 on the `wg` bridge interface, or explain why wg-easy's per-client filter alone is acceptable there).
- SSH: primary's `/etc/ssh/sshd_config` has `#PasswordAuthentication yes` (commented = default = password auth ON).
Check `/etc/ssh/sshd_config.d/` and server-01's sshd config the same way (read-only). Opening SSH to the phone makes
key-only auth important. Stage (as LAST, separate, optional owner steps): Termius key generation on the phone →
append its public key to `~/.ssh/authorized_keys` on both hosts → prove key login from phone AND laptop → only then a
drop-in `/etc/ssh/sshd_config.d/10-keys-only.conf` (`PasswordAuthentication no`, `KbdInteractiveAuthentication no`)
validated with `sudo sshd -t` before `sudo systemctl reload ssh`. Warn loudly about lock-out and keep a console path.
- DOCKER-USER mistakes can cut container networking (design risk 4): every firewall apply step must have a timed
auto-revert (check whether `at` is installed; otherwise `systemd-run --on-active=5min <flush command>`), and the
owner cancels the revert only after the tests pass.
- Health (design §Q8): host script on a 5-min systemd timer: container healthy; `wg show` listen port inside the
container; `dig +short fuppzd3f9v.reverseproxyserver.net @1.1.1.1` == current public IP (from
`https://1.1.1.1/cdn-cgi/trace`); cert days left on 10.42.42.1:443 (alert ≤ 21). Alerts → ntfy, dedicated topic +
bot user (read `/home/administrator/.claude/projects/-opt-appdata-docker/memory/playbook_ntfy_users_topics.md` for
the topic/user/token procedure; creating the user and token is an OWNER/main-session step in the runbook; the
token goes to Vault `secret/ntfy/wireguard-health` field `token`). Alert only on change of state (no 5-min spam);
normal priority (3), not urgent.
- The wireguard project has `wireguard-secretspec-resolver.timer` running `compose up -d` on the LIVE compose every
15 min — never edit live files; stage in `docker-compose/wireguard/deploy/w4/`. Installed secretspec accepts ONLY
`revision = "1.0"`.
## Scope allowlist
- CREATE `docker-compose/wireguard/deploy/w4/` (firewall rules file(s), apply/revert scripts, systemd unit(s) for the
backstop and the health timer, health script, per-client filter values file) and `docker-compose/wireguard/W4_RUNBOOK.md`.
- APPEND `docker-compose/wireguard/.claude/context.md`.
- No live edits, no container starts/stops, no firewall changes, no sudo, no wg-easy UI/API writes, no commits.
## Steps
1. Read the design sections, W3 runbook + staged files, live compose, `docker network inspect wg` (no Env), current
`iptables -S DOCKER-USER` only if readable without sudo (otherwise list it as an owner preflight command), sshd
configs, the ntfy playbook, wg-easy firewall code + docs (WebFetch, cite).
2. Stage the files.
3. VALIDATE with real tools: `bash -n` + `shellcheck` (if installed) on scripts; `systemd-analyze verify` on units;
`iptables-restore --test` / `nft -c -f` on the rules file if the binary runs without root (else record it as an
owner preflight step); run the health script once in a no-alert dry-run mode against the live system (read-only)
and record its output.
4. W4_RUNBOOK.md — commands IN FULL, no `…`, one command per block, why + expected output: preflight → ntfy bot/topic
+ Vault token (main session) → per-client filters in the UI (exact values per client) → backstop apply with timed
auto-revert → tests (from the phone: Termius SSH to 192.168.1.88 and 192.168.1.90 works, https grafana works,
192.168.1.88:5432 and :8200 time out; from the tablet: SSH times out, https works; from the laptop: SSH + UI work)
→ cancel the revert → enable the persistent unit → health timer → optional SSH key-only steps → rollback.
5. Wrap-up.
## Wrap-up (always)
Append the dated handoff to `docker-compose/wireguard/.claude/context.md`. Run
`python3 /opt/appdata/docker/.claude/scripts/embed_memory_dir.py --only-recent 3`. Never print secrets or
`.Config.Env`. Final message = JSON:
```
{"status":"succeeded|partially_succeeded|failed","project":"#274 W4 build prep","actions_taken":[],"actions_failed":[],
"files_touched":[],"containers_restarted":[],"validation":{},"docs_fetched":[],"acl_paths":{},"owner_decisions":[],
"risks":[],"unverified":[],"next_step":"","notes":""}
```