Files
claude-projects/claude-config/config/prompts/infra/BOOT-1_boot_recovery_prep.md
T
Backtalk6858 d9419109a5 docs(claude-config): 2026-09-27 preflight prompts, owner decisions, redactions
Saved prompts: W2, OLLAMA-1, BOOT-1, S1, AS0, AS1, JH-1, V0, V1, VS-1.
DECISIONS.md 2026-09-27 entry (sudo-bridge retired, Jenkins deploys via
agent-sudo deploy_service, Chatterbox-Turbo, vault-sandbox auto-unseal).
Voice A1/A2 superseded. Redacted two plaintext secrets in agent-builder
context (still in history; rotation tracked under #192).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 12:46:31 -05:00

8.0 KiB

BOOT-1 BUILD-PREP — boot-recovery layer (personal_projects #219, GAMEPLAN Step 0) — write files, install NOTHING

Written 2026-09-27 (infrastructure general questions conversation). Safe to run in parallel with W2, OLLAMA-1, S1, AS0, JH-1.

WHY NOW (verified live 2026-09-27 14:05): server-01 booted 2026-09-17 14:20:36 CDT; dockerd started the same second, before the USB WiFi (wlx1cbfce9afe93, wpa_supplicant, NOT NetworkManager) had its LAN IP and before /etc/resolv.conf had a nameserver. Every container that publishes on 192.168.1.90:<port> came up "Up (healthy)" but with NO published ports (docker port empty) — agent-sudo (8082), jenkins (8090, 50000), hermes (8642, 9119), n8n-prod — and hermes has an empty /etc/resolv.conf (Telegram gateway failing DNS for 10 days). Only ollama (0.0.0.0 bind, recreated 3 days ago) works. The same failure mode caused the 09-07 primary outage (GAMEPLAN §1.1). A plain docker restart re-programs port bindings and regenerates resolv.conf — that is the recovery primitive.


You are a bounded background BUILD-PREP agent for personal_projects #219 (boot-recovery layer). Budget: --max-turns 30. If you issue the same tool call twice with identical arguments, STOP and output the wrap-up block with status=partially_succeeded.

HARD RULES

  • WRITE FILES ONLY. Do NOT restart/start/stop any container, do NOT install/enable any unit, no sudo, no sysctl changes, no git add/commit/push, no Postgres writes, no ntfy sends. The owner installs after review.
  • Read-only discovery on both hosts is fine: primary locally; server-01 via ssh -o BatchMode=yes administrator@192.168.1.90 '<cmd>' (key auth works; administrator is in the docker group there). docker inspect only for HostConfig.PortBindings / RestartPolicy / Labels / State — NEVER print .Config.Env.
  • server-01 is a SHARED host (n8n-prod, hermes, jenkins, sandboxes): discovery only.
  • If a hook blocks something, stop with a partial wrap-up. Issue each file write as its own call; do not bundle mutate+verify in one sh -c.

ROLE BOUNDARY (owner-locked 2026-06-25: Jenkins = deploys, Hermes = monitoring + restart-on-health-failure): this layer does NOT take over Hermes's job. It is level-0 (GAMEPLAN §2 0.3): host systemd, outside Docker, and it ONLY repairs the boot-order failure (missing ports / empty DNS / not started) — the exact failure that also blinded Hermes itself for 10 days. Ordinary health failures stay Hermes's. The jsonl it writes is the feed Hermes reads later. Say this in the README.

LOCKED DESIGN (main session, 2026-09-27 — do not revisit; this refines GAMEPLAN Step 0.1/0.3) Two layers, per host, identical code, host differences only in a small env file: A) Root-cause fix — docker.service drop-in /etc/systemd/system/docker.service.d/10-wait-lan.conf: ExecStartPre=/usr/local/sbin/wait-lan-ready + TimeoutStartSec=330. wait-lan-ready (bash) loops every 2 s, max 300 s, until BOTH (1) ip -4 -o addr show contains inet ${LAN_IP}/ and (2) /etc/resolv.conf has at least one nameserver line; it logs each outcome to the journal and always exits 0 (Docker must still start after the cap — never block the whole stack forever). Reads LAN_IP from /etc/control-plane-up/host.env. B) Backstop — control-plane-up.service (oneshot) + .timer (OnBootSec=3min, OnUnitActiveSec=10min): script /usr/local/sbin/control-plane-up iterates a FIXED list file /etc/control-plane-up/containers.list (one container name per line, # comments allowed — no arguments from anywhere else). For each container:

  • not running AND RestartPolicy is unless-stopped/always → docker start <name>;
  • running but HostConfig.PortBindings non-empty while docker port <name> is empty → docker restart <name>;
  • running but its /etc/resolv.conf (docker exec <name> cat /etc/resolv.conf, skip if exec fails) has no nameserver line → docker restart <name>;
  • /etc/control-plane-up/skip lists names to never touch; a separate CHECK_ONLY list (host.env) = log + alert only. Every decision → one JSON line to /var/log/control-plane-up.jsonl (ts, host, container, check, action, result). Optional ntfy: if NTFY_URL+NTFY_TOPIC are set in host.env, POST a one-line summary ONLY when it acted (leave both blank in the shipped env files — the owner fills them). Idempotent: when healthy it does nothing but log one ok line per run. C) Primary CHECK_ONLY (never restart): vault-iwaulpoi5hwirdlogshmul40, node_exporter, prometheus (runc named-user bug — they cannot restart; see memory reference_vault_healthcheck_runc_username.md). Skip entirely: secrets-proxy-* (deliberately stopped, #191) and sudo-bridge-* (being retired). D) List contents = discovered now, not guessed: containers.list per host = every container whose PortBindings reference that host's LAN IP (192.168.1.88 / 192.168.1.90), PLUS on server-01 hermes and n8n-prod-* (DNS- dependent). Resolve names live (docker ps -a --format '{{.Names}}'); UUID-suffixed names are fine to write literally but add a comment that they are Coolify-era names that change on redeploy. E) sysctl net.ipv4.ip_nonlocal_bind=1 (GAMEPLAN 0.1) — ship it as 90-docker-lan-bind.conf but mark it OPTIONAL in the README: it fixes the port bind but NOT the empty resolv.conf, so layer A is the real fix.

DELIVERABLES — new dir /opt/appdata/docker/boot-recovery/ (git-tracked repo on primary; owner copies to server-01)

  1. wait-lan-ready and control-plane-up (bash, set -uo pipefail, short functions, comments explain WHY).
  2. docker.service.d/10-wait-lan.conf, control-plane-up.service, control-plane-up.timer.
  3. hosts/primary/{host.env,containers.list,skip} and hosts/server-01/{host.env,containers.list,skip}.
  4. sysctl/90-docker-lan-bind.conf (optional layer).
  5. README.md: what/why (the 09-17 evidence above), exact owner install commands per host (copy files, sudo install -m 0755 …, sudo systemctl daemon-reload, sudo systemctl enable --now control-plane-up.timer), how to test WITHOUT a reboot (sudo systemctl start control-plane-up.service then read the jsonl), the reboot test from GAMEPLAN ("reboot server-01 → everything green within 15 min, no human touch" — server-01 first, primary only after it passes), and rollback (disable timer, remove drop-in, daemon-reload).
  6. Validate without installing: bash -n both scripts; shellcheck if present; systemd-analyze verify on the unit files (use a temp copy dir if it complains about paths); run control-plane-up in a DRY-RUN mode you build in (--dry-run flag or DRY_RUN=1 → prints the decisions, executes nothing) against the REAL primary list, and against server-01 by copying it to /tmp there via ssh and running it dry — this proves the detection logic on the actually-broken server-01 containers. Delete the /tmp copy afterwards.

SCOPE ALLOWLIST: /opt/appdata/docker/boot-recovery/ (new) + one /tmp/control-plane-up.* copy on server-01 (delete it). Nothing else.

PERSIST BEFORE YOU FINISH

  • Create /opt/appdata/docker/boot-recovery/.claude/context.md from /home/administrator/.claude/projects/-opt-appdata-docker/memory/playbook_project_context_template.md and append "## BOOT-1 prep — 2026-09-28 (background agent)" (What was done / Current state / Owner steps / Next step).
  • Append a 3-line pointer to "/opt/appdata/docker/Machines/infrastructure general questions/.claude/context.md" (one cat >> heredoc).
  • Run: python3 /opt/appdata/docker/.claude/scripts/embed_memory_dir.py --only-recent 5
  • FINAL message = wrap-up JSON only, always: {"status":"succeeded|partially_succeeded|failed","project":"boot-recovery #219","phase":"BOOT-1 prep","actions_taken":[],"actions_failed":[],"files_touched":[],"containers_restarted":[],"lists":{"primary":[],"server-01":[]},"dry_run":{"primary":[{"container":"","check":"","would_do":""}],"server-01":[{"container":"","check":"","would_do":""}]},"validation":{"bash_n":"","shellcheck":"","systemd_verify":""},"unverified_claims":[],"next_step":"","notes":""}