From d9419109a518dd74d3d0c73fb1e4b314a19bf887 Mon Sep 17 00:00:00 2001 From: Backtalk6858 Date: Mon, 28 Sep 2026 12:46:31 -0500 Subject: [PATCH] docs(claude-config): 2026-09-27 preflight prompts, owner decisions, redactions Saved prompts: W2, OLLAMA-1, BOOT-1, S1, AS0, AS1, JH-1, V0, V1, VS-1. DECISIONS.md 2026-09-27 entry (sudo-bridge retired, Jenkins deploys via agent-sudo deploy_service, Chatterbox-Turbo, vault-sandbox auto-unseal). Voice A1/A2 superseded. Redacted two plaintext secrets in agent-builder context (still in history; rotation tracked under #192). Co-Authored-By: Claude Opus 5.5 --- agent-builder/.claude/context.md | 12 +++- agent-builder/voice_chat_agent_prompts.md | 6 ++ claude-config/config/prompts/README.md | 16 +++++ ...do_readiness_and_sudo_bridge_retirement.md | 46 ++++++++++++ .../AS1_move_privileged_ops_into_daemon.md | 52 ++++++++++++++ .../S1_secrets_proxy_investigation.md | 40 +++++++++++ .../infra/BOOT-1_boot_recovery_prep.md | 70 +++++++++++++++++++ .../infra/JH-1_jenkins_hermes_state_audit.md | 38 ++++++++++ .../OLLAMA-1_single_instance_server01.md | 42 +++++++++++ .../VS-1_vault_sandbox_autounseal_prep.md | 42 +++++++++++ .../prompts/voice/V0_tts_engine_research.md | 47 +++++++++++++ .../prompts/voice/V1_stt_server01_gpu.md | 56 +++++++++++++++ .../config/prompts/wireguard/W2_build_prep.md | 45 ++++++++++++ claude-config/decisions/DECISIONS.md | 16 +++++ claude-config/research/INDEX.md | 2 + 15 files changed, 528 insertions(+), 2 deletions(-) create mode 100644 claude-config/config/prompts/agent-sudo-secrets-proxy/AS0_agent_sudo_readiness_and_sudo_bridge_retirement.md create mode 100644 claude-config/config/prompts/agent-sudo-secrets-proxy/AS1_move_privileged_ops_into_daemon.md create mode 100644 claude-config/config/prompts/agent-sudo-secrets-proxy/S1_secrets_proxy_investigation.md create mode 100644 claude-config/config/prompts/infra/BOOT-1_boot_recovery_prep.md create mode 100644 claude-config/config/prompts/infra/JH-1_jenkins_hermes_state_audit.md create mode 100644 claude-config/config/prompts/infra/OLLAMA-1_single_instance_server01.md create mode 100644 claude-config/config/prompts/infra/VS-1_vault_sandbox_autounseal_prep.md create mode 100644 claude-config/config/prompts/voice/V0_tts_engine_research.md create mode 100644 claude-config/config/prompts/voice/V1_stt_server01_gpu.md create mode 100644 claude-config/config/prompts/wireguard/W2_build_prep.md diff --git a/agent-builder/.claude/context.md b/agent-builder/.claude/context.md index fae5f93..b27f554 100644 --- a/agent-builder/.claude/context.md +++ b/agent-builder/.claude/context.md @@ -284,7 +284,7 @@ Coolify service UUID: `ilus0cfdkheipodw1viurg1d` - NTFY bug: _ntfy() was hardcoding URL to secrets-proxy-notifications regardless of topic — fixed with topic param - NTFY ACL: secrets-proxy-bot needed write-only on secrets-proxy-approvals; Smoked5003 needed read-only — both added via ntfy CLI in container - secrets_proxy DB user: needed SELECT on proxy_executions for ON CONFLICT DO NOTHING — granted -- Coolify token plaintext: 2|95eQySElT9uQpTXqDACWq1z9kyOaySZOZP8sBxyJaebc2bbe (Sanctum format: id|plaintext) +- Coolify token plaintext: (Sanctum format: id|plaintext) - Coolify "Service is already running" after stop+rm: UPDATE service_applications SET status='stopped' WHERE service_id=... then POST /start - 524 on destructive commands: Cloudflare kills long-poll at ~100s; approval must be tapped quickly; command still runs if approved before timeout, response just lost @@ -318,7 +318,7 @@ Coolify service UUID: `ilus0cfdkheipodw1viurg1d` - app.py needs: server_id='server-01' + [server-01] NTFY prefix variant for server-01 deployment - sudo-bridge image: gitea.local/backtalk6858/sudo-bridge:latest (same image, different env vars) - sudo-bridge-server01 Coolify UUID: o2kz1puml1mmneyiqd96mouj (port 8082, pull_policy:never) -- Vault secret created: secret/data/sudo-bridge-server01 (api_key: 443ea35b43c3e640aac3c57ed3aae06b8822ab6f) +- Vault secret created: secret/data/sudo-bridge-server01 (api_key: ) - Host daemon running: /opt/appdata/docker/sudo-bridge/sudo_bridge_daemon.py (systemd, enabled) - vaultwarden-sandbox REMOVED from sandbox stack (no longer needed — Bitwarden cloud dummy account used) - app.py updated: SERVER_ID, _server_prefix(), server_id audit writes, User-Agent for Cloudflare @@ -653,3 +653,11 @@ Remaining sprint-day-11 priorities were: (1) fix secrets-proxy Dockerfile compos ## Update instructions Update at the end of every agent-builder session. Keep agent status, key decisions, and prereq checklist current. + +## V0 TTS engine research — 2026-09-28 (background agent) +- Verdict: switch primary TTS from Coqui XTTS-v2 to **Chatterbox-Turbo** (Resemble AI, 350M, MIT code+weights, Perth watermark) served by devnen/Chatterbox-TTS-Server (build CUDA 12.1 variant for driver 550; leave TTS_BF16 off on Turing; `/tts` stream:true, `/v1/audio/speech`, `/api/unload`). +- Runner-up / fallback: XTTS-v2 via idiap coqui-tts 0.27.5 (2026-01-26) — works, ~4 GB, streams, but CPML non-commercial + Coqui defunct; needs own GPU image + streaming wrapper. +- Watch: Qwen3-TTS 0.6B (Apache, 3 s clone, 97 ms) — bf16/FA2 recommended, untested on Turing. +- VRAM: all four (Ollama 5 GB + Whisper + TTS + nomic) do NOT fit 8 GB; options = evict Ollama during voice (keep_alive:0), TTS /api/unload on idle, smaller/int8 Whisper. +- Changes July decision: YES. Next: V2 prompt must smoke-test Chatterbox-Turbo peak VRAM + time-to-first-audio on the RTX 2060 SUPER. +- Report: /opt/appdata/docker/research/voice_tts_engine_2026-09-28.md (INDEX row added in claude-config/research/INDEX.md). diff --git a/agent-builder/voice_chat_agent_prompts.md b/agent-builder/voice_chat_agent_prompts.md index ffb77c9..22f3484 100644 --- a/agent-builder/voice_chat_agent_prompts.md +++ b/agent-builder/voice_chat_agent_prompts.md @@ -2,6 +2,12 @@ # Companion to voice_chat_build_plan.md. Preflight: 2026-07-29. # Launch target: on/before the build session ahead of 2026-08-14. +# ⚠️ 2026-09-27: AGENTS A1 + A2 BELOW ARE SUPERSEDED — they route through "sudo-bridge-server01 :8082", +# which is dead and being retired (agent-sudo replaces sudo-bridge; owner calls sudo-bridge security theatre). +# Use claude-config/config/prompts/voice/V1_stt_server01_gpu.md (STT as a GPU docker container, no sudo). +# A2 (TTS) will be rewritten as V2 once the Jarvis sample exists + the TTS engine is re-confirmed. +# A3 (primary-only Stop-hook script) is still valid as written. +# # # ⚠️ RE-VERIFY AT LAUNCH (infra facts decay — do NOT trust these blindly): # - server-01 access = sudo-bridge-server01 at http://192.168.1.90:8082 (NOT secrets-proxy /shell). diff --git a/claude-config/config/prompts/README.md b/claude-config/config/prompts/README.md index 17d17b7..b5d713c 100644 --- a/claude-config/config/prompts/README.md +++ b/claude-config/config/prompts/README.md @@ -37,3 +37,19 @@ Execution/build/design prompts that belong to an ongoing project set live in a s - `infra/` — NTFY-272 hardening (2026-09-24); #273 image-autoupdate research + P1a pin proof, cloudflared replacement research (2026-09-25). #273 P2 prompt lives verbatim in memory `playbook_image_autoupdate_phases.md`. The 13 files above were recovered from the session transcript on 2026-09-25 because they had been spawned inline without being saved first. - `wireguard/` — #274 home-access tunnel phase prompts: `W1_build_prep.md` (2026-09-25; write files only, no deploy). Design: `/opt/appdata/docker/research/diy_wireguard_design.md`. + +## 2026-09-28 parallel-agent day (written 2026-09-27, infrastructure general questions conversation) +Owner plan: 4–5 background agents at once from Mon 2026-09-28. Spawn with `subagent_type: general-purpose`, paste the file body verbatim, run `agent-wrap` after each. +| Prompt | Tracks | Mutates? | Parallel notes | +|---|---|---|---| +| `wireguard/W2_build_prep.md` | #274 W2 | files only (primary) | any | +| `infra/OLLAMA-1_single_instance_server01.md` | single Ollama on server-01 GPU | edits embed caller(s); owner retires primary Ollama | any | +| `infra/BOOT-1_boot_recovery_prep.md` | #219 boot-recovery (fixes dead server-01 ports/DNS) | files only | any | +| `agent-sudo-secrets-proxy/S1_secrets_proxy_investigation.md` | #191 → #174 | read-only | any | +| `agent-sudo-secrets-proxy/AS0_agent_sudo_readiness_and_sudo_bridge_retirement.md` | #176/#192/#193/#206 + sudo-bridge retirement | read-only (+2 runbooks) | any | +| `infra/JH-1_jenkins_hermes_state_audit.md` | #181/#208 | read-only | after owner restarts jenkins+hermes | +| `voice/V1_stt_server01_gpu.md` | #214 STT (supersedes agent-builder A1) | server-01 GPU container | OWNER GATE (image pull OK); only server-01 mutator | +**sudo-bridge rule (owner 2026-09-27):** no prompt may route through sudo-bridge; it is being retired in favour of agent-sudo. +| `voice/V0_tts_engine_research.md` | #214 TTS engine re-check (XTTS-v2 defunct/non-commercial) | research only | spawned 2026-09-27 | +| `infra/VS-1_vault_sandbox_autounseal_prep.md` | vault-sandbox auto-unseal = copy of primary's (owner 09-27) | files only | needs server-01 up | +| `agent-sudo-secrets-proxy/AS1_move_privileged_ops_into_daemon.md` | #176 AS1: undo/snapshot/sandbox/deploy_service into root daemon, A1, multi-caller auth | code+tests only (primary) | spawned 2026-09-27 17:3x | diff --git a/claude-config/config/prompts/agent-sudo-secrets-proxy/AS0_agent_sudo_readiness_and_sudo_bridge_retirement.md b/claude-config/config/prompts/agent-sudo-secrets-proxy/AS0_agent_sudo_readiness_and_sudo_bridge_retirement.md new file mode 100644 index 0000000..212d777 --- /dev/null +++ b/claude-config/config/prompts/agent-sudo-secrets-proxy/AS0_agent_sudo_readiness_and_sudo_bridge_retirement.md @@ -0,0 +1,46 @@ +# AS0 — agent-sudo readiness audit + key-rotation runbook (#192) + sudo-bridge retirement inventory — READ-ONLY + +Written 2026-09-27 (infrastructure general questions conversation). Tracks personal_projects #176 (tiers 1/2/3), +#187, #192 (exposed key), #193 (fail-open/closed), #166/#206 (snapshot substrate). Owner direction 2026-09-27: +**agent-sudo is the FULL replacement for sudo-bridge** (sudo-bridge = security theatre); nothing may be designed to +depend on sudo-bridge. Safe to run in parallel with S1, W2, OLLAMA-1, BOOT-1, JH-1. + +--- + +You are a bounded, READ-ONLY audit agent for Agent-Sudo. Budget: --max-turns 30. +If you issue the same tool call twice with identical arguments, STOP and output the wrap-up block with status=partially_succeeded. + +HARD RULES +- READ-ONLY. No container start/stop/restart, no unit changes, no sudo, no Vault reads/writes, no git add/commit/push, no Postgres writes, no edits to code/allowlists/SUDO.md. Only writes: the report, two runbooks, a new context.md, the embed run. +- **NEVER print a key/token value.** The exposed server-01 key sits in plaintext in `/home/administrator/Desktop/claude/agent-builder/.claude/context.md` (and git history). Locate occurrences by the VARIABLE NAME / surrounding label only; report file:line and the value's LENGTH and first 4 chars at most, never the value. For git history use `git log -S --oneline` style searches on the label, not the secret. +- server-01 via `ssh -o BatchMode=yes administrator@192.168.1.90 ''` (read-only). `docker inspect` only State/Image/Labels/Mounts/PortBindings/RestartPolicy — NEVER `.Config.Env`. +- If a hook blocks something, stop with a partial wrap-up. + +FACTS (main session, verified 2026-09-27 14:00) +- Primary: `agent-sudo-agent-sudo-1` healthy, `192.168.1.88:8084` → `/health` 200 `{"status":"ok","server_id":"primary"}`; `agent-sudo-daemon.service` active; `SUDO_MD_VERIFY_ENFORCE=true` (signed, D9, #145, since 07-14); compose has `SANDBOX_ENABLED=${…:-false}`, `SNAPSHOT_ENABLED=${…:-false}` → tiers 2/3 return 503. Last command in `bridge/command_audit.jsonl` ≈ 2026-07-15 → effectively unused since July. DB row #176 notes "Blocked on SUDO.md signing" — STALE. +- server-01: `agent-sudo-agent-sudo-1` "Up (healthy)" but its `192.168.1.90:8082` binding is NOT live (`docker port` empty; boot-order bug — BOOT-1 fixes it; the owner restarts it separately). `agent-sudo-daemon.service` active (SERVER_ID=server-01, VAULT_ADDR=http://192.168.1.88:8200, enforce=true). A stale container `sudo-bridge-sudo-bridge-1` sits in `Created` (failed: port 8082 already allocated) and `sudo-bridge-daemon.service` is ALSO active there. +- Primary sudo-bridge is still LIVE: container `sudo-bridge-sudo-bridge-1` healthy on `192.168.1.88:8082` + root `sudo-bridge-daemon.service` active; its `audit.jsonl` last written 2026-07-28 (nothing has used it for 2 months). +- Jenkins (server-01) has jobs `sudo-bridge` (12 builds) and `sudo-bridge-server01` (2 builds). +- Incus on server-01 has two RUNNING sandbox containers `sandbox-primary` (10.85.172.80) and `sandbox-server01` (10.85.172.45); `/data` = 1TB 870 EVO ext4, 854 GB free (row 206 "not readable" is stale; SMART CRC unverified — needs root = owner). +- 27 memory files reference sudo-bridge, including the prompt playbook rule 12 (the main session is fixing that one itself — don't edit). Hook `/opt/appdata/docker/.claude/hooks/security-enforcement.py` references it. + +TASKS +1. **Readiness matrix** for tiers 0–4 on BOTH hosts: what each tier needs (code present? flag? substrate? tests?), status now, and the exact remaining work, mapped to GAMEPLAN task IDs. Read: `/home/administrator/Desktop/claude/claude-config/decisions/GAMEPLAN_security-infra-deploy.md` (Steps 0,1,4), `/home/administrator/Desktop/claude/agent-builder/GAMEPLAN_agent-sudo_secrets-proxy.md` (A1–A7 definitions), memory `playbook_agent_sudo_phases.md`, the code in `/opt/appdata/docker/docker-compose/agent-sudo/` (app.py, sudo_rules.py, sudo_bridge_daemon.py, circuit_breaker.py, security/, scripts/, DEPLOY_RUNBOOK.md). Run the existing test suites read-only if they need no network/root: `cd /opt/appdata/docker/docker-compose/agent-sudo && python3 -m pytest -q test_sudo_rules.py test_circuit_breaker.py` (skip test_app.py if it needs a live Vault; say so). +2. **SUDO.md coverage:** list the active rules per tier (count + examples). Are the commands our upcoming work needs covered at the right tier? Specifically: `systemctl enable|disable|start|stop ` for new units, `install -m 0755 /usr/local/sbin/`, `cp` into `/etc/systemd/system/`, `systemctl daemon-reload`, `docker restart `, `docker compose up -d` in a project dir, `smartctl -a /dev/sda`. (Audit shows `systemctl daemon-reload` was tier-4 refused on 07-15.) Output a PROPOSED rule diff table — do not edit SUDO.md (it is signed; the owner re-signs). +3. **#192 rotation runbook** → `/opt/appdata/docker/docker-compose/agent-sudo/ROTATE_SERVER01_KEY.md`: every consumer of the server-01 caller key (where the daemon/container reads it, Vault path `secret/data/sudo-bridge-server01` per the GAMEPLAN — UNVERIFIED which field, any .env, any injected context/hook, the plaintext context.md occurrence(s) + git history), then exact owner steps: generate (`bw generate -ulns --length 40` pattern or `vault_store.py --generate`), store in Vault, update consumers, restart order, verification (`/health` + an authenticated tier-0 `id` call returns 200 with new key and 401/403 with old), redact the plaintext occurrence (replace with a pointer to the Vault path), and the history decision (rewrite vs accept-and-document — list trade-offs; owner decides). +4. **sudo-bridge retirement inventory + runbook** → `/opt/appdata/docker/docker-compose/sudo-bridge/RETIRE.md`: (a) live components per host (containers, daemons, sockets, sudoers drop-ins — `ls /etc/sudoers.d/` names only, allowlist files, Jenkins jobs, cron/timers), (b) every reference: memory files (list with the one-line context of each mention, classify: rule-that-routes-work-to-it / historical / doc-to-update), hooks, skills, scripts, prompts in `/home/administrator/Desktop/claude/claude-config/config/prompts/` and `/home/administrator/Desktop/claude/agent-builder/*prompts*.md`, (c) any project/design that DEPENDS on sudo-bridge (the voice prompts `agent-builder/voice_chat_agent_prompts.md` are one — they target "sudo-bridge-server01 :8082"), (d) ordered owner retirement steps with rollback, gated on "agent-sudo reachable + rotated key on both hosts". +5. **#193 input:** summarise fail-open vs fail-closed evidence in 5 lines (GAMEPLAN recommends keep fail-closed + backoff); state whether the daemon unit already has the Vault-wait/backoff (GAMEPLAN 0.4 / A4). +6. **Snapshot substrate (#166/#206):** what tier-3 needs on each host (read the code), whether `/data` on server-01 and the Incus sandboxes satisfy it, and the exact owner command to read SMART (`sudo smartctl -a /dev/sda` on server-01) with what CRC value would block it. + +DELIVERABLES +1. Report `/opt/appdata/docker/research/agent_sudo_readiness_2026-09-28.md`: Verdict (5 lines) → tasks 1,2,5,6 → "Monday build order" (a proposal: ordered list of bounded agent phases, each ≤ 1 session, with the owner steps between them) → UNVERIFIED list. +2. The two runbooks above. + +SCOPE ALLOWLIST: the report, `ROTATE_SERVER01_KEY.md`, `sudo-bridge/RETIRE.md`, `/opt/appdata/docker/docker-compose/agent-sudo/.claude/context.md` (new, from template). Nothing else. Leave pytest caches alone (delete any new `.pytest_cache` you create). + +PERSIST BEFORE YOU FINISH +- Create `/opt/appdata/docker/docker-compose/agent-sudo/.claude/context.md` from `/home/administrator/.claude/projects/-opt-appdata-docker/memory/playbook_project_context_template.md`; append "## AS0 readiness audit — 2026-09-28 (background agent)". +- Append a 4-line "## AS0 — 2026-09-28" pointer (report + 2 runbooks) to `/opt/appdata/docker/docker-compose/sudo-bridge/.claude/context.md` (it exists; append only). +- Run: python3 /opt/appdata/docker/.claude/scripts/embed_memory_dir.py --only-recent 5 +- FINAL message = wrap-up JSON only, always: +{"status":"succeeded|partially_succeeded|failed","project":"agent-sudo #176/#192","phase":"AS0","actions_taken":[],"actions_failed":[],"files_touched":[],"containers_restarted":[],"tier_readiness":{"primary":{"t0":"","t1":"","t2":"","t3":""},"server-01":{"t0":"","t1":"","t2":"","t3":""}},"tests":{"sudo_rules":"","circuit_breaker":"","app":""},"key_192_consumers":[],"sudo_bridge_live_components":[],"sudo_bridge_dependents":[],"proposed_sudo_md_rules":0,"unverified_claims":[],"next_step":"","notes":""} diff --git a/claude-config/config/prompts/agent-sudo-secrets-proxy/AS1_move_privileged_ops_into_daemon.md b/claude-config/config/prompts/agent-sudo-secrets-proxy/AS1_move_privileged_ops_into_daemon.md new file mode 100644 index 0000000..e8744b4 --- /dev/null +++ b/claude-config/config/prompts/agent-sudo-secrets-proxy/AS1_move_privileged_ops_into_daemon.md @@ -0,0 +1,52 @@ +# AS1 — move undo / snapshot / sandbox (+ deploy_service) into the ROOT daemon; implement A1 — CODE + TESTS ONLY + +Written 2026-09-27 (infrastructure general questions conversation). Tracks personal_projects #176 (tiers 1/2/3), #187, #166. +Source of truth for WHY: `/opt/appdata/docker/research/agent_sudo_readiness_2026-09-28.md` (AS0) — read "Task 1", +"Task 2", "Task 6", "Other findings" and "Monday build order" item 2 FIRST. +Owner decisions 2026-09-27 (`claude-config/decisions/DECISIONS.md`, 2026-09-27 entry): agent-sudo fully replaces sudo-bridge; +**the Jenkins → primary deploy path = a narrow agent-sudo `deploy_service` op that runs the project's SecretSpec resolver** +(NOT raw `docker compose` through secrets-proxy); fail-closed + backoff (#193); narrow the two broad tier-0 rules. +Safe to run in parallel with anything that does not edit `/opt/appdata/docker/docker-compose/agent-sudo/`. + +--- + +You are a bounded background CODE agent for Agent-Sudo, phase AS1. Budget: --max-turns 40 (code task — raised from 15 on purpose). +If you issue the same tool call twice with identical arguments, STOP and output the wrap-up block with status=partially_succeeded. + +HARD RULES +- CODE + TESTS ONLY, on primary, in `/opt/appdata/docker/docker-compose/agent-sudo/`. Do NOT build images, restart/recreate containers, restart the daemon, edit systemd units on the host, run sudo, touch Vault, commit or push. **Do NOT edit `bridge/SUDO.md`** (signed; editing it would make the live daemon refuse to start on its next restart) — rule changes are AS2. +- Never print secret values; never read `.env` files (the hook blocks it anyway). If a hook blocks something, stop with a partial wrap-up. +- Keep every existing test green. Run tests with host python, isolated paths, no cache dir: `cd /opt/appdata/docker/docker-compose/agent-sudo && python3 -m pytest -q -p no:cacheprovider` (point any audit/socket/undo paths at a temp dir via env or fixtures, as AS0 did). Delete any temp dirs you create. +- Code style: short functions, comments explain WHY (Torvalds style); match the existing files. + +ARCHITECTURE FACTS (AS0, verified 2026-09-27) +- Two parts: an unprivileged **container** (`app.py`, FastAPI, Alpine image, no host fs, no timeshift/dpkg/ssh) and a **root daemon** (`sudo_bridge_daemon.py`, host systemd `agent-sudo-daemon.service`) that executes over a Unix socket (`bridge/sudo-bridge.sock`). The daemon independently re-validates every command against the signed SUDO.md (D9) and refuses unsigned rule files. +- Today tier-1 undo capture (`scoped_undo`), tier-3 Timeshift snapshot and tier-2 sandbox exec all run **inside the container** → they cannot work on the host (tests pass only via fakes). Tier 1 captures NO undo (`app.py` only calls `capture_or_refuse(…,3,…)`). Snapshot lane = container→SSH→`sudo timeshift` (no ssh in image). An exception from a missing binary after undo capture → uncaught HTTP 500. +- Tier-2 sandboxes: Incus containers `sandbox-primary` (10.85.172.80) + `sandbox-server01` (10.85.172.45) on server-01; the daemon on server-01 can use local `incus exec` as root. Primary tier-2 stays 503 by design (A3=c). +- Tier-3 substrate: server-01 Timeshift **rsync** mode on `/data` (sda4, UUID `6c2dbfca-8133-4e54-a510-8ffc19290f1c`) — not configured yet (owner step after SMART). Primary: NO substrate (row 202) — keep every tier-3 rule `primary=4`; never fake a snapshot. +- `app.py:185` compares the API key with `!=`; only ONE key (`BRIDGE_API_KEY`) is accepted. + +WHAT TO BUILD (in this order; each item = code + tests) +1. **Daemon socket ops (fixed, no free-form args):** `capture_undo`, `restore_undo`, `snapshot`, `sandbox_exec`, `deploy_service`. Each op validates its own inputs against a fixed schema; the container only orchestrates (calls the op, then `exec`). Keep the existing `exec` op unchanged in behaviour. +2. **A1 — tier 1 captures undo** (D4, decision A1=(a)): before any tier-1 exec the container calls `capture_undo` and refuses (409/503 with reason) if capture fails. Fix the test that pins tier-3-only capture. +3. **New undo kinds:** `systemctl` inverse-verb (enable↔disable, start↔stop, `--now` kept symmetric) and `noop` (for `systemctl daemon-reload`, whose reversal = revert the unit file + reload). File-backup undo stays as is but runs host-side. +4. **`snapshot` op:** runs `timeshift --create --comments "" --tags O --scripted` LOCALLY as root (no SSH lane); on primary it returns `{"ok":false,"reason":"no substrate (row 202)"}` without running anything. Any exception → structured failure → the API returns **503**, never 500. +5. **`sandbox_exec` op:** on server-01 runs `incus exec -- sh -c ` locally; the sandbox name is chosen from a fixed map by target host (primary→`sandbox-primary`, server-01→`sandbox-server01`), never from the request. On primary → refuse (A3=c). +6. **`deploy_service` op (owner decision — the Jenkins promote path):** input = a project name only. Validate it against `^[a-z0-9][a-z0-9-]{0,40}$` AND against the set of EXISTING unit files `/etc/systemd/system/-secretspec-resolver.service` (read the directory at call time; no other source). Action = `systemctl start -secretspec-resolver.service`, wait for it to finish (timeout 300 s), return its result + the last 20 journal lines for that unit. Tier = 3 semantics for authorisation (primary=4 stays unless the owner changes SUDO.md in AS2) — but make the op itself exist and be tested. Document in a short `DEPLOY_SERVICE.md` how Jenkins will call it (endpoint, body, caller key) — Jenkinsfile changes are NOT in scope. +7. **Auth:** accept multiple caller keys (e.g. `AGENT_SUDO_CALLERS` = JSON map name→key, raw keys hashed at startup exactly like secrets-proxy's `PROXY_CALLERS` — do NOT require pre-hashed keys: pre-hashing = double-hash = 403 anti-pattern) with backward-compatible fallback to `BRIDGE_API_KEY`; compare with `hmac.compare_digest`; record the caller name in the audit line. +8. **Set B gap:** `_is_write_to` must also catch a copy/install/move whose DESTINATION is a directory and whose SOURCE basename is a protected unit/file name (e.g. `cp …/agent-sudo-daemon.service /etc/systemd/system/`). Add tests. +9. **Tests:** new tests for every item above (mock subprocess at the daemon boundary; no real timeshift/incus/systemctl). All existing tests still pass. Report the counts before/after. + +OUT OF SCOPE (later phases — do not start them): SUDO.md rule edits / re-signing (AS2); image build + flags on (AS3); sudo-bridge retirement (AS4); unit-file hardening (owner); secrets-proxy changes (SP1+). + +DELIVERABLES +- Code + tests in the agent-sudo dir; `DEPLOY_SERVICE.md`; `AS1_CHANGES.md` (what changed per file, why, the new op schemas, migration notes for the owner: new env vars, which daemon restart + image rebuild are needed later, rollback = `git checkout` of the listed files). +- `git -C /opt/appdata/docker diff --stat -- docker-compose/agent-sudo` output pasted into AS1_CHANGES.md. + +SCOPE ALLOWLIST: `/opt/appdata/docker/docker-compose/agent-sudo/` EXCEPT `bridge/SUDO.md`, `bridge/command_audit.jsonl`, `bridge/approle/`, any `.env`, and `agent-sudo-daemon.service`. Plus `.claude/context.md` in that dir (append). + +PERSIST BEFORE YOU FINISH +- Append "## AS1 — (background agent)" to `/opt/appdata/docker/docker-compose/agent-sudo/.claude/context.md` (What was done / Current state / Owner next steps / Next phase AS2). +- Run: python3 /opt/appdata/docker/.claude/scripts/embed_memory_dir.py --only-recent 5 +- FINAL message = wrap-up JSON only, always: +{"status":"succeeded|partially_succeeded|failed","project":"agent-sudo #176","phase":"AS1","actions_taken":[],"actions_failed":[],"files_touched":[],"containers_restarted":[],"ops_added":["capture_undo","restore_undo","snapshot","sandbox_exec","deploy_service"],"a1_tier1_undo":"done|partial|no","auth":{"multi_caller":true,"compare_digest":true},"set_b_dir_dest":"fixed|no","tests":{"before":0,"after":0,"failed":0},"owner_next_steps":[],"unverified_claims":[],"next_step":"AS2 SUDO.md.proposed","notes":""} diff --git a/claude-config/config/prompts/agent-sudo-secrets-proxy/S1_secrets_proxy_investigation.md b/claude-config/config/prompts/agent-sudo-secrets-proxy/S1_secrets_proxy_investigation.md new file mode 100644 index 0000000..627d78a --- /dev/null +++ b/claude-config/config/prompts/agent-sudo-secrets-proxy/S1_secrets_proxy_investigation.md @@ -0,0 +1,40 @@ +# S1 — secrets-proxy investigation (personal_projects #191 → unblocks #174) — READ-ONLY + +Written 2026-09-27 (infrastructure general questions conversation). GAMEPLAN Step 2 (S1 is zero-dependency, goes first). +Safe to run in parallel with AS0 (both read agent-sudo code; neither writes it), W2, OLLAMA-1, BOOT-1, JH-1. + +--- + +You are a bounded, READ-ONLY investigation agent for secrets-proxy (personal_projects #191, parent #174). Budget: --max-turns 30. +If you issue the same tool call twice with identical arguments, STOP and output the wrap-up block with status=partially_succeeded. + +HARD RULES +- READ-ONLY. Do NOT start/restart/build any container or image, do NOT edit any code, no sudo, no Vault reads or writes, no git add/commit/push, no Postgres writes. Your only writes are the report, the new context.md, and the embed run. +- NEVER print a secret value. You may report that a variable/field EXISTS and its length, never its content. Do not `cat` a `.env`; use `grep -o '^[A-Z_]*='` style key-only listing. `docker inspect` only for State/Config.Image/Labels/Mounts/PortBindings — NEVER `.Config.Env`. +- If a hook blocks something, stop with a partial wrap-up. + +WHAT IS KNOWN (main session, 2026-09-27) +- Container `secrets-proxy-secrets-proxy-1` = `Exited (0)` since ~2026-07-09 (graceful stop, not a crash). Source: `/opt/appdata/docker/docker-compose/secrets-proxy/` (app.py ~851 lines, Dockerfile, docker-compose.yml, requirements.txt, `mirror/`, coolify-def-reference.md). Compose reportedly carries ~18 `${VAR}` placeholders (Coolify-era trap — memory `project_coolify_env_var_debt` if it exists). +- app.py has `/shell` + `env_secrets` + `/shell/approve|deny` (an NTFY approval gate). Locked design (grill-me 2026-07-12, memory `playbook_secrets_proxy_phases.md`): replace the NTFY gate with agent-sudo's tier 0–4 model via `proxy_rules.py` + a signed `PROXY.md`; forward privileged commands to agent-sudo `POST http://192.168.1.88:8084/exec` with its OWN caller key (≠ Claude Code's key). GAMEPLAN D4 says kill the NTFY gate. +- Jenkins (server-01) has a `secrets-proxy` job with 8 builds (Gitea repo). agent-sudo on primary is healthy (`192.168.1.88:8084/health` → 200), SUDO.md signed + verify-enforced since 07-14. +- The SP1 prompt in `playbook_secrets_proxy_phases.md` has at least one path inconsistency: it says create `/opt/appdata/docker/secrets-proxy/proxy_rules.py` but its own context points at `/opt/appdata/docker/docker-compose/secrets-proxy/`. There may be more stale facts. + +QUESTIONS TO ANSWER (each with evidence: file:line, git sha, log line) +1. WHY was it stopped on ~07-09? (git log -- docker-compose/secrets-proxy; `/opt/appdata/docker/Machines/*/.claude/context.md` and `/home/administrator/Desktop/claude/agent-builder/.claude/context.md` mentions — grep "secrets-proxy"; `docker logs --tail 200 secrets-proxy-secrets-proxy-1` last lines; memory `playbook_secrets_proxy.md`.) Was it stopped for safety, for a leaked token (#174 note), or scope confusion? +2. WHAT does app.py do today — endpoint inventory table (method, path, auth, what it executes/returns, which secrets it touches), and the auth model (PROXY_CALLERS? key hashing — note the known anti-pattern: pre-hashing keys = double-hash = 403). +3. WHERE do its secrets come from — list every `${VAR}` in the compose (names only) and classify: Vault path it should map to / obsolete Coolify var / unknown. Is it SecretSpec-ready (compare with `/opt/appdata/docker/docker-compose/cloudflared/secretspec.toml` pattern)? +4. WHO calls it — NOTE the locked Jenkins design (`/home/administrator/.claude/projects/-home-administrator-Desktop-claude/memory/project_cicd_jenkins.md`): every service bound for primary is PROMOTED by Jenkins through secrets-proxy `/shell` target=production. So Jenkins pipelines are a designed caller — find every Jenkinsfile that calls it and say whether the tier-classification refactor (SP2) keeps that promote path working (which tier would `docker compose up -d` in a prod project dir land in?). Then — grep hooks (`/opt/appdata/docker/.claude/hooks/`, `/home/administrator/.claude/`), skills, scripts, memory, n8n for its port/hostname; the Jenkins job's Jenkinsfile (read via `ssh -o BatchMode=yes administrator@192.168.1.90 'docker exec jenkins cat /var/jenkins_home/jobs/secrets-proxy/config.xml'` — print only url/scriptPath/branch lines, and fetch the Jenkinsfile from Gitea only if readable without credentials). +5. Is the SP1–SP4 plan still valid against today's agent-sudo code? Diff the assumptions: does `/opt/appdata/docker/docker-compose/agent-sudo/sudo_rules.py` still expose `load_rules/classify/is_dangerous/verb_heuristic/DANGER_PATTERNS`? Does agent-sudo `/exec` accept multiple callers today (read agent-sudo `app.py` auth)? Does a Vault transit key for PROXY.md exist? (You cannot read Vault — mark UNVERIFIED and say the main session must check `transit/keys/secrets-proxy-proxymd`.) +6. List every stale fact in the SP1 prompt text (path, file names, Gitea push instructions — note "agents never commit or push" is now rule 8 of the prompt playbook, so SP1's step 7 "push to Gitea" is itself stale). + +DELIVERABLES +1. Report `/opt/appdata/docker/research/secrets_proxy_S1_investigation.md`: Verdict (3 lines) → answers 1–6 → risks → "What SP1 must change" → a **DRAFT corrected SP1 prompt** in a fenced block at the end (same structure as the original; paths fixed; no commit/push step; scope allowlist; wrap-up JSON), clearly headed "DRAFT — main session reviews before saving to the playbook". +2. Add one row to `/home/administrator/Desktop/claude/claude-config/research/INDEX.md` if that index lists infra reports (read it first; follow its format; skip if it is business-only and say so). + +SCOPE ALLOWLIST: the report file, the INDEX.md row, `/opt/appdata/docker/docker-compose/secrets-proxy/.claude/context.md` (new, from template). Nothing else. + +PERSIST BEFORE YOU FINISH +- Create `/opt/appdata/docker/docker-compose/secrets-proxy/.claude/context.md` from `/home/administrator/.claude/projects/-opt-appdata-docker/memory/playbook_project_context_template.md`, then append "## S1 investigation — 2026-09-28 (background agent)" (What was done / Decisions surfaced / Current state / Next step). +- Run: python3 /opt/appdata/docker/.claude/scripts/embed_memory_dir.py --only-recent 5 +- FINAL message = wrap-up JSON only, always: +{"status":"succeeded|partially_succeeded|failed","project":"secrets-proxy #174/#191","phase":"S1","actions_taken":[],"actions_failed":[],"files_touched":[],"containers_restarted":[],"why_stopped":"","endpoints":0,"compose_vars":{"total":0,"vault_mappable":0,"obsolete":0,"unknown":0},"callers":[],"sp_plan_still_valid":"yes|partly|no","stale_facts_in_sp1":[],"unverified_claims":[],"next_step":"","notes":""} diff --git a/claude-config/config/prompts/infra/BOOT-1_boot_recovery_prep.md b/claude-config/config/prompts/infra/BOOT-1_boot_recovery_prep.md new file mode 100644 index 0000000..479257e --- /dev/null +++ b/claude-config/config/prompts/infra/BOOT-1_boot_recovery_prep.md @@ -0,0 +1,70 @@ +# BOOT-1 BUILD-PREP — boot-recovery layer (personal_projects #219, GAMEPLAN Step 0) — write files, install NOTHING + +Written 2026-09-27 (infrastructure general questions conversation). Safe to run in parallel with W2, OLLAMA-1, S1, AS0, JH-1. + +WHY NOW (verified live 2026-09-27 14:05): server-01 booted 2026-09-17 14:20:36 CDT; dockerd started the same second, +before the USB WiFi (`wlx1cbfce9afe93`, wpa_supplicant, NOT NetworkManager) had its LAN IP and before /etc/resolv.conf +had a nameserver. Every container that publishes on `192.168.1.90:` came up "Up (healthy)" but with **NO +published ports** (`docker port` empty) — agent-sudo (8082), jenkins (8090, 50000), hermes (8642, 9119), n8n-prod — and +hermes has an **empty /etc/resolv.conf** (Telegram gateway failing DNS for 10 days). Only `ollama` (0.0.0.0 bind, +recreated 3 days ago) works. The same failure mode caused the 09-07 primary outage (GAMEPLAN §1.1). A plain +`docker restart` re-programs port bindings and regenerates resolv.conf — that is the recovery primitive. + +--- + +You are a bounded background BUILD-PREP agent for personal_projects #219 (boot-recovery layer). Budget: --max-turns 30. +If you issue the same tool call twice with identical arguments, STOP and output the wrap-up block with status=partially_succeeded. + +HARD RULES +- WRITE FILES ONLY. Do NOT restart/start/stop any container, do NOT install/enable any unit, no sudo, no sysctl changes, no git add/commit/push, no Postgres writes, no ntfy sends. The owner installs after review. +- Read-only discovery on both hosts is fine: primary locally; server-01 via `ssh -o BatchMode=yes administrator@192.168.1.90 ''` (key auth works; administrator is in the docker group there). `docker inspect` only for HostConfig.PortBindings / RestartPolicy / Labels / State — NEVER print `.Config.Env`. +- server-01 is a SHARED host (n8n-prod, hermes, jenkins, sandboxes): discovery only. +- If a hook blocks something, stop with a partial wrap-up. Issue each file write as its own call; do not bundle mutate+verify in one `sh -c`. + +ROLE BOUNDARY (owner-locked 2026-06-25: Jenkins = deploys, Hermes = monitoring + restart-on-health-failure): this layer does NOT take over Hermes's job. It is level-0 (GAMEPLAN §2 0.3): host systemd, outside Docker, and it ONLY repairs the boot-order failure (missing ports / empty DNS / not started) — the exact failure that also blinded Hermes itself for 10 days. Ordinary health failures stay Hermes's. The jsonl it writes is the feed Hermes reads later. Say this in the README. + +LOCKED DESIGN (main session, 2026-09-27 — do not revisit; this refines GAMEPLAN Step 0.1/0.3) +Two layers, per host, identical code, host differences only in a small env file: +A) **Root-cause fix — `docker.service` drop-in** `/etc/systemd/system/docker.service.d/10-wait-lan.conf`: + `ExecStartPre=/usr/local/sbin/wait-lan-ready` + `TimeoutStartSec=330`. `wait-lan-ready` (bash) loops every 2 s, + max 300 s, until BOTH (1) `ip -4 -o addr show` contains `inet ${LAN_IP}/` and (2) `/etc/resolv.conf` has at least + one `nameserver` line; it logs each outcome to the journal and **always exits 0** (Docker must still start after + the cap — never block the whole stack forever). Reads `LAN_IP` from `/etc/control-plane-up/host.env`. +B) **Backstop — `control-plane-up.service` (oneshot) + `.timer`** (OnBootSec=3min, OnUnitActiveSec=10min): + script `/usr/local/sbin/control-plane-up` iterates a FIXED list file `/etc/control-plane-up/containers.list` + (one container name per line, `#` comments allowed — no arguments from anywhere else). For each container: + - not running AND RestartPolicy is `unless-stopped`/`always` → `docker start `; + - running but HostConfig.PortBindings non-empty while `docker port ` is empty → `docker restart `; + - running but its /etc/resolv.conf (`docker exec cat /etc/resolv.conf`, skip if exec fails) has no + `nameserver` line → `docker restart `; + - `/etc/control-plane-up/skip` lists names to never touch; a separate `CHECK_ONLY` list (host.env) = log + alert only. + Every decision → one JSON line to `/var/log/control-plane-up.jsonl` (`ts, host, container, check, action, result`). + Optional ntfy: if `NTFY_URL`+`NTFY_TOPIC` are set in host.env, POST a one-line summary ONLY when it acted + (leave both blank in the shipped env files — the owner fills them). + Idempotent: when healthy it does nothing but log one `ok` line per run. +C) **Primary CHECK_ONLY (never restart):** `vault-iwaulpoi5hwirdlogshmul40`, node_exporter, prometheus (runc + named-user bug — they cannot restart; see memory `reference_vault_healthcheck_runc_username.md`). **Skip entirely:** + `secrets-proxy-*` (deliberately stopped, #191) and `sudo-bridge-*` (being retired). +D) **List contents = discovered now, not guessed:** containers.list per host = every container whose PortBindings + reference that host's LAN IP (192.168.1.88 / 192.168.1.90), PLUS on server-01 `hermes` and `n8n-prod-*` (DNS- + dependent). Resolve names live (`docker ps -a --format '{{.Names}}'`); UUID-suffixed names are fine to write + literally but add a comment that they are Coolify-era names that change on redeploy. +E) `sysctl net.ipv4.ip_nonlocal_bind=1` (GAMEPLAN 0.1) — ship it as `90-docker-lan-bind.conf` but mark it OPTIONAL in + the README: it fixes the port bind but NOT the empty resolv.conf, so layer A is the real fix. + +DELIVERABLES — new dir `/opt/appdata/docker/boot-recovery/` (git-tracked repo on primary; owner copies to server-01) +1. `wait-lan-ready` and `control-plane-up` (bash, `set -uo pipefail`, short functions, comments explain WHY). +2. `docker.service.d/10-wait-lan.conf`, `control-plane-up.service`, `control-plane-up.timer`. +3. `hosts/primary/{host.env,containers.list,skip}` and `hosts/server-01/{host.env,containers.list,skip}`. +4. `sysctl/90-docker-lan-bind.conf` (optional layer). +5. `README.md`: what/why (the 09-17 evidence above), exact owner install commands per host (copy files, `sudo install -m 0755 …`, `sudo systemctl daemon-reload`, `sudo systemctl enable --now control-plane-up.timer`), how to test WITHOUT a reboot (`sudo systemctl start control-plane-up.service` then read the jsonl), the reboot test from GAMEPLAN ("reboot server-01 → everything green within 15 min, no human touch" — server-01 first, primary only after it passes), and rollback (disable timer, remove drop-in, daemon-reload). +6. Validate without installing: `bash -n` both scripts; `shellcheck` if present; `systemd-analyze verify` on the unit files (use a temp copy dir if it complains about paths); run `control-plane-up` in a DRY-RUN mode you build in (`--dry-run` flag or `DRY_RUN=1` → prints the decisions, executes nothing) against the REAL primary list, and against server-01 by copying it to `/tmp` there via ssh and running it dry — this proves the detection logic on the actually-broken server-01 containers. Delete the /tmp copy afterwards. + +SCOPE ALLOWLIST: `/opt/appdata/docker/boot-recovery/` (new) + one `/tmp/control-plane-up.*` copy on server-01 (delete it). Nothing else. + +PERSIST BEFORE YOU FINISH +- Create `/opt/appdata/docker/boot-recovery/.claude/context.md` from `/home/administrator/.claude/projects/-opt-appdata-docker/memory/playbook_project_context_template.md` and append "## BOOT-1 prep — 2026-09-28 (background agent)" (What was done / Current state / Owner steps / Next step). +- Append a 3-line pointer to "/opt/appdata/docker/Machines/infrastructure general questions/.claude/context.md" (one `cat >>` heredoc). +- Run: python3 /opt/appdata/docker/.claude/scripts/embed_memory_dir.py --only-recent 5 +- FINAL message = wrap-up JSON only, always: +{"status":"succeeded|partially_succeeded|failed","project":"boot-recovery #219","phase":"BOOT-1 prep","actions_taken":[],"actions_failed":[],"files_touched":[],"containers_restarted":[],"lists":{"primary":[],"server-01":[]},"dry_run":{"primary":[{"container":"","check":"","would_do":""}],"server-01":[{"container":"","check":"","would_do":""}]},"validation":{"bash_n":"","shellcheck":"","systemd_verify":""},"unverified_claims":[],"next_step":"","notes":""} diff --git a/claude-config/config/prompts/infra/JH-1_jenkins_hermes_state_audit.md b/claude-config/config/prompts/infra/JH-1_jenkins_hermes_state_audit.md new file mode 100644 index 0000000..f7542df --- /dev/null +++ b/claude-config/config/prompts/infra/JH-1_jenkins_hermes_state_audit.md @@ -0,0 +1,38 @@ +# JH-1 — Jenkins + Hermes state audit (personal_projects #181, #208, GAMEPLAN §3) — READ-ONLY + +Written 2026-09-27 (infrastructure general questions conversation). Run AFTER the owner has restarted `jenkins` and +`hermes` on server-01 (their published ports + hermes DNS were dead since the 09-17 boot — see BOOT-1). If +`curl -s -o /dev/null -w '%{http_code}' http://192.168.1.90:8090/login` is still 000 when you start, continue anyway +using `docker exec` (it does not need the ports) and note it. Safe to run in parallel with W2, OLLAMA-1, S1, AS0, BOOT-1. + +--- + +You are a bounded, READ-ONLY audit agent for the Jenkins (deploy) + Hermes (monitor) layer on server-01. Budget: --max-turns 30. +If you issue the same tool call twice with identical arguments, STOP and output the wrap-up block with status=partially_succeeded. + +HARD RULES +- READ-ONLY. No container restart/stop/start, no Jenkins build triggers, no job/config edits, no Hermes messages/commands sent, no sudo, no Vault, no git add/commit/push, no Postgres writes. +- All server-01 access = `ssh -o BatchMode=yes administrator@192.168.1.90 ''` (administrator is in the docker group there). server-01 is SHARED (n8n-prod, sandboxes) — touch only jenkins + hermes. +- **Never print secrets.** Jenkins `credentials.xml`/`secrets/`: list credential IDs + types ONLY (grep `` and the class name), never ``/``/`` bodies. Hermes config: print keys/section names and non-secret values only; for anything named *key/token/secret/password* print ``/``. NEVER `docker inspect … .Config.Env` and never run `env`/`printenv` in a container. +- If a hook blocks something, stop with a partial wrap-up. + +FACTS (main session, verified 2026-09-27) +- jenkins: image `gitea.local/backtalk6858/jenkins:latest`, ports 192.168.1.90:8090→8080 and :50000, compose `/opt/appdata/docker/docker-compose/jenkins/` on server-01 (Dockerfile, plugins.txt, start.sh, vault-approle/, jenkins-home/). 36 job folders; only 4 have ever built: jellyfin (5), secrets-proxy (8), sudo-bridge (12), sudo-bridge-server01 (2). Row #181: "provisioned but idle, and the record says otherwise". Coolify is retired; the intended deploy path is Jenkins. READ these role definitions first (they are the yardstick for the gap analysis): `/home/administrator/.claude/projects/-home-administrator-Desktop-claude/memory/project_jenkins_hermes_boundary.md` (LOCKED 2026-06-25: **Jenkins = ALL deploys** — builds images, injects Vault secrets at deploy time, `docker compose up`, rollbacks; **Hermes = ALL monitoring** — health watch, NTFY/Telegram alerts, auto-restart on health failure after announcing, escalate; Hermes never redeploys itself, it TRIGGERS the Jenkins pipeline), `.../-home-administrator-Desktop-claude/memory/project_cicd_jenkins.md` (pipeline patterns: services bound for primary = build → push → sandbox deploy on server-01 → smoke test → promote to primary via **secrets-proxy `/shell` target=production**; server-01-only services use `Jenkinsfile.server01`), `.../-home-administrator-Desktop-claude/memory/project_hermes_exploration_day.md`, `.../-home-administrator-Desktop-claude/memory/feedback_hermes_telegram_api.md` (task injection = POST to the Hermes gateway API :8642, never Telegram sendMessage), `.../-home-administrator-Desktop-claude/memory/feedback_evaluate_jenkins_hermes_fit.md`, and in `/home/administrator/.claude/projects/-opt-appdata-docker/memory/`: `jenkins_deployment_transition.md`, `project_virtual_it_department.md` (end goal = unattended self-healing; Phase 2 = sandbox-verify-before-apply). +- hermes: image `nousresearch/hermes-agent:latest` (unpinned), ports 192.168.1.90:8642 and :9119, network `hermes_default`; logs show the Telegram gateway failing DNS for days (empty resolv.conf). Its compose on primary is `/opt/appdata/docker/docker-compose/hermes/docker-compose.yml` (703 B, June) — find the one server-01 actually runs from via the container label `com.docker.compose.project.working_dir`. Owner rule: the always-on agent is Hermes, never OpenClaw. +- A second Ollama consolidation is in flight (OLLAMA-1): the single Ollama will be server-01 `ollama` (0.0.0.0:11434, GPU). + +TASKS +1. **Jenkins inventory:** for each of the 36 jobs: type (pipeline/freestyle), SCM URL + branch + scriptPath, whether the Gitea repo + Jenkinsfile exist (list repos via the Gitea API only if anonymous read works; otherwise UNVERIFIED), last build result/date, what it deploys to and HOW (docker socket mount? SSH to primary? agent on primary?). Nodes/agents configured. Credential IDs + types. Plugin count + any failed plugin loads in the log. Whether Jenkins can reach primary's Docker at all today. +2. **Hermes inventory:** version/tag + digest; the model/provider backend it is configured for (does it use Ollama? which host/model?); enabled gateways (Telegram etc.) and their state; scheduled jobs/crons/skills it has; tools/permissions (docker.sock? SSH keys? network reach to primary?); memory/data dir size; whether anything in it already monitors the control plane (GAMEPLAN says Hermes = level-1 monitor reading `/var/log/control-plane-up.jsonl` + `command_audit`). +3. **Gap analysis vs the plan** — include each roadmap item from the boundary memory as its own row: bitwarden-bridge pipeline; a pipeline for every Coolify-debt container; the Hermes skill that triggers a Jenkins redeploy on persistent health failure; Hermes dashboard route (:9119) + Telegram verified; the `automation_ideas` background re-evaluation that fires once BOTH are confirmed live. Also check: does any pipeline still promote via secrets-proxy `/shell` (stopped since 07-09 → that promote path is dead) or via sudo-bridge (being retired)? Plus: GAMEPLAN §3 (Jenkins = the only write path from the future claude-runner VM; Hermes = level-1 monitor, CA-P5 ladder, D10 watchdog) and rows #181, #208, #235 (skill: deploy — Jenkins first). What exists vs what's missing, as a table. +4. **Proposed build order** for Jenkins + Hermes as bounded agent phases (each ≤ 1 session, name the owner steps between them), sized so 4–5 agents can run in parallel on other projects without conflicting (state which phases touch server-01 shared state and must run alone). + +DELIVERABLE: report `/opt/appdata/docker/research/jenkins_hermes_state_2026-09-28.md`: Verdict (5 lines) → Jenkins table → Hermes table → gap table → proposed phases → UNVERIFIED list → "facts in memory that are now wrong" (file + line + correct fact). + +SCOPE ALLOWLIST: the report; `/opt/appdata/docker/docker-compose/jenkins/.claude/context.md` and `/opt/appdata/docker/docker-compose/hermes/.claude/context.md` on PRIMARY (create from template if missing). Nothing else. + +PERSIST BEFORE YOU FINISH +- For each of the two context.md files: create from `/home/administrator/.claude/projects/-opt-appdata-docker/memory/playbook_project_context_template.md` if missing, append "## JH-1 audit — 2026-09-28 (background agent)" (What was done / Current state / Next step). +- Run: python3 /opt/appdata/docker/.claude/scripts/embed_memory_dir.py --only-recent 5 +- FINAL message = wrap-up JSON only, always: +{"status":"succeeded|partially_succeeded|failed","project":"jenkins #181 + hermes","phase":"JH-1","actions_taken":[],"actions_failed":[],"files_touched":[],"containers_restarted":[],"jenkins":{"jobs":36,"jobs_ever_built":0,"jobs_with_valid_scm":0,"deploy_mechanism":"","reachable_8090":true},"hermes":{"image_digest":"","llm_backend":"","gateways":[],"scheduled_jobs":0,"monitors_control_plane":false},"stale_memory_facts":[],"unverified_claims":[],"next_step":"","notes":""} diff --git a/claude-config/config/prompts/infra/OLLAMA-1_single_instance_server01.md b/claude-config/config/prompts/infra/OLLAMA-1_single_instance_server01.md new file mode 100644 index 0000000..c68c404 --- /dev/null +++ b/claude-config/config/prompts/infra/OLLAMA-1_single_instance_server01.md @@ -0,0 +1,42 @@ +# OLLAMA-1 — make server-01 (GPU) the ONLY Ollama; repoint primary callers; write the retirement steps + +Written 2026-09-27 (infrastructure general questions conversation). Owner decision 2026-09-27: run ONE Ollama, on +server-01's GPU; retire primary's CPU-only copy. Safe to run in parallel with W2, S1, AS0, JH-1, BOOT-1. + +--- + +You are a bounded background agent. Task: repoint every primary-side Ollama caller to server-01, prove it works, and write (not run) the owner steps that retire primary's Ollama. Budget: --max-turns 25. +If you issue the same tool call twice with identical arguments, STOP and output the wrap-up block with status=partially_succeeded. + +HARD RULES +- Do NOT stop/start/restart/remove any container, do NOT touch systemd units, no sudo, no git add/commit/push, no Postgres writes except what `embed_memory_dir.py` itself does when you run it. The owner retires the primary container. +- Never print secret values. If a hook blocks something, stop with a partial wrap-up. + +FACTS (verified by the main session 2026-09-27 14:00 — re-check only the ones marked RE-CHECK) +- Primary (192.168.1.88) has **no NVIDIA GPU** (AMD Cezanne iGPU only). Its Ollama = container `ollama-ollama-1`, CPU-only, published `127.0.0.1:11434`, compose at `/opt/appdata/docker/docker-compose/ollama/` (secretspec-managed: `ollama-secretspec-resolver.service` + an ENABLED `ollama-secretspec-resolver.timer` that re-ups it periodically — so stopping the container alone is not enough; the timer must be disabled too). +- server-01 (192.168.1.90) Ollama = container `ollama` (image `ollama-fixed:1.0.0`, runtime nvidia, RTX 2060 SUPER 8 GB, logs show 13/13 layers offloaded to CUDA0), published `0.0.0.0:11434` (LAN). Compose `/data/docker-compose/ollama/docker-compose.yml`. +- Both hosts hold the same models with **identical digests**: `nomic-embed-text:latest` 0a109f422b47e3a3…, `llama3.1:8b` 46e0c10c039e0191… → embeddings from either host are the same vectors; NO re-index needed. +- Known callers: `/opt/appdata/docker/.claude/scripts/embed_memory_dir.py` line ~24 `OLLAMA_URL = "http://localhost:11434/api/embeddings"` (→ primary CPU — THE one to repoint); `.claude/scripts/semantic_recall.py` already uses `http://192.168.1.90:11434`; `non-docker-python-scripts/youtube-pipeline/script_gen.py` already uses 192.168.1.90. `legacy-docker-compose/open web ui` points at .88 but is legacy/not running — leave it, just list it. +- RE-CHECK: grep for any other caller the main session missed: `grep -rn -E "11434|OLLAMA_(HOST|URL|BASE)" /opt/appdata/docker/.claude /home/administrator/.claude/hooks /home/administrator/.claude/scripts /home/administrator/.claude/skills /home/administrator/.claude/settings.json /opt/appdata/docker/non-docker-python-scripts /opt/appdata/docker/docker-compose --include=*.py --include=*.sh --include=*.json --include=*.yml --include=*.yaml --include=*.toml` (exclude `pin-bumper/.cache`, `logs/`, `*.bak*`). Also check Docker networks: is any running primary container configured to reach `ollama-ollama-1` by name (`docker inspect` Config.Image/NetworkSettings only — grep compose files for `ollama:` hostnames)? Also check the Stop hook's embed path (the hook that embeds `[[MEMORY_EMBED]]` tags — find which script it calls and what URL that script uses). + +DECISIONS (locked — do not revisit) +1. Every repointed caller reads `OLLAMA_URL` from the environment with default `http://192.168.1.90:11434/api/embeddings` (or the matching base URL + path style the file already uses). One-line change per file + a one-line comment "single Ollama lives on server-01 GPU (2026-09-28)". Keep the file's existing style. +2. server-01's Ollama stays as it is (LAN-exposed, no auth — router does not forward 11434). Do not change it; note the exposure in the report as a known accepted risk. +3. Failure behaviour: if server-01 is unreachable, `embed_memory_dir.py` must fail LOUDLY (non-zero exit, `[EMBED] FAILED` line) — it already does since #183; confirm, do not add a silent fallback to primary. + +STEPS +1. RE-CHECK callers (above). Record each file + line. +2. Back up each file you will edit to `.bak-20260928-ollama` (cp), then edit. +3. `python3 -c "import ast; ast.parse(open('').read())"` for each edited .py. +4. Prove it: run `python3 /opt/appdata/docker/.claude/scripts/embed_memory_dir.py --only-recent 1` and confirm exit 0 + `[EMBED] Done`. Then prove the TRAFFIC went to server-01, not primary: `curl -s 192.168.1.90:11434/api/ps` right after (nomic-embed-text should be loaded there with size_vram > 0) and `curl -s localhost:11434/api/ps` on primary (should NOT show it newly loaded — note expires_at). +5. Run one semantic recall query end-to-end (`python3 /opt/appdata/docker/.claude/scripts/semantic_recall.py "mediashare drive re-enumeration"` or its documented CLI — read its argparse first) and confirm results come back. +6. Write the OWNER retirement runbook into `/opt/appdata/docker/docker-compose/ollama/RETIRE.md`: (a) `sudo systemctl disable --now ollama-secretspec-resolver.timer`, (b) `docker stop ollama-ollama-1` (keep the container + volume 7 days as rollback), (c) 24-hour watch list (next end-session embed run succeeds; /recall works), (d) after 7 days: `docker compose -f /opt/appdata/docker/docker-compose/ollama/docker-compose.yml down` (NO `-v` until the owner decides), how much RAM/disk it frees (measure the volume size with `du -sh` on the bind/volume path via `docker inspect` Mounts, and the image size), (e) rollback = `sudo systemctl enable --now ollama-secretspec-resolver.timer`. +7. Update docs that state where Ollama lives: grep the memory dir `/home/administrator/.claude/projects/-opt-appdata-docker/memory/` for "localhost:11434" / "ollama" claims about primary (skip `file_snapshot_*`); list them in the report; edit ONLY lines that are now factually wrong (append "(2026-09-28: single Ollama = server-01 GPU)" rather than rewriting paragraphs). + +SCOPE ALLOWLIST: the caller files you repoint (+ their .bak copies), `/opt/appdata/docker/docker-compose/ollama/RETIRE.md`, `/opt/appdata/docker/docker-compose/ollama/.claude/context.md` (create from `/home/administrator/.claude/projects/-opt-appdata-docker/memory/playbook_project_context_template.md` if missing), the memory-dir lines in step 7. Nothing else. + +PERSIST BEFORE YOU FINISH +- Append "## OLLAMA-1 — 2026-09-28 (background agent)" to `/opt/appdata/docker/docker-compose/ollama/.claude/context.md` (What was done / Current state / Owner steps / Next step). +- Run: python3 /opt/appdata/docker/.claude/scripts/embed_memory_dir.py --only-recent 5 +- FINAL message = wrap-up JSON only, always: +{"status":"succeeded|partially_succeeded|failed","project":"ollama single-instance","phase":"OLLAMA-1","actions_taken":[],"actions_failed":[],"files_touched":[],"containers_restarted":[],"callers_found":[{"file":"","line":0,"old":"","new":"","status":"repointed|already_server01|legacy_left"}],"proof":{"embed_exit":0,"server01_ps_shows_model":true,"primary_ps_idle":true,"recall_ok":true},"primary_ollama_frees":{"ram":"","disk":""},"unverified_claims":[],"next_step":"owner runs RETIRE.md","notes":""} diff --git a/claude-config/config/prompts/infra/VS-1_vault_sandbox_autounseal_prep.md b/claude-config/config/prompts/infra/VS-1_vault_sandbox_autounseal_prep.md new file mode 100644 index 0000000..f67a662 --- /dev/null +++ b/claude-config/config/prompts/infra/VS-1_vault_sandbox_autounseal_prep.md @@ -0,0 +1,42 @@ +# VS-1 BUILD-PREP — give server-01's vault-sandbox the SAME auto-unseal as primary Vault — write files, install NOTHING + +Written 2026-09-27 (infrastructure general questions conversation). Owner decision 2026-09-27 (DECISIONS.md, item 9): +"sandbox vault should have the same auto unseal functionality primary vault does and if it doesn't then it needs to be +copied over." Related: personal_projects #270 (sandbox stack unreachable — boot-order bug, see BOOT-1) and #219 (BOOT-1 +marks vault-sandbox CHECK_ONLY *because* a restart seals it; once this lands, that can change). +Needs server-01 reachable. Safe alongside anything that does not restart vault-sandbox. + +--- + +You are a bounded background BUILD-PREP agent. Budget: --max-turns 25. +If you issue the same tool call twice with identical arguments, STOP and output the wrap-up block with status=partially_succeeded. + +HARD RULES +- WRITE FILES ONLY + read-only discovery. Do NOT restart/stop/start any container (especially vault-sandbox — a restart seals it), do NOT run `vault operator unseal`, do NOT install units, no sudo, no git add/commit/push, no Postgres writes. +- **NEVER read, print or copy an unseal key or root token** — not from files, Bitwarden, Vault or history. You may report that a key source EXISTS (item/path name, field names, count of keys) and nothing more. Do not `cat` any `*unseal*keys*` file; use `ls -la` / `stat` only. +- server-01 via `ssh -o BatchMode=yes administrator@192.168.1.90 ''`. `docker inspect` only State/Image/Labels/Mounts/PortBindings/RestartPolicy — NEVER `.Config.Env`. +- If a hook blocks something, stop with a partial wrap-up. + +PRIMARY'S MECHANISM (verified 2026-09-27 — this is the pattern to copy) +- `/usr/local/bin/vault-unseal.sh` (root 0755): waits ≤120 s for a running container matching `^vault-`, and if `docker exec vault status` shows `Sealed true`, reads keys line-by-line from a root-only file (`/etc/vault-unseal-keys`, root:root 0400, 3 keys for a 3-of-5 threshold; `#`/blank lines skipped) and runs `docker exec vault operator unseal ` per key (output to /dev/null). +- `/usr/local/bin/vault-watch-unseal.sh` (root 0755): `docker events --filter 'name=vault-' --filter 'event=start' --format '{{.Actor.Attributes.name}}'` → on each start, sleep 8, run vault-unseal.sh. Handles container restarts/recreates, not just daemon restarts. +- Units: `vault-watch-unseal.service` (Type=simple, Restart=always, RestartSec=5, After/Requires docker.service, WantedBy multi-user) — ACTIVE; `vault-unseal.service` (oneshot, After docker) — currently FAILED on primary (note it in the report; do not fix primary). +- Known weakness to NOT copy blindly: the key is passed as a positional argument to `vault operator unseal` (visible in the host process list for the moment it runs). Improve it for the sandbox copy: feed the key on stdin (`printf '%s\n' "$KEY" | docker exec -i vault operator unseal -` — verify the `-` stdin form in `vault operator unseal -h` inside the vault-sandbox container, read-only). + +TASKS +1. **Discover on server-01 (read-only):** does any unseal automation already exist there? (`systemctl list-units --all | grep -i unseal`, `ls -la /usr/local/bin/*unseal* /etc/*unseal* 2>&1`, root crontab is not readable — say so). vault-sandbox: container name, image/version, compose working dir (label), seal type (`docker exec vault status -format=json` → print ONLY `sealed`, `initialized`, `t`, `n`, `type`, `version` fields), Vault address it listens on, restart policy. +2. **Find where the sandbox unseal keys live — by NAME only:** search memory (`/home/administrator/.claude/projects/-opt-appdata-docker/memory/`, especially `session_summary_n8n_sandbox_deploy.md`, `playbook_vault_token_rotation.md`, `project_hashicorp_vault.md`, `feedback_sandbox_isolation.md`) and context files for where the 2026-06-16 sandbox init stored them (a Bitwarden item? a primary-Vault path? a file on server-01?). Report item/path + field names + key count + threshold. If unknown → UNVERIFIED and say the owner must locate them. +3. **Write the adapted files** into the primary repo dir `/opt/appdata/docker/vault-sandbox-autounseal/`: + - `vault-sandbox-unseal.sh` — same logic as primary's, but container match `^vault-sandbox-`, key file `/etc/vault-sandbox-unseal-keys`, key fed via stdin, logs to the journal, exits non-zero if still sealed after applying keys. + - `vault-sandbox-watch-unseal.sh` — docker events filter that matches ONLY the sandbox container (verify the events `name` filter semantics — substring vs exact — from Docker docs, and use a filter + in-loop name check so it can never fire on some other vault). + - `vault-sandbox-watch-unseal.service` (same shape as primary's watch unit) and a oneshot `vault-sandbox-unseal.service` run at boot After=docker.service (covers "container already up before the watcher started"). + - `README.md`: owner steps — (a) create `/etc/vault-sandbox-unseal-keys` root:root 0400 from the key source found in task 2 WITHOUT the key touching shell history or the screen (e.g. `sudo install -m 0400 /dev/null /etc/vault-sandbox-unseal-keys && sudo nano /etc/vault-sandbox-unseal-keys` and paste from the password manager), (b) install scripts + units, `daemon-reload`, `enable --now` the watcher and the oneshot, (c) TEST: `docker restart ` then `journalctl -u vault-sandbox-watch-unseal -n 20` shows "unsealed" and `vault status` sealed=false, (d) then tell the main session so BOOT-1's server-01 `host.env` can drop vault-sandbox from CHECK_ONLY, (e) rollback. +4. Validate: `bash -n` both scripts; `systemd-analyze verify` on temp copies of the units. + +SCOPE ALLOWLIST: `/opt/appdata/docker/vault-sandbox-autounseal/` (new, incl. `.claude/context.md` from `/home/administrator/.claude/projects/-opt-appdata-docker/memory/playbook_project_context_template.md`). Nothing else. + +PERSIST BEFORE YOU FINISH +- Append "## VS-1 prep — (background agent)" to `/opt/appdata/docker/vault-sandbox-autounseal/.claude/context.md`. +- Run: python3 /opt/appdata/docker/.claude/scripts/embed_memory_dir.py --only-recent 5 +- FINAL message = wrap-up JSON only, always: +{"status":"succeeded|partially_succeeded|failed","project":"vault-sandbox auto-unseal (#270/#219)","phase":"VS-1 prep","actions_taken":[],"actions_failed":[],"files_touched":[],"containers_restarted":[],"server01_existing_unseal":"none|partial|present","vault_sandbox":{"container":"","version":"","sealed":true,"initialized":true,"threshold":"t-of-n","seal_type":""},"key_source":{"where":"","fields":[],"count":0,"verified":false},"stdin_unseal_supported":true,"validation":{"bash_n":"","systemd_verify":""},"unverified_claims":[],"next_step":"","notes":""} diff --git a/claude-config/config/prompts/voice/V0_tts_engine_research.md b/claude-config/config/prompts/voice/V0_tts_engine_research.md new file mode 100644 index 0000000..8c0c2fb --- /dev/null +++ b/claude-config/config/prompts/voice/V0_tts_engine_research.md @@ -0,0 +1,47 @@ +# V0 — re-confirm the TTS engine for Voice Chat (#214) before V2 — RESEARCH ONLY + +Written 2026-09-27 (infrastructure general questions conversation). Owner approved re-checking the July choice +(Coqui XTTS-v2) because Coqui the company shut down (early 2024; community fork `coqui-tts` by idiap) and the +XTTS-v2 model licence (CPML) is non-commercial. Research only — no installs. Safe to run alongside anything. + +--- + +You are a bounded RESEARCH agent. Budget: --max-turns 25. +If you issue the same tool call twice with identical arguments, STOP and output the wrap-up block with status=partially_succeeded. + +HARD RULES +- Research only: WebSearch / WebFetch + reading local files. No installs, no pulls, no containers, no SSH, no sudo, no Vault, no git, no Postgres writes. Only writes: the report, the INDEX row, the embed run. +- Every claim about a project (licence, last release date, VRAM, cloning ability, Docker image) must cite a URL you actually fetched in this run. Anything you could not verify = mark UNVERIFIED. Today is 2026-09-27; prefer sources from 2026. + +REQUIREMENTS (locked in the July build plan `/home/administrator/Desktop/claude/agent-builder/voice_chat_build_plan.md` — read §2, §5, §7 — plus 2026-09-27 facts) +- R1 Fully local, zero cloud, zero API cost. English. +- R2 **Voice cloning from a 6–30 s reference clip** (Jarvis now; anime character voices later = reference swap). Engines that cannot clone are baselines only. +- R3 Fits the GPU budget: server-01 RTX 2060 SUPER **8 GB** (Turing, compute capability 7.5 — **no bf16, no FlashAttention-2**; check each engine's minimum), shared with Ollama llama3.1:8b (~5 GB when loaded, unloads after 5 min idle), nomic-embed-text (~0.3 GB) and a faster-whisper STT (~1–2 GB, V1). Target: TTS ≤ ~3 GB VRAM, or state what must be unloaded. Driver 550 (CUDA 12.4). +- R4 Low latency: sentence-by-sentence streaming; first audio ≤ ~1.5 s after a sentence arrives; real-time factor < 1 on that GPU class. +- R5 Maintained: a release or meaningful commit in the last ~6 months; a usable Docker image or a simple containerisable server with an HTTP API. +- R6 Licence: personal use is fine; ALSO state whether commercial use is allowed (the owner's "virtual IT department" idea #29 may become a product — a non-commercial licence is a flag, not a disqualifier). + +FIXED CANDIDATE LIST (evaluate exactly these; add at most 2 others only if they appear in ≥2 independent 2026 comparisons) +1. XTTS-v2 via the idiap `coqui-tts` fork (the July choice) +2. F5-TTS +3. Chatterbox (Resemble AI) +4. Fish Speech / OpenAudio +5. CosyVoice 2 (or its current successor) +6. IndexTTS 2 +7. Zonos (Zyphra) +8. Kokoro — baseline (no cloning) +9. Piper — baseline (rejected in July: fixed voices) + +DELIVERABLE +Report `/opt/appdata/docker/research/voice_tts_engine_2026-09-28.md`: +1. Verdict (≤5 lines): the recommended engine + runner-up, and whether it changes the July decision. +2. Scoring table: rows = candidates; columns = R1–R6 (pass/fail/partial + one-line evidence + source URL), VRAM (measured-by-someone vs claimed), cloning clip length needed, streaming support, Docker image (exact name) + API style, last release date, licence (code / weights separately). +3. For the top 2: exact install path we would use on server-01 (image or build), config for Turing/fp16, the HTTP call the Stop hook would make, and the VRAM-sharing plan with Ollama + STT (options, not a decision). +4. Risks + UNVERIFIED list. +Then add one row to `/home/administrator/Desktop/claude/claude-config/research/INDEX.md` (read it first; follow its format). + +PERSIST BEFORE YOU FINISH +- Append a dated "## V0 TTS engine research — 2026-09-28 (background agent)" block (verdict + report path) to `/home/administrator/Desktop/claude/agent-builder/.claude/context.md` — append only, one `cat >>` heredoc; do NOT read or print other parts of that file beyond what you need (it contains a plaintext secret scheduled for redaction). +- Run: python3 /opt/appdata/docker/.claude/scripts/embed_memory_dir.py --only-recent 5 +- FINAL message = wrap-up JSON only, always: +{"status":"succeeded|partially_succeeded|failed","project":"voice chat #214","phase":"V0 TTS research","actions_taken":[],"actions_failed":[],"files_touched":[],"containers_restarted":[],"recommended":"","runner_up":"","changes_july_decision":true,"top2":[{"engine":"","vram_gb":0,"clone_clip_s":"","streaming":true,"image":"","licence_code":"","licence_weights":"","commercial_ok":false,"turing_ok":true}],"unverified_claims":[],"next_step":"main session writes V2 prompt for the chosen engine","notes":""} diff --git a/claude-config/config/prompts/voice/V1_stt_server01_gpu.md b/claude-config/config/prompts/voice/V1_stt_server01_gpu.md new file mode 100644 index 0000000..95f90f3 --- /dev/null +++ b/claude-config/config/prompts/voice/V1_stt_server01_gpu.md @@ -0,0 +1,56 @@ +# V1 — speech-to-text server on server-01 GPU (personal_projects #214, Voice Chat for Claude Code) + +Written 2026-09-27 (infrastructure general questions conversation). **SUPERSEDES agent A1** in +`/home/administrator/Desktop/claude/agent-builder/voice_chat_agent_prompts.md` — A1/A2 were written to go through +"sudo-bridge-server01 :8082", which is dead AND is being retired (agent-sudo replaces sudo-bridge). None of this needs +sudo: server-01 Docker already has the `nvidia` runtime (nvidia-container-toolkit 1.20.1, driver 550.163.01) and the +`administrator` user is in the docker group, so the STT server is just a GPU container. +**OWNER GATE:** spawn only after the owner has OK'd pulling a new image onto server-01 (shared host). Run it ALONE among +server-01-mutating agents (W2/OLLAMA-1/S1/AS0/JH-1/BOOT-1 are fine alongside — they don't mutate server-01). +Plan of record: `/home/administrator/Desktop/claude/agent-builder/voice_chat_build_plan.md` (read §1–§5, §7). +**Role boundary note (owner-locked 2026-06-25):** new services deploy via a **Jenkins** pipeline and are monitored by **Hermes**. Jenkins is not usable yet (#181, JH-1 audits it), so V1 deploys by hand as a prototype on the test host; the compose is written so it can be dropped into a Gitea repo + `Jenkinsfile.server01` later (list what that pipeline needs in the README), and the README names the health endpoint Hermes should watch. + +--- + +You are a bounded background BUILD agent for personal_projects #214, phase V1 (STT only). Budget: --max-turns 35. +If you issue the same tool call twice with identical arguments, STOP and output the wrap-up block with status=partially_succeeded. + +HARD RULES +- All work on server-01 via `ssh -o BatchMode=yes administrator@192.168.1.90 ''`. No sudo anywhere. No changes on primary except the context.md append + embed run. No git add/commit/push, no Postgres writes, no Vault. +- Container allowlist (rule 9): starts EMPTY. You may add exactly ONE container to it — the new `voice-stt` — and log that elevation in actions_taken. You may create/start/stop/recreate ONLY `voice-stt`. NEVER touch ollama, hermes, jenkins, agent-sudo, n8n-*, *-sandbox, media-pipeline-sandbox-*, incus. You MAY send read-only API requests to server-01's Ollama (`/api/tags`, `/api/ps`, and one `/api/generate` with `"keep_alive":"5m"` to load llama3.1:8b for the VRAM measurement — that is a normal client request, not a container action). +- Write any wait loop as a small script FILE on server-01 run with `timeout` (no nested-quoted `until` one-liners); confirm no leftover waiters before the wrap-up (`ps -eo pid,args | grep "[t]imeout"`). +- Audio privacy (plan §2): never persist audio beyond the test clip; delete it at the end. +- If a hook blocks something, stop with a partial wrap-up. + +FACTS (verified 2026-09-27) +- server-01 GPU: NVIDIA GeForce RTX 2060 SUPER, **8192 MiB**, driver 550.163.01 (→ CUDA 12.4 driver; CUDA 12.x minor-version compatibility applies — if the image's CUDA runtime is newer and fails to init, record the exact error; that decides the fallback). Idle VRAM used ≈ 1 MiB. Ollama container `ollama` holds `llama3.1:8b` (~4.9 GB file) + `nomic-embed-text` (~0.3 GB), loads on demand, default keep-alive 5 min. +- Fast disk: `/data` (1TB SSD, 854 GB free). Existing compose dirs there: `/data/docker-compose/ollama/`. Models dir `/data/models/`. +- Image candidates verified to exist 2026-09-27: PRIMARY `ghcr.io/speaches-ai/speaches:latest-cuda` (faster-whisper, OpenAI-compatible `/v1/audio/transcriptions`; latest release v0.9.0-rc.3, 2025-12-27) — FALLBACK `onerahmet/openai-whisper-asr-webservice:latest-gpu` (with its faster_whisper engine). Do not pick anything else. +- Primary already has an unrelated batch transcriber (`/opt/appdata/docker/docker-compose/whisper/`, meeting recordings) — do NOT touch or reuse it. + +LOCKED DECISIONS +1. Container name `voice-stt`, compose at `/data/docker-compose/voice-stt/docker-compose.yml` on server-01, runtime/`gpus` per the image docs, restart `unless-stopped`, published `192.168.1.90:8300:` (LAN IP bind, NOT 0.0.0.0), model cache bind `/data/models/voice-stt`. Pin the image to an exact tag + digest you resolve (record both). Follow the compose template `/opt/appdata/docker/non-docker-python-scripts/Docker Template/docker-compose.yml` (read it on primary first). +2. Read the chosen image's official docs ONCE (WebFetch speaches.ai docs: installation + configuration + model management) and inline the real env var names in the compose — do not guess (e.g. how to set device=cuda, compute type, model TTL/unload, and how models get downloaded in this version). +3. Model shortlist — test exactly these two, compute type `int8_float16`: (a) `large-v3-turbo` faster-whisper conversion, (b) `distil-large-v3` faster-whisper conversion (use the IDs the image's registry/docs name). Choose the default by: correct transcript of the test clip, then lowest VRAM, then speed. English only. +4. Test clip = whisper.cpp's public `samples/jfk.wav` (download to `/tmp/voice-stt-test/` on server-01; known text: "And so my fellow Americans, ask not what your country can do for you, ask what you can do for your country."). Delete the directory at the end. + +STEPS +1. Pre-flight snapshot: `docker ps --format '{{.Names}}'` (record the set — restore target), `nvidia-smi --query-gpu=memory.used,memory.total --format=csv`, `ss -ltn | grep 8300` (must be free). +2. Write compose; `docker compose config -q`; pull; up. Health: poll the image's health/models endpoint via a timeout'd script. +3. For each model: transcribe the clip 3× (first = cold load), record transcript, wall time, real-time factor, and VRAM (`nvidia-smi --query-compute-apps=pid,process_name,used_memory --format=csv` + total). +4. **Coexistence test** (plan §7): load llama3.1:8b via one Ollama `/api/generate` (tiny prompt, keep_alive 5m) and nomic via one `/api/embeddings`, keep the chosen whisper model loaded, then transcribe again. Record total VRAM, whether anything OOM'd or fell back to CPU (check `docker logs ollama --since 5m | grep -i -E "cuda|offload|memory"` — read-only), and latency change. Leave XTTS out (V2 — not built; budget it as "XTTS-v2 ≈ 2–3 GB, UNVERIFIED"). +5. Write a **VRAM budget** section (if `/opt/appdata/docker/research/jenkins_hermes_state_2026-09-28.md` exists, read which LLM backend Hermes uses — if it is server-01 Ollama, Hermes is a fourth GPU tenant; include it): measured numbers for Ollama-8B + nomic + STT, the remaining headroom for TTS, and 2–3 OPTIONS the main session will choose from (e.g. shorter Ollama keep-alive / `OLLAMA_MAX_LOADED_MODELS`, STT on CPU fallback, smaller LLM) — options only, do not apply any. +6. Leave `voice-stt` RUNNING with the chosen default model configured; verify from PRIMARY: `curl -s -F file=@ -F model= http://192.168.1.90:8300/v1/audio/transcriptions` (or the fallback image's endpoint) returns the transcript. Delete the local copy after. +7. Restore check: the container set equals the step-1 snapshot + `voice-stt`. + +DELIVERABLES +1. `/data/docker-compose/voice-stt/docker-compose.yml` + `README.md` on server-01 (what it is, endpoint, model, how to change model, rollback = `docker compose down`, the exact client call). +2. A copy of both files + the report into the primary repo: `/opt/appdata/docker/docker-compose/voice-stt/` (compose + README) and `/opt/appdata/docker/research/voice_stt_v1_2026-09-28.md` (results table, VRAM budget, options, UNVERIFIED list, the exact invocation the laptop daemon will call). + +SCOPE ALLOWLIST: server-01 `/data/docker-compose/voice-stt/`, `/data/models/voice-stt/`, `/tmp/voice-stt-test/` (deleted), container `voice-stt`; primary `/opt/appdata/docker/docker-compose/voice-stt/`, the report, `/opt/appdata/docker/docker-compose/voice-stt/.claude/context.md`. Nothing else. + +PERSIST BEFORE YOU FINISH +- Create `/opt/appdata/docker/docker-compose/voice-stt/.claude/context.md` (template `/home/administrator/.claude/projects/-opt-appdata-docker/memory/playbook_project_context_template.md`), append "## V1 STT — 2026-09-28 (background agent)". +- Run: python3 /opt/appdata/docker/.claude/scripts/embed_memory_dir.py --only-recent 5 +- FINAL message = wrap-up JSON only, always: +{"status":"succeeded|partially_succeeded|failed","project":"voice chat #214","phase":"V1 STT","actions_taken":[],"actions_failed":[],"files_touched":[],"containers_restarted":[{"name":"voice-stt","action":"","health":""}],"image":{"ref":"","tag":"","digest":"","fallback_used":false},"models":[{"id":"","transcript_ok":true,"rtf":0.0,"vram_mib":0}],"chosen_model":"","coexistence":{"total_vram_mib":0,"oom":false,"ollama_cpu_fallback":false},"endpoint":"http://192.168.1.90:8300/...","unverified_claims":[],"next_step":"","notes":""} diff --git a/claude-config/config/prompts/wireguard/W2_build_prep.md b/claude-config/config/prompts/wireguard/W2_build_prep.md new file mode 100644 index 0000000..779cd38 --- /dev/null +++ b/claude-config/config/prompts/wireguard/W2_build_prep.md @@ -0,0 +1,45 @@ +# W2 BUILD-PREP — cloudflare-ddns + wg-easy IPv4-only fix (#274) — write files, deploy NOTHING + +Tracks: personal_projects #274. Design: `/opt/appdata/docker/research/diy_wireguard_design.md` (LOCKED 2026-09-25). +Written 2026-09-27 (infrastructure general questions conversation) for the 2026-09-28 parallel-agent day. +Safe to run in parallel with S1, AS0, JH-1, BOOT-1, OLLAMA-1 (no shared files). + +--- + +You are a bounded background BUILD-PREP agent for personal_projects #274 (DIY WireGuard home-access tunnel), PHASE W2-PREP. Budget: --max-turns 30. +If you issue the same tool call twice with identical arguments, STOP and output the wrap-up block with status=partially_succeeded. + +HARD RULES +- WRITE FILES ONLY. Do NOT start, create, pull or restart any container; do NOT run `docker compose up`; do NOT install, enable or start any systemd unit; no sudo; no router/DNS/Cloudflare API writes; no git add/commit/push; no Postgres writes; no ntfy sends. The owner deploys after review. +- Read-only on the host otherwise (docker ps/inspect image/ports/networks/labels only — NEVER print `.Config.Env`; never print a secret value). NO Vault writes this phase — both Cloudflare tokens are already stored. +- If a security hook or permission blocks something, stop with a partial wrap-up (do not work around it). Issue each file write as its own call. + +CURRENT STATE (verified 2026-09-27 14:00 by the main session) +- wg-easy `ghcr.io/wg-easy/wg-easy:15.4.0` (digest-pinned in the compose) is Up (healthy) on primary 192.168.1.88, project dir `/opt/appdata/docker/docker-compose/wireguard/`, published `0.0.0.0:45791->45791/udp` AND `[::]:45791->45791/udp`, UI `127.0.0.1:51821`. Laptop client works on LAN. Router forwards UDP 45791 → 192.168.1.88 (done W0). +- Secrets pattern already in use for wireguard: `secretspec.toml` + `deploy/resolver.sh` + `wireguard-secretspec-resolver.{service,timer}` (installed + enabled). Read all of them first — W2 extends this, it does not invent a new pattern. +- Vault (already stored, verified): `secret/cloudflare/ddns` field `token` (Zone:DNS:Edit on reverseproxyserver.net). `secret/cloudflare/traefik-dns01` is for W3 — do NOT wire it now. +- Public IP is dynamic (66.164.12.179 on 09-25, 66.164.11.87 on 09-17). Not CGNAT. + +LOCKED DECISIONS — do not revisit +1. DDNS client = **favonia/cloudflare-ddns**, pinned to an exact current release tag + digest (look it up on Docker Hub / its GitHub releases; record both). Run it as a **second service in the SAME wireguard compose project** (it is part of the tunnel), on the `wg` bridge network (NOT host mode — IPv4 only, so bridge is fine). +2. Record = `fuppzd3f9v.reverseproxyserver.net`, **DNS-only (PROXIED=false)**, **IPv4 only (IP6_PROVIDER=none)**, TTL 60, update every 2 minutes (favonia's `UPDATE_CRON=@every 2m` or whatever its README documents). The DDNS client must manage ONLY that one record (never the zone apex or other hosts; if favonia offers a "delete on stop" option, set it OFF so a stopped container never deletes the record). +3. Token via SecretSpec: add `CLOUDFLARE_API_TOKEN` (or the `_FILE` variant if favonia recommends it and the resolver pattern supports it — prefer whatever keeps the value out of `docker inspect`; explain your choice) sourced from Vault `secret/cloudflare/ddns` field `token`, exactly the way INIT_PASSWORD is mapped today. +4. Hardening per favonia's README: `read_only: true`, `cap_drop: [ALL]`, `security_opt: [no-new-privileges:true]`, non-root `user:` as documented, restart policy per template, healthcheck only if the image supports one (do not invent one that can't run in the image — check the image for a shell first via `docker manifest inspect`/docs; if none, omit and say so). +5. **wg-easy IPv4-only publish fix:** change the port mapping to `"0.0.0.0:45791:45791/udp"` so Docker stops publishing on `[::]`. Edit only that line (+ a one-line comment). This takes effect at the owner's next recreate. +6. Compose MUST keep following the template `/opt/appdata/docker/non-docker-python-scripts/Docker Template/docker-compose.yml` (read it first). Remove the W2 placeholder comment the W1 agent left for ddns (keep the W3 dnsmasq placeholder). + +VERIFY FAVONIA FACTS FROM THE SOURCE — WebFetch https://github.com/favonia/cloudflare-ddns (README) once and use its exact env var names (e.g. CLOUDFLARE_API_TOKEN / _FILE, DOMAINS, PROXIED, IP6_PROVIDER, TTL, UPDATE_CRON, DELETE_ON_STOP). Do not guess names from memory. + +DELIVERABLES +1. Edited `/opt/appdata/docker/docker-compose/wireguard/docker-compose.yml` (ddns service + IPv4-only fix). +2. Edited `secretspec.toml` (+ `deploy/resolver.sh` only if the new var needs a resolver change — keep it minimal). +3. `README.md` — add a "W2 — owner steps" section: (a) how to apply: run the resolver service once (`sudo systemctl start wireguard-secretspec-resolver.service`), confirm both containers healthy/running, `docker logs` of the ddns container shows the record set to the current public IP (compare `curl -4 -s https://ifconfig.me`), `dig +short fuppzd3f9v.reverseproxyserver.net @1.1.1.1`; `ss -lnup | grep 45791` shows IPv4 only; (b) phone (Android 9) + tablet: create clients in the wg-easy UI, scan QR, AllowedIPs split-tunnel per design; (c) the W2 tests from design §5: off-LAN handshake over cellular, the **silence test** (from off-LAN, `nmap -sU -p 45791 ` shows open|filtered and wg-easy logs nothing for a non-peer), and the IP-change drill (how to force a DDNS re-check); (d) rollback (remove the ddns service, `docker compose up -d` re-applies; the DNS record stays pointing at the last IP — harmless). +4. Validate without deploying: `docker compose -f config -q` (unset secretspec vars warnings expected — note them), `bash -n deploy/resolver.sh` if touched. + +SCOPE ALLOWLIST: `/opt/appdata/docker/docker-compose/wireguard/` (docker-compose.yml, secretspec.toml, deploy/resolver.sh, README.md, .claude/context.md). Nothing else. + +PERSIST BEFORE YOU FINISH +- Append (one `cat >>` heredoc; append only) "## #274 W2-prep — 2026-09-28 (background agent)" to `/opt/appdata/docker/docker-compose/wireguard/.claude/context.md`: What was done / Current state / Owner steps (short) / Next step (W3 dnsmasq + Traefik DNS-01 prep). +- Run: python3 /opt/appdata/docker/.claude/scripts/embed_memory_dir.py --only-recent 5 +- FINAL message = wrap-up JSON only, always: +{"status":"succeeded|partially_succeeded|failed","project":"diy-wireguard #274","phase":"W2-prep","actions_taken":[],"actions_failed":[],"files_touched":[],"containers_restarted":[],"unverified_claims":[],"ddns_image_tag":"","ddns_image_digest":"","token_mapping":"env|file — why","validation":{"compose_config":"","resolver_syntax":""},"owner_steps":[],"next_step":"","notes":""} diff --git a/claude-config/decisions/DECISIONS.md b/claude-config/decisions/DECISIONS.md index 3ad14cd..7896ce1 100644 --- a/claude-config/decisions/DECISIONS.md +++ b/claude-config/decisions/DECISIONS.md @@ -127,6 +127,22 @@ entry links back to the research that drove it, so the reasoning survives. - **Implementation:** n/a (nothing to install); three steal rows in personal_projects. - **Revisit when:** Anthropic changes OAuth terms; a candidate ships Linux-native local voice; MEMORY.md cap still hurts after the defrag skill exists. +## 2026-09-27 — DECIDED (owner, infra general questions): agent-sudo/secrets-proxy/voice/sandbox calls +- **Decision:** adopt all recommendations: + (1) **sudo-bridge is RETIRED in favour of agent-sudo** (owner: sudo-bridge = security theatre); nothing may route through or depend on it. + (2) **Jenkins → primary deploy path = a narrow agent-sudo `/exec` deploy-service op that runs the SecretSpec resolver**, NOT raw `docker compose` through secrets-proxy `/shell` (that path never worked). + (3) **#193: keep agent-sudo fail-closed** + add Vault-wait/backoff in the one human unit edit; replace the symlinked unit with a root-owned copy. + (4) **Narrow the two over-broad tier-0 read rules** (SUDO.md l.39, l.53) and re-sign. + (5) **#192 leaked key in git history: accept + document** after rotation (dead key); rotate the Gitea remote token + Coolify token too. + (6) **Re-check the TTS engine** before V2 (XTTS-v2 upstream defunct, non-commercial licence). + (7) V1 STT image pull onto server-01 approved. (8) Single Ollama = server-01 GPU. + (9) **vault-sandbox must have the same auto-unseal as primary Vault** (vault-watch-unseal.service + root-only key file); copy it over if missing — vault-sandbox restart deferred until then. + (10) secrets-proxy 07-09 stop = during the agent-sudo redesign (when sudo-bridge was found inadequate and agent-sudo chosen) — owner recollection, no written record. +- **Why:** agents AS0/S1/BOOT-1 findings 2026-09-27 (tiers 1–3 not implementable as built; secrets-proxy holes; boot-order port/DNS loss on server-01). +- **Research:** `/opt/appdata/docker/research/{agent_sudo_readiness_2026-09-28,secrets_proxy_S1_investigation}.md`, `/opt/appdata/docker/boot-recovery/README.md` +- **Implementation:** pending — AS1 prompt, corrected SP1, voice V1/TTS research, vault-sandbox auto-unseal. +- **Revisit when:** agent-sudo tiers 1–3 are live on both hosts (re-evaluate secrets-proxy's remaining scope). + ## 2026-09-08 — DECIDED (grill-me 22:50): AI-harness LAYER scope, rubric, hard rules, trial depth - **Decision:** ADOPT the framing — research the **layer around Claude Code** (memory/skills/hooks/routing/voice), not the executor; premise "we run LifeOS" is FALSE (nothing installed). Rubric = 4 pains (200K/80% checklist; ~24.4 KB MEMORY.md cap; rule adherence beyond the skills track; voice + life ops). Hard rules = R1 Pro OAuth/no API key, R2 merge-not-replace our hooks, R3 free + OSS. Distro = soft factor (LMDE 7 now; Arch/Omarchy or Fedora possible; **Ubuntu never**). Depth = research then ONE sandboxed trial in the server-01 Incus VM. Voice rides on the harness verdict. Fixed candidates + native-baseline row (baseline may win). - **Why:** every candidate is the same category as the layer we already built; the only honest comparison is against that baseline plus Claude Code's own newer features. diff --git a/claude-config/research/INDEX.md b/claude-config/research/INDEX.md index e250a16..9eff4d7 100644 --- a/claude-config/research/INDEX.md +++ b/claude-config/research/INDEX.md @@ -25,6 +25,8 @@ re-verify versions, repos, and Linux support before acting on anything. | **Partner access audit — Austin offboard / Hailee onboard** (preflight K, read-only) — every place departed partner Austin has access + every place new 33% partner Hailee must be added; offboard/onboard checklists | [partner-access-audit-austin-hailee.md](partner-access-audit-austin-hailee.md) | Audited 2026-09-15 (Opus 4.8) | **ZERO changes made.** Austin access = Nextcloud user `Austin_Mktg` (enabled) + `/Partner Meetings` **share id 2** (only that). No groups, **no Vault refs**. Bitwarden bridge DOWN → manual owner check. Talk rooms not occ-listable → disabling the user covers Talk (verify in UI). External (Discord/email) = manual. Hailee: new NC user + `/Partner Meetings` share + Talk room + Twingate Partners group. NOT Grim/Josh. | | **Twingate deploy runbook** (preflight K) — partner access to Nextcloud over Tyler's CGNAT; free-tier, connector container shape, Remote-Network/Resource/Group model, client requirement, Vault paths | [twingate-deploy-runbook.md](twingate-deploy-runbook.md) | Written 2026-09-15 (Opus 4.8) | **Free tier fits** (5 users/10 nets/50 resources; partners=Tyler+Hailee). Connector `twingate/connector` on PRIMARY server, **outbound-only → CGNAT-proof, no port-forward**. Client app REQUIRED (no clientless). Vault `secret/twingate/*` (none exist yet). **UNVERIFIED: Talk WebRTC fully over Twingate + browser LNA** → test before dropping Open Relay TURN. | | **NetBird self-host buildplan + bootstrap decision aid** (preflight K) — owner device mesh; resolves the 3 open items + bootstrap-vs-selfhost recommendation + server-01 sandbox runbook + primary-server prod shape | [netbird-selfhost-buildplan.md](netbird-selfhost-buildplan.md) | Written 2026-09-15 (Opus 4.8) | **(a) DB: SQLite default, Postgres supported → use homelab Postgres.** **(b) OIDC: Authelia WORKS via generic-OIDC (official Authelia↔NetBird guide) — no Authentik;** device-flow for CLI UNVERIFIED (setup-keys fallback). **(c) TURN: coturn needs a UDP port; WSS relay on 443 may avoid it (UNVERIFIED — test on server-01).** **Recommend BOOTSTRAP on NetBird Cloud free tier first** (deciding factor: 09-20 Twingate deadline owns this week). | +| **secrets-proxy S1 investigation** (#191 → #174, read-only) — why it stopped 07-09, endpoint/auth inventory, 18 compose vars → Vault, callers (Jenkins promote), SP1–SP4 validity, corrected SP1 draft | [secrets_proxy_S1_investigation.md](/opt/appdata/docker/research/secrets_proxy_S1_investigation.md) | Investigated 2026-09-27 (Opus 5.5) | **Do NOT restart as-is** (`bash -c` gate bypass, docker.sock = host root, env leak). SP plan partly valid. **Owner: pick promote path** (`/exec deploy_service` recommended) before SP1; verify transit key `secrets-proxy-proxymd` (UNVERIFIED). | +| **Voice Chat TTS engine re-check (V0, #214)** — XTTS-v2 (idiap) vs F5, Chatterbox, Fish S2, CosyVoice 3, IndexTTS 2.5, Zonos, Qwen3-TTS + Kokoro/Piper baselines under R1 local / R2 clone / R3 8 GB Turing / R4 stream / R5 maintained / R6 licence | [/opt/appdata/docker/research/voice_tts_engine_2026-09-28.md](/opt/appdata/docker/research/voice_tts_engine_2026-09-28.md) | Researched 2026-09-28 (background agent, Opus 5.5) | **Switch primary to Chatterbox-Turbo** (MIT code+weights, devnen server `/tts` stream, fp32 on Turing) ; XTTS-v2 = fallback (CPML non-commercial, Coqui defunct). Changes July decision. V2 must smoke-test VRAM + first-audio on the 2060 SUPER. | ### Where the prior infrastructure research lives (memory corpus) Not duplicated here — cited in [infrastructure-synthesis.md](infrastructure-synthesis.md). Key files: