docs(prompts): 09-29 agent prompts — AS2, VP-1, W3 build prep, V2 + V2b Chatterbox
Held back: L_twingate-partner-swap-deploy.md (standing holdback). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,70 @@
|
||||
# AS2 — agent-sudo SUDO.md rule proposal + classifier tests (files only)
|
||||
|
||||
--max-turns 15. You are one agent of a batch; the main session reports all results together after the whole batch
|
||||
finishes — do not address the owner directly in your final message; write for the main session.
|
||||
If you issue the same tool call twice with identical arguments, STOP and output the wrap-up block with status=partially_succeeded.
|
||||
|
||||
## Goal
|
||||
Write `/opt/appdata/docker/docker-compose/agent-sudo/bridge/SUDO.md.proposed` — a full copy of the current
|
||||
`bridge/SUDO.md` with the AS0 Task-2 rule changes applied — plus classifier tests proving each new/narrowed rule
|
||||
classifies as intended. The owner reviews, edits, re-signs (`security/sudo_sign.py`) and a human restarts the daemon
|
||||
later. You do NOT edit SUDO.md, sign, restart anything, or commit.
|
||||
|
||||
## Background (verified 2026-09-29 by the main session)
|
||||
- Rule table format in SUDO.md: `| pattern | tier | state | host_override | verb | reason |`; states `active|proposed`.
|
||||
- Source of the changes: `/opt/appdata/docker/research/agent_sudo_readiness_2026-09-28.md` lines 30–55 (Task 2).
|
||||
Read it. Rules to ADD, all `state=active`, both hosts unless noted:
|
||||
- t1 exact-name: `systemctl enable --now control-plane-up.timer`, `systemctl start *-secretspec-resolver.service`,
|
||||
`systemctl enable --now *-secretspec-resolver.timer`, `systemctl disable --now *-secretspec-resolver.timer`.
|
||||
- t1 `systemctl daemon-reload` (undo = noop).
|
||||
- t1 `install -m 0755 -o root -g root /opt/appdata/docker/docker-compose/*/deploy/* /usr/local/sbin/`.
|
||||
- t1 `cp /opt/appdata/docker/docker-compose/*/deploy/*-secretspec-resolver.service /etc/systemd/system/`, same for
|
||||
`.timer`, and `cp /opt/appdata/docker/boot-recovery/control-plane-up.* /etc/systemd/system/` (NOTE: AS0 wrote
|
||||
`…/deploy/control-plane-up.*` but the real files live in `/opt/appdata/docker/boot-recovery/` — use the real path).
|
||||
- t3 `docker compose -f /opt/appdata/docker/docker-compose/*/docker-compose.yml up -d`, host_override `primary=4`.
|
||||
- t0 `smartctl -a /dev/sd?`, `smartctl -a /dev/nvme?n1`, `smartctl -H /dev/*`; t0 `id`.
|
||||
- NARROW `cat /etc/*` (t0 active) → replace with explicit t0 rules: `cat /etc/systemd/system/*`,
|
||||
`cat /etc/sysctl.d/*`, `cat /etc/fstab`, `cat /etc/control-plane-up/*`, `cat /etc/docker/daemon.json`; and add
|
||||
`cat /etc/shadow*`, `cat /etc/sudoers*`, `cat /etc/agent-sudo/*` → tier 4 ABOVE them (first match wins — verify
|
||||
the classifier's match order in `sudo_rules.py` before relying on it).
|
||||
- NARROW `docker inspect *` (t0, exposes every container's Config.Env secrets): propose
|
||||
`docker inspect --format * *` stays t0 only if the classifier can refuse `.Config.Env`; otherwise move plain
|
||||
`docker inspect *` to t4 and add a t0 `docker inspect --format '{{.State.Status}}' *` style set. Pick one, justify.
|
||||
- DECISION for you to document, not make: whether `systemctl start *-secretspec-resolver.service` on PRIMARY is t1
|
||||
(unlocks Jenkins `/deploy` there). Write it as t1 with a `reason` flag "OWNER DECIDES primary" and list it in the
|
||||
wrap-up `owner_decisions`.
|
||||
- AS1 (uncommitted, 157 tests passing) added daemon-side undo, tier-1 undo capture, `/deploy`, multi-caller auth. Read
|
||||
`AS1_CHANGES.md` first; do not modify AS1 code except to add undo kinds strictly required by a new rule (`noop` for
|
||||
daemon-reload, inverse-verb for systemctl enable/disable/start) — if those already exist, reuse them.
|
||||
- Set B gap: `_is_write_to` matches a literal unit path; `cp …/x.service /etc/systemd/system/` (dir destination) is not
|
||||
caught. Keep patterns name-anchored and ADD a test showing the gap is closed or explicitly still open.
|
||||
|
||||
## Scope allowlist (touch nothing else)
|
||||
- CREATE `bridge/SUDO.md.proposed`, `AS2_CHANGES.md`, `test_as2_rules.py` in `/opt/appdata/docker/docker-compose/agent-sudo/`.
|
||||
- EDIT only if a required undo kind is missing: `security/scoped_undo.py`, `daemon_ops.py` (+ their tests).
|
||||
- APPEND `/opt/appdata/docker/docker-compose/agent-sudo/.claude/context.md` (dated handoff).
|
||||
- Do NOT touch `bridge/SUDO.md`, signatures, `.env`, the running daemon/container, or anything outside that dir.
|
||||
|
||||
## Steps
|
||||
1. Read SUDO.md, sudo_rules.py (matching order + how tiers/host_override resolve), AS1_CHANGES.md, the AS0 report lines 30–55.
|
||||
2. Write SUDO.md.proposed (copy + changes; mark every changed row's reason with `AS2:`).
|
||||
3. Write test_as2_rules.py: for each added/narrowed pattern, a positive case (expected tier per host) and a near-miss
|
||||
negative (e.g. `cat /etc/shadow` → t4, `systemctl enable --now evil.timer` → NOT t1, dir-destination cp).
|
||||
Tests must load rules from SUDO.md.proposed (not SUDO.md) — if the loader verifies signatures, load unsigned via
|
||||
the existing test path the current tests use; read test_sudo_rules.py for the pattern.
|
||||
4. Run the FULL suite: `cd /opt/appdata/docker/docker-compose/agent-sudo && python3 -m pytest -q`. Report pass/fail counts.
|
||||
All 157 prior tests must still pass.
|
||||
5. Write AS2_CHANGES.md: table of every rule change (old → new, tier, why), the owner decisions, and the exact owner
|
||||
steps to adopt (review diff → copy over SUDO.md → `python3 security/sudo_sign.py` usage as found in the code → daemon
|
||||
restart happens in the hardening pass, not now).
|
||||
6. Wrap-up.
|
||||
|
||||
## Wrap-up (always, success or failure)
|
||||
Append the dated handoff to the project `.claude/context.md` (What was done / Decisions / Current state / Next step).
|
||||
Run `python3 /opt/appdata/docker/.claude/scripts/embed_memory_dir.py --only-recent 3`. No commits, no pushes, no
|
||||
restarts, never print secrets. Final message = this JSON:
|
||||
```
|
||||
{"status":"succeeded|partially_succeeded|failed","project":"agent-sudo #176 AS2","actions_taken":[],"actions_failed":[],
|
||||
"files_touched":[],"containers_restarted":[],"rules_added":<int>,"rules_narrowed":<int>,"tests":{"passed":<int>,"failed":<int>},
|
||||
"owner_decisions":[],"set_b_gap":"closed|open — why","next_step":"","notes":""}
|
||||
```
|
||||
@@ -0,0 +1,61 @@
|
||||
# VP-1 — port the proven VS-1 auto-unseal to PRIMARY Vault (#182) — files only, owner installs
|
||||
|
||||
--max-turns 15. You are one agent of a batch; the main session reports all results together after the whole batch
|
||||
finishes — do not address the owner directly in your final message; write for the main session.
|
||||
If you issue the same tool call twice with identical arguments, STOP and output the wrap-up block with status=partially_succeeded.
|
||||
|
||||
## Goal
|
||||
Stage replacement unseal scripts + units for primary's production Vault, built from the VS-1 design that was proven
|
||||
twice on server-01's vault-sandbox (2026-09-28), plus an owner runbook. You change NOTHING live.
|
||||
|
||||
## Background (verified 2026-09-29)
|
||||
- Reference (proven): `/opt/appdata/docker/vault-sandbox-autounseal/` (README.md, vault-sandbox-unseal.sh/.service,
|
||||
vault-sandbox-watch-unseal.sh/.service). Method: unseal key sent via HTTP API from INSIDE the container over stdin:
|
||||
`printf '{"key":"%s"}' "$key" | docker exec -i <c> sh -c "wget -q -O - --header 'Content-Type: application/json' --post-file=/dev/stdin http://127.0.0.1:8200/v1/sys/unseal"`
|
||||
— key never on argv. Watcher filter: `docker events --filter label=com.docker.compose.service=<svc>` plus an in-loop
|
||||
name check (the `name=` events filter key is silently ignored — that is the bug). Seal state from `vault status`
|
||||
exit code: 0 unsealed, 2 sealed, 1 not ready.
|
||||
- Primary live (read only): `/usr/local/bin/vault-unseal.sh` (+ `.bak`), `/usr/local/bin/vault-watch-unseal.sh`
|
||||
(uses the broken `name=` filter), units `/etc/systemd/system/vault-unseal.service` (boot oneshot, FAILED since
|
||||
2026-09-07) and `vault-watch-unseal.service` (running). Repo copies: `/opt/appdata/docker/docker-compose/vault/vault-unseal.sh`,
|
||||
`vault-unseal.service`. Container `vault-iwaulpoi5hwirdlogshmul40`, compose service label `vault`, project `vault`,
|
||||
`init: true` applied 09-28. Vault UI at 192.168.1.88:8200.
|
||||
- Read the existing scripts to learn WHERE the key comes from — do NOT read, print, copy or log the key itself, and do
|
||||
not open any file that holds it. Reference it only by path. Find out why the boot oneshot fails (journalctl -u
|
||||
vault-unseal.service --no-pager -n 50 — output may mention paths, never values).
|
||||
- ⚠ `vault-selfheal.timer` runs `docker compose up -d` on `/opt/appdata/docker/docker-compose/vault/docker-compose.yml`
|
||||
unattended — never edit that compose. You don't need to.
|
||||
- Primary is production: no restarts, no `docker restart vault…`, no unseal tests against live Vault. The owner runs
|
||||
the install and the restart test.
|
||||
|
||||
## Scope allowlist
|
||||
- CREATE `/opt/appdata/docker/docker-compose/vault/deploy/autounseal/` with: `vault-unseal.sh`, `vault-watch-unseal.sh`,
|
||||
`vault-unseal.service`, `vault-watch-unseal.service`, `README.md` (owner runbook).
|
||||
- APPEND `/opt/appdata/docker/docker-compose/vault/.claude/context.md` (create from the project template if missing).
|
||||
- Nothing else. No sudo, no systemctl mutations, no container actions.
|
||||
|
||||
## Steps
|
||||
1. Read the VS-1 reference files and README, primary's live scripts/units (`systemctl cat vault-unseal.service
|
||||
vault-watch-unseal.service`), and the failing oneshot's journal.
|
||||
2. Write the four staged files: same key source as today's primary scripts, VS-1 HTTP-stdin method, label filter
|
||||
`com.docker.compose.service=vault` + in-loop check that the container name starts with `vault-` and is NOT
|
||||
`vault-sandbox`, exit-code seal detection, bounded retries, logs to journal without the key.
|
||||
3. Validate with the real tools: `bash -n` each script, `shellcheck` if installed, and
|
||||
`systemd-analyze verify <staged unit files>` (it may warn about missing ExecStart paths since files aren't installed
|
||||
— record warnings, fix real errors).
|
||||
4. README.md runbook, commands IN FULL, no `…`: backup current scripts/units (`.pre-vp1`), `sudo install` the staged
|
||||
files, `sudo systemctl daemon-reload`, `sudo systemctl restart vault-watch-unseal.service`, then the proof test
|
||||
(owner-run): `docker restart vault-iwaulpoi5hwirdlogshmul40`, wait 30 s,
|
||||
`curl -s http://192.168.1.88:8200/v1/sys/health` → `"sealed":false`; journal check; rollback commands; note that
|
||||
this unblocks the BOOT-1 primary reboot test (#219).
|
||||
5. Wrap-up.
|
||||
|
||||
## Wrap-up (always)
|
||||
Append the dated handoff to the vault project `.claude/context.md`. Run
|
||||
`python3 /opt/appdata/docker/.claude/scripts/embed_memory_dir.py --only-recent 3`. No commits/pushes/restarts; never
|
||||
print secrets. Final message = JSON:
|
||||
```
|
||||
{"status":"succeeded|partially_succeeded|failed","project":"#182 VP-1 primary vault auto-unseal","actions_taken":[],
|
||||
"actions_failed":[],"files_touched":[],"containers_restarted":[],"oneshot_failure_cause":"","key_source":"<path only>",
|
||||
"validation":{"bash_n":"","shellcheck":"","systemd_analyze":""},"next_step":"","notes":""}
|
||||
```
|
||||
@@ -21,13 +21,15 @@ PRIMARY'S MECHANISM (verified 2026-09-27 — this is the pattern to copy)
|
||||
- `/usr/local/bin/vault-unseal.sh` (root 0755): waits ≤120 s for a running container matching `^vault-`, and if `docker exec <c> vault status` shows `Sealed true`, reads keys line-by-line from a root-only file (`/etc/vault-unseal-keys`, root:root 0400, 3 keys for a 3-of-5 threshold; `#`/blank lines skipped) and runs `docker exec <c> vault operator unseal <key>` per key (output to /dev/null).
|
||||
- `/usr/local/bin/vault-watch-unseal.sh` (root 0755): `docker events --filter 'name=vault-' --filter 'event=start' --format '{{.Actor.Attributes.name}}'` → on each start, sleep 8, run vault-unseal.sh. Handles container restarts/recreates, not just daemon restarts.
|
||||
- Units: `vault-watch-unseal.service` (Type=simple, Restart=always, RestartSec=5, After/Requires docker.service, WantedBy multi-user) — ACTIVE; `vault-unseal.service` (oneshot, After docker) — currently FAILED on primary (note it in the report; do not fix primary).
|
||||
- Known weakness to NOT copy blindly: the key is passed as a positional argument to `vault operator unseal` (visible in the host process list for the moment it runs). Improve it for the sandbox copy: feed the key on stdin (`printf '%s\n' "$KEY" | docker exec -i <c> vault operator unseal -` — verify the `-` stdin form in `vault operator unseal -h` inside the vault-sandbox container, read-only).
|
||||
- **ALREADY PRESENT ON server-01 (personal_projects #182, verified 2026-07-15):** `/usr/local/bin/vault-sandbox-unseal.sh` (0750) with keys in `/etc/vault-sandbox-unseal-keys` (0600 root, **1-of-1**). Task 1 must establish what EXISTS vs what is MISSING (the watcher unit? enabled? does it fire on container restart?) — the job is to fill the gaps to match primary, not rebuild from scratch.
|
||||
- **DO NOT use `vault operator unseal -` (stdin).** #182: tried 2026-07-15 on both hosts, rolled back the same day — Vault 2.0.x parses the literal "-" as the key (400 "key must be a valid hex or base64 string"). Known-dead; never retry it. The recorded CANDIDATE (untested) is the HTTP API: `printf '{"key":"%s"}' "$KEY" | docker exec -i <c> sh -c "curl -s -X PUT -d @- http://127.0.0.1:8200/v1/sys/unseal"` (check curl/wget exists in the image first). Write it as the new script's method but mark it UNTESTED; the owner's README test (restart vault-sandbox → watcher fires → sealed=false) is the proof, sandbox FIRST, primary only after (#182).
|
||||
- Known weakness to NOT copy blindly: the key is passed as a positional argument to `vault operator unseal` (visible in the host process list for the moment it runs). Fix it for the sandbox copy with the HTTP-API candidate above (NOT the CLI `-` form).
|
||||
|
||||
TASKS
|
||||
1. **Discover on server-01 (read-only):** does any unseal automation already exist there? (`systemctl list-units --all | grep -i unseal`, `ls -la /usr/local/bin/*unseal* /etc/*unseal* 2>&1`, root crontab is not readable — say so). vault-sandbox: container name, image/version, compose working dir (label), seal type (`docker exec <c> vault status -format=json` → print ONLY `sealed`, `initialized`, `t`, `n`, `type`, `version` fields), Vault address it listens on, restart policy.
|
||||
2. **Find where the sandbox unseal keys live — by NAME only:** search memory (`/home/administrator/.claude/projects/-opt-appdata-docker/memory/`, especially `session_summary_n8n_sandbox_deploy.md`, `playbook_vault_token_rotation.md`, `project_hashicorp_vault.md`, `feedback_sandbox_isolation.md`) and context files for where the 2026-06-16 sandbox init stored them (a Bitwarden item? a primary-Vault path? a file on server-01?). Report item/path + field names + key count + threshold. If unknown → UNVERIFIED and say the owner must locate them.
|
||||
3. **Write the adapted files** into the primary repo dir `/opt/appdata/docker/vault-sandbox-autounseal/`:
|
||||
- `vault-sandbox-unseal.sh` — same logic as primary's, but container match `^vault-sandbox-`, key file `/etc/vault-sandbox-unseal-keys`, key fed via stdin, logs to the journal, exits non-zero if still sealed after applying keys.
|
||||
- `vault-sandbox-unseal.sh` — same logic as primary's, but container match `^vault-sandbox-`, key file `/etc/vault-sandbox-unseal-keys`, key sent via the HTTP API JSON body (never argv, never CLI `-`), logs to the journal, exits non-zero if still sealed after applying keys.
|
||||
- `vault-sandbox-watch-unseal.sh` — docker events filter that matches ONLY the sandbox container (verify the events `name` filter semantics — substring vs exact — from Docker docs, and use a filter + in-loop name check so it can never fire on some other vault).
|
||||
- `vault-sandbox-watch-unseal.service` (same shape as primary's watch unit) and a oneshot `vault-sandbox-unseal.service` run at boot After=docker.service (covers "container already up before the watcher started").
|
||||
- `README.md`: owner steps — (a) create `/etc/vault-sandbox-unseal-keys` root:root 0400 from the key source found in task 2 WITHOUT the key touching shell history or the screen (e.g. `sudo install -m 0400 /dev/null /etc/vault-sandbox-unseal-keys && sudo nano /etc/vault-sandbox-unseal-keys` and paste from the password manager), (b) install scripts + units, `daemon-reload`, `enable --now` the watcher and the oneshot, (c) TEST: `docker restart <vault-sandbox>` then `journalctl -u vault-sandbox-watch-unseal -n 20` shows "unsealed" and `vault status` sealed=false, (d) then tell the main session so BOOT-1's server-01 `host.env` can drop vault-sandbox from CHECK_ONLY, (e) rollback.
|
||||
@@ -39,4 +41,4 @@ PERSIST BEFORE YOU FINISH
|
||||
- Append "## VS-1 prep — <date> (background agent)" to `/opt/appdata/docker/vault-sandbox-autounseal/.claude/context.md`.
|
||||
- Run: python3 /opt/appdata/docker/.claude/scripts/embed_memory_dir.py --only-recent 5
|
||||
- FINAL message = wrap-up JSON only, always:
|
||||
{"status":"succeeded|partially_succeeded|failed","project":"vault-sandbox auto-unseal (#270/#219)","phase":"VS-1 prep","actions_taken":[],"actions_failed":[],"files_touched":[],"containers_restarted":[],"server01_existing_unseal":"none|partial|present","vault_sandbox":{"container":"","version":"","sealed":true,"initialized":true,"threshold":"t-of-n","seal_type":""},"key_source":{"where":"","fields":[],"count":0,"verified":false},"stdin_unseal_supported":true,"validation":{"bash_n":"","systemd_verify":""},"unverified_claims":[],"next_step":"","notes":""}
|
||||
{"status":"succeeded|partially_succeeded|failed","project":"vault-sandbox auto-unseal (#270/#219)","phase":"VS-1 prep","actions_taken":[],"actions_failed":[],"files_touched":[],"containers_restarted":[],"server01_existing_unseal":"none|partial|present","vault_sandbox":{"container":"","version":"","sealed":true,"initialized":true,"threshold":"t-of-n","seal_type":""},"key_source":{"where":"","fields":[],"count":0,"verified":false},"existing_on_server01":{"script":"","keys_file":"","watcher_unit":""},"http_api_unseal_tool_in_image":"curl|wget|none","validation":{"bash_n":"","systemd_verify":""},"unverified_claims":[],"next_step":"","notes":""}
|
||||
|
||||
@@ -0,0 +1,73 @@
|
||||
# V2 — Chatterbox-Turbo TTS on server-01 (build + measure, prototype deploy) — #214
|
||||
|
||||
--max-turns 25 (reason: image build + model download over WiFi + measurements; stop at 25 regardless).
|
||||
You are one agent of a batch; the main session reports all results together after the whole batch finishes — do not
|
||||
address the owner directly in your final message; write for the main session.
|
||||
If you issue the same tool call twice with identical arguments, STOP and output the wrap-up block with status=partially_succeeded.
|
||||
If a download/build makes no progress for 10 minutes, stop it and report partially_succeeded (do not sit idle).
|
||||
|
||||
## Goal
|
||||
A running `voice-tts` container on server-01 serving Chatterbox-Turbo with the owner-approved JARVIS reference voice,
|
||||
plus MEASURED VRAM/latency numbers that decide whether it can share the GPU. Decision is pre-made by the owner:
|
||||
**Option A time-share** — Ollama keeps its default idle unload (5 min); Chatterbox is the LOWEST-priority GPU tenant.
|
||||
If it cannot fit without changing other services, report "PARK recommended" — never tune Ollama/STT to make room.
|
||||
|
||||
## Background (verified 2026-09-29)
|
||||
- server-01 = 192.168.1.90, `ssh administrator@192.168.1.90` (admin is in the docker group; no sudo needed or allowed).
|
||||
RTX 2060 SUPER 8 GB, NVIDIA driver 550 (CUDA ≤ 12.4). Right now 436 MiB used (models idle-unloaded after a reboot).
|
||||
Other GPU tenants: `ollama` (image ollama-fixed:1.0.0, models llama3.1:8b + nomic-embed ≈ 5.5 GB when loaded, default
|
||||
keep_alive 5 min) and `voice-stt` (speaches, distil-large-v3 ≈ 922 MiB, `/data/docker-compose/voice-stt/`, :8300).
|
||||
Shared host (hermes, jenkins, n8n-prod, sandbox stack, agent-sudo) — touch none of them.
|
||||
- server-01 is on flaky USB WiFi: large copies from primary corrupt. **server-01 must pull/build everything itself**
|
||||
(git clone + docker build ON server-01). Only the 1 MB reference WAV is copied from primary (verify sha256 after).
|
||||
- Engine research + install path: `/opt/appdata/docker/research/voice_tts_engine_2026-09-28.md` §1 and §3a — read it.
|
||||
Server: github.com/devnen/Chatterbox-TTS-Server, use its **CUDA 12.1** NVIDIA Dockerfile/compose variant (NOT
|
||||
12.8/13.0 — driver 550). Pin the repo at the current commit (record the SHA). fp32 default (bf16 is opt-in; Turing
|
||||
has no bf16). Endpoints per research: `/tts` (stream:true), OpenAI `/v1/audio/speech`, `/api/unload`; port 8004
|
||||
UNVERIFIED — read the server config. Every output carries the Perth watermark (fine for personal use).
|
||||
- Reference voice (owner-approved): primary `/opt/appdata/docker/Machines/infrastructure general questions/jarvis_ref_11s.wav`
|
||||
(alt `jarvis_ref_16s.wav`). PERSONAL USE ONLY — Marvel/Paul Bettany voice; note that in the README.
|
||||
- Pattern to copy: `/opt/appdata/docker/docker-compose/voice-stt/README.md` (hand-deployed prototype at
|
||||
`/data/docker-compose/<name>/` on server-01, repo copy on primary at `/opt/appdata/docker/docker-compose/<name>/`,
|
||||
LAN-IP bind only, models cached under `/data/models/<name>`, pinned image, README with endpoint/health/rollback).
|
||||
- Before building, WebFetch the devnen repo README + its docker docs and inline the exact env/config keys you use.
|
||||
|
||||
## Container allowlist (rule 9 — starts EMPTY)
|
||||
You may create/start/stop/recreate ONLY `voice-tts` (after logging the elevation in actions_taken). You may send normal
|
||||
API requests to `ollama` (e.g. load llama3.1:8b + nomic to measure coexistence) and `voice-stt` (a transcription), but
|
||||
never restart, recreate or reconfigure them. Anything else = surface, don't touch.
|
||||
|
||||
## Scope allowlist
|
||||
- server-01: `/data/docker-compose/voice-tts/` (clone + compose), `/data/models/voice-tts/`, `/data/voices/jarvis/`.
|
||||
- primary: CREATE `/opt/appdata/docker/docker-compose/voice-tts/` (docker-compose.yml copy, README.md, `.claude/context.md`);
|
||||
CREATE `/opt/appdata/docker/research/voice_tts_v2_2026-09-29.md` (measurements report).
|
||||
- Nothing else. No commits, no pushes, no sudo, no Gitea image push (list it as a follow-up).
|
||||
|
||||
## Steps
|
||||
1. Read the research §1/§3a, voice-stt README, WebFetch devnen docs. `nvidia-smi` baseline on server-01.
|
||||
2. On server-01: clone at a pinned commit into `/data/docker-compose/voice-tts/src`, write our compose (LAN bind
|
||||
192.168.1.90 only, GPU reservation, model cache volume, voices volume, `restart: unless-stopped`, healthcheck),
|
||||
copy the WAV (scp, then sha256 on both ends), build (`docker compose build`), `docker compose up -d`.
|
||||
3. Validate: `docker compose config -q`; health endpoint; synthesize 3 test sentences with the JARVIS voice via
|
||||
`/v1/audio/speech` and `/tts` stream:true; save outputs under `/data/voices/jarvis/tests/`; transcribe one output
|
||||
through voice-stt (:8300) as an intelligibility check.
|
||||
4. MEASURE (write every number): VRAM idle / model loaded / during synthesis; time-to-first-audio-chunk and total time
|
||||
for a short and a ~30-word sentence; RTF. Coexistence: (a) STT + TTS loaded, Ollama idle; (b) load llama3.1:8b +
|
||||
nomic via the Ollama API while TTS is loaded — does anything OOM, does Ollama offload to CPU (`ollama ps` /
|
||||
`/api/ps` shows size_vram vs size), timings; (c) `/api/unload` frees how much and how fast; reload time.
|
||||
5. Verdict: FITS (time-share works with Ollama's 5-min idle unload) / FITS-WITH-UNLOAD (voice client must call
|
||||
/api/unload after a session) / PARK (real contention). Leave the container running only if FITS or FITS-WITH-UNLOAD;
|
||||
if PARK, `docker compose down` voice-tts and say so.
|
||||
6. Write README (commands in FULL), the research report (numbers table + verdict + UNVERIFIED list), context handoff.
|
||||
|
||||
## Wrap-up (always)
|
||||
Append the dated handoff to `/opt/appdata/docker/docker-compose/voice-tts/.claude/context.md` (create it from the
|
||||
project template if missing). Run `python3 /opt/appdata/docker/.claude/scripts/embed_memory_dir.py --only-recent 3`.
|
||||
Confirm no leftover processes (`ps -eo pid,args | grep -E "[d]ocker build|[t]imeout"` on server-01). Never print
|
||||
secrets. Final message = JSON:
|
||||
```
|
||||
{"status":"succeeded|partially_succeeded|failed","project":"#214 V2 Chatterbox","actions_taken":[],"actions_failed":[],
|
||||
"files_touched":[],"containers_restarted":[],"repo_commit":"","image":"","endpoint":"","vram_mib":{"idle":0,"loaded":0,
|
||||
"synth_peak":0,"with_ollama_loaded":0},"latency_s":{"first_chunk_short":0,"total_30_words":0,"rtf":0},
|
||||
"coexistence":"","verdict":"FITS|FITS-WITH-UNLOAD|PARK","unverified":[],"next_step":"","notes":""}
|
||||
```
|
||||
@@ -0,0 +1,82 @@
|
||||
# V2b — Chatterbox-Turbo: build on PRIMARY, pull + run + measure on server-01 — #214
|
||||
|
||||
--max-turns 25 (reason: image build + push + pull retries + measurements; stop at 25 regardless).
|
||||
You are one agent of a batch; the main session reports all results together after the whole batch finishes — do not
|
||||
address the owner directly in your final message; write for the main session.
|
||||
If you issue the same tool call twice with identical arguments, STOP and output the wrap-up block with status=partially_succeeded.
|
||||
If a build/push/pull makes no progress for 10 minutes, stop it and report partially_succeeded.
|
||||
|
||||
## Why this re-run exists (verified 2026-09-29 20:20 by the main session)
|
||||
V2 failed: its image build ON server-01 died on a pip hash mismatch. Root cause is NOT the network: server-01 has a
|
||||
**bad RAM region** — the same 1.5 GB local file read through the page cache gave 20/20 wrong sha256 (9 different
|
||||
values) while direct reads were 5/5 identical. So on server-01 anything big that sits in RAM (pip downloads, image
|
||||
extraction, HF model downloads) can silently corrupt. **Rule for this run: server-01 downloads/builds NOTHING big
|
||||
unverified.** Build everything on primary (Ethernet, healthy RAM), move it only via `docker pull` from Gitea (Docker
|
||||
verifies every layer digest + diff-id and fails loudly → retry). An overnight kernel memtest fix follows later (not yours).
|
||||
|
||||
## Goal
|
||||
`voice-tts` running on server-01 from a primary-built, Gitea-hosted image with the Chatterbox-Turbo weights BAKED IN,
|
||||
the owner-approved JARVIS voice, and MEASURED VRAM/latency. Owner decision (binding): Option A time-share — Ollama keeps
|
||||
its default 5-min idle unload; Chatterbox is the LOWEST-priority GPU tenant. If it can't fit without changing other
|
||||
services → verdict PARK (never tune Ollama/STT).
|
||||
|
||||
## What V2 already left (reuse it — read it, don't redo it)
|
||||
- server-01 `/data/docker-compose/voice-tts/`: `src/` (devnen/Chatterbox-TTS-Server @ 915ae289340e10c6047f27f47e22eae9bf350c32),
|
||||
`src/Dockerfile.cu121` (base nvidia/cuda:12.1.1-runtime-ubuntu22.04, upstream requirements-nvidia.txt = torch 2.5.1+cu121,
|
||||
chatterbox-v2 pinned cc0357396d9c73fc1e6c544ee40bb596020edd09), `docker-compose.yml` (LAN bind 192.168.1.90:8004, GPU
|
||||
reservation, TTS_BF16=off, python healthcheck), `config.yaml` (upstream default, chatterbox-turbo, port 8004),
|
||||
`voices/jarvis/jarvis.wav` (sha256 e8b30592…64ab, mounted at /app/reference_audio), `measure.py` (written, never run).
|
||||
- Primary: `/opt/appdata/docker/docker-compose/voice-tts/{docker-compose.yml,.claude/context.md}` (read the V2 handoff).
|
||||
- Known: after `/api/unload` the model only comes back via POST `/restart_server`.
|
||||
- Driver 550 on server-01 = CUDA ≤ 12.4 → keep the 12.1.1 base.
|
||||
|
||||
## Steps
|
||||
1. PRIMARY build context at `/opt/appdata/docker/docker-compose/voice-tts/build/`: `git clone` devnen/Chatterbox-TTS-Server
|
||||
and `git checkout 915ae289340e10c6047f27f47e22eae9bf350c32`. Copy ONLY the small text files `Dockerfile.cu121`,
|
||||
`config.yaml`, `docker-compose.yml`, `measure.py` from server-01 with scp, then verify each: sha256 on primary ==
|
||||
`dd if=<file> iflag=direct status=none | sha256sum` on server-01 (direct read bypasses server-01's bad RAM).
|
||||
2. Bake the model: read `config.yaml` + the server's model-loading code to find the HF repo id(s) and cache path. Add a
|
||||
build step to the Dockerfile (on primary) that downloads those weights into the image at the path the server reads
|
||||
(e.g. `HF_HOME=/app/hf_cache` + `huggingface_hub.snapshot_download`). If a token is required, STOP → wrap-up with
|
||||
the repo id (no guessing, no tokens from anywhere). In the server-01 compose, do NOT mount a volume over the baked path.
|
||||
3. Build on primary: `docker build -f Dockerfile.cu121 -t gitea.local/backtalk6858/voice-tts:915ae28-cu121 .` (log to a
|
||||
file; primary is RAM-oversubscribed — build only, no other heavy work in parallel). Then a quick CPU-only sanity:
|
||||
`docker run --rm --entrypoint python3 <image> -c "import torch, chatterbox; print(torch.__version__)"`.
|
||||
4. Push (Gitea keeps ONE credential per registry — follow exactly, never print token values):
|
||||
`docker logout gitea.local`; login with the WRITE token:
|
||||
`grep -i '^Password=' /opt/appdata/docker/docker-compose/gitea/Docker-Registry-Token.txt | cut -d= -f2 | docker login gitea.local --username "$(grep -i '^Username=' /opt/appdata/docker/docker-compose/gitea/Docker-Registry-Token.txt | cut -d= -f2)" --password-stdin`;
|
||||
`docker push gitea.local/backtalk6858/voice-tts:915ae28-cu121` (record the pushed digest); `docker logout gitea.local`;
|
||||
re-login with the READ token (same command with `/opt/appdata/docker/docker-compose/gitea/Coolify-Pull.txt`).
|
||||
Each login/push/logout = its own command (no chained mutate+verify).
|
||||
5. server-01 pull (it is already logged in to gitea.local read-only): write a small retry script on server-01
|
||||
(`/data/docker-compose/voice-tts/pull.sh`: up to 5 attempts of `docker pull gitea.local/backtalk6858/voice-tts:915ae28-cu121`,
|
||||
30 s apart), run it with `timeout 3600`. Then confirm `docker image inspect --format '{{index .RepoDigests 0}}'` == the pushed digest.
|
||||
6. server-01 compose: replace the `build:` section with `image: gitea.local/backtalk6858/voice-tts:915ae28-cu121` +
|
||||
`pull_policy: never`; `docker compose config -q`; `docker compose up -d`. Copy the updated compose back to primary's repo copy.
|
||||
7. Validate + MEASURE with `measure.py` (fix it if needed): health; 3 JARVIS test sentences via `/v1/audio/speech` and
|
||||
`/tts` stream:true (outputs to `voices/jarvis/tests/`); one output transcribed through voice-stt (:8300) as an
|
||||
intelligibility check — garbled audio or nonsense transcript may be the bad-RAM problem, say so; VRAM idle/loaded/
|
||||
synth peak; first-chunk + total latency (short + ~30 words), RTF; coexistence with llama3.1:8b + nomic loaded via the
|
||||
Ollama API (OOM? CPU offload per `/api/ps` size_vram vs size?); `/api/unload` freed MiB + `/restart_server` reload time.
|
||||
8. Verdict FITS / FITS-WITH-UNLOAD / PARK. Leave running only for the first two; PARK → `docker compose down`.
|
||||
9. Write `/opt/appdata/docker/docker-compose/voice-tts/README.md` (commands in FULL) and
|
||||
`/opt/appdata/docker/research/voice_tts_v2_2026-09-29.md` (numbers table, verdict, UNVERIFIED list).
|
||||
|
||||
## Allowlists
|
||||
- Containers (rule 9, starts EMPTY): only `voice-tts` on server-01 (log the elevation). API requests (not restarts) to
|
||||
`ollama` and `voice-stt` are allowed. Never touch anything else.
|
||||
- Files: primary `/opt/appdata/docker/docker-compose/voice-tts/**`, `/opt/appdata/docker/research/voice_tts_v2_2026-09-29.md`;
|
||||
server-01 `/data/docker-compose/voice-tts/**`. Registry: only the one tag above. No sudo, no commits/pushes to git.
|
||||
|
||||
## Wrap-up (always)
|
||||
Append the dated handoff to `/opt/appdata/docker/docker-compose/voice-tts/.claude/context.md`. Run
|
||||
`python3 /opt/appdata/docker/.claude/scripts/embed_memory_dir.py --only-recent 3`. Confirm no leftover build/pull/timeout
|
||||
processes on either host (`ps -eo pid,args | grep -E "[d]ocker (build|pull|push)|[t]imeout"`) — never `pkill -f` a pattern
|
||||
that matches your own shell. Confirm primary's docker is logged in to gitea.local with the READ token again. Never print
|
||||
secrets. Final message = JSON:
|
||||
```
|
||||
{"status":"succeeded|partially_succeeded|failed","project":"#214 V2b Chatterbox","actions_taken":[],"actions_failed":[],
|
||||
"files_touched":[],"containers_restarted":[],"image":"","pushed_digest":"","pull_attempts":0,"endpoint":"",
|
||||
"vram_mib":{"idle":0,"loaded":0,"synth_peak":0,"with_ollama_loaded":0},"latency_s":{"first_chunk_short":0,"total_30_words":0,"rtf":0},
|
||||
"coexistence":"","audio_ok":"yes|no|garbled","verdict":"FITS|FITS-WITH-UNLOAD|PARK","unverified":[],"next_step":"","notes":""}
|
||||
```
|
||||
@@ -0,0 +1,64 @@
|
||||
# W3 — split-DNS + VPN-only HTTPS (build prep, files only) — #274
|
||||
|
||||
--max-turns 15. You are one agent of a batch; the main session reports all results together after the whole batch
|
||||
finishes — do not address the owner directly in your final message; write for the main session.
|
||||
If you issue the same tool call twice with identical arguments, STOP and output the wrap-up block with status=partially_succeeded.
|
||||
|
||||
## Goal
|
||||
Stage everything W3 needs so the owner can apply it in one sitting: tunnel clients resolve
|
||||
`grafana.reverseproxyserver.net` and `cadvisor.reverseproxyserver.net` to 10.42.42.1 and get a valid Let's Encrypt
|
||||
cert via Traefik DNS-01; off-VPN public paths unchanged. Design (LOCKED): `/opt/appdata/docker/research/diy_wireguard_design.md`
|
||||
— read §W3 row and all W3 details. First group = grafana + cadvisor only.
|
||||
|
||||
## Background (verified 2026-09-29)
|
||||
- wg-easy 15.4 on primary, container `wg-easy`, docker network `wg` = 10.42.42.0/24, gateway 10.42.42.1 (the host),
|
||||
wg-easy at 10.42.42.42, tunnel 10.8.0.0/24, UDP 45791. Client defaults: DNS `1.1.1.1`, AllowedIPs
|
||||
`10.8.0.0/24,10.42.42.0/24`. Clients: laptop 10.8.0.2, phone 10.8.0.3, tablet 10.8.0.4 (all proven off-LAN).
|
||||
wg-easy stores config in SQLite (`/etc/wireguard/wg-easy.db`, table `user_configs_table.default_dns`) — changing
|
||||
default DNS affects NEW clients; existing clients need per-client DNS set in the UI or a re-import. Put the exact UI
|
||||
steps in the runbook.
|
||||
- Today primary host services bind 192.168.1.88 (unreachable from the tunnel); only 0.0.0.0 binds answer at 10.42.42.1.
|
||||
- Traefik = container `coolify-proxy` (traefik:v3.6), compose at `/opt/appdata/docker/docker-compose/traefik/`
|
||||
(`docker-compose.yml`, `traefik.yml`, backup `traefik.yml.bak-20260917-231205`), network `coolify`. It is LIVE for
|
||||
~31 sites — no edits to live files. No traefik secretspec-resolver timer exists, but STAGE anyway in
|
||||
`docker-compose/traefik/deploy/w3/`. The wireguard project HAS `wireguard-secretspec-resolver.timer` (runs
|
||||
`compose up -d` every 15 min) — stage its changes in `docker-compose/wireguard/deploy/w3/`, never the live file.
|
||||
- Cloudflare DNS-01 token: Vault `secret/cloudflare/traefik-dns01` field `token` (AppRole files in
|
||||
`/opt/appdata/docker/docker-compose/vault/approle/role-id` + `secret-id`, Vault http://192.168.1.88:8200). You do
|
||||
not need to read it; wire it via the secretspec pattern used by other projects (see
|
||||
`docker-compose/wireguard/secretspec.toml` for the format; installed secretspec accepts ONLY `revision = "1.0"`).
|
||||
- Before writing Traefik config, WebFetch the official Traefik v3 docs for the ACME DNS challenge with the Cloudflare
|
||||
provider and for entrypoint address binding, and inline the exact option names you use in the README (cite URLs).
|
||||
- Grafana/cadvisor current routing: read their compose labels/dynamic config (`docker inspect grafana cadvisor` —
|
||||
pipe `.Config.Labels` only, never `.Config.Env`) to find how Traefik routes them today.
|
||||
|
||||
## Scope allowlist
|
||||
- CREATE `docker-compose/traefik/deploy/w3/` (staged docker-compose.yml, traefik.yml, dynamic `vpn-https.yml`, any
|
||||
secretspec.toml) and `docker-compose/wireguard/deploy/w3/` (dnsmasq service + config, staged compose/secretspec).
|
||||
- CREATE `docker-compose/wireguard/W3_RUNBOOK.md`. APPEND `docker-compose/wireguard/.claude/context.md`.
|
||||
- No live edits, no container starts/stops, no DNS record changes, no sudo.
|
||||
|
||||
## Steps
|
||||
1. Read the design doc (W3), both projects' live compose/config, grafana/cadvisor labels, the docs (WebFetch).
|
||||
2. Stage: dnsmasq on network `wg` (fixed IP in 10.42.42.0/24, answering `grafana|cadvisor.reverseproxyserver.net` →
|
||||
10.42.42.1, forwarding everything else to 1.1.1.1); Traefik `cfdns` resolver + `https-vpn` entrypoint bound
|
||||
10.42.42.1:443 (publish only on that IP) + `vpn-https.yml` routers for the two hosts with `tls.certresolver=cfdns`.
|
||||
3. VALIDATE with real tools: `docker compose -f <staged file> config -q` for each staged compose (run from a copy
|
||||
dir if relative paths need it); `secretspec check` equivalent if available, else confirm `revision = "1.0"`;
|
||||
`docker run --rm -v <staged traefik.yml>:/t.yml traefik:v3.6 traefik --configFile=/t.yml --help >/dev/null` is NOT a
|
||||
validator — instead parse YAML with python3 and cross-check option names against the fetched docs. Record results.
|
||||
4. W3_RUNBOOK.md (commands IN FULL, no `…`): order = token into secretspec path → dnsmasq up → Traefik backup + swap
|
||||
+ ONE scheduled proxy recreate (warn: all 31 sites blip) → set client DNS to the dnsmasq IP (UI steps per client)
|
||||
→ tests (on VPN `dig grafana.reverseproxyserver.net @<dnsmasq ip>` = 10.42.42.1; browser cert valid; off VPN
|
||||
public path unchanged; all sites still OK) → rollback (restore `traefik.yml.bak`, remove port + file, recreate).
|
||||
5. Wrap-up.
|
||||
|
||||
## Wrap-up (always)
|
||||
Append the dated handoff to `docker-compose/wireguard/.claude/context.md`. Run
|
||||
`python3 /opt/appdata/docker/.claude/scripts/embed_memory_dir.py --only-recent 3`. No commits/pushes/restarts, never
|
||||
print secrets or `.Config.Env`. Final message = JSON:
|
||||
```
|
||||
{"status":"succeeded|partially_succeeded|failed","project":"#274 W3 build prep","actions_taken":[],"actions_failed":[],
|
||||
"files_touched":[],"containers_restarted":[],"validation":{},"docs_fetched":[],"owner_decisions":[],
|
||||
"risks":[],"next_step":"","notes":""}
|
||||
```
|
||||
Reference in New Issue
Block a user