# Voice Chat — Background Agent Prompts (decomposed build) # Companion to voice_chat_build_plan.md. Preflight: 2026-07-29. # Launch target: on/before the build session ahead of 2026-08-14. # # ⚠️ RE-VERIFY AT LAUNCH (infra facts decay — do NOT trust these blindly): # - server-01 access = sudo-bridge-server01 at http://192.168.1.90:8082 (NOT secrets-proxy /shell). # Confirm the bridge is up and the endpoint/port before launching. # - Vault: resolve IP dynamically via docker inspect vault-iwaulpoi5hwirdlogshmul40; AppRole only. # - infrastructure-auth-map: secret/data/claude/infrastructure-auth-map. # - server-01 file writes: use `docker exec python3 -c` (tier-1) for unattended writes; # bare `python3 -c *` may be tier-2 (NTFY approval). Confirm allowlist state first. # - These 3 agents touch server-01 GPU/models independently → safe to run in PARALLEL, # EXCEPT A1 and A2 both measure VRAM; run A1 first, then A2, so VRAM numbers are additive/coherent. # - A3 runs on PRIMARY (writes a hook script), independent of A1/A2 → fully parallel. # # NONE of these agents register the global hook, install the laptop daemon, or make design # decisions. Those are main-session (see §MAIN-SESSION CHECKLIST at the bottom). --- ## AGENT A1 — Install + smoke-test faster-whisper on server-01 ``` You are a setup agent operating on server-01 through the sudo-bridge-server01 endpoint (http://192.168.1.90:8082 — VERIFY it is reachable before doing anything else). You will install and smoke-test faster-whisper for GPU speech-to-text. You will NOT touch primary, the laptop, Ollama's config, or any global hook. ## Task Get faster-whisper transcribing on server-01's GPU and report exact VRAM headroom while Ollama is running, so we know STT can coexist with Ollama + (later) XTTS. ## Credentials you will need - server-01 privileged ops: sudo-bridge-server01 at http://192.168.1.90:8082 — auth key at Vault secret/data/claude/infrastructure-auth-map (field: sudo-bridge-server01). RE-VERIFY field name. - AppRole bootstrap (if Vault read needed): VAULT_IP=$(docker inspect vault-iwaulpoi5hwirdlogshmul40 --format '{{json .NetworkSettings.Networks}}' | python3 -c "import sys,json; nets=json.load(sys.stdin); print(list(nets.values())[0]['IPAddress'])") ROLE_ID=$(cat /opt/appdata/docker/docker-compose/vault/approle/role-id | tr -d '[:space:]') SECRET_ID=$(cat /opt/appdata/docker/docker-compose/vault/approle/secret-id | tr -d '[:space:]') # login → token; ALWAYS revoke-self after use. ## Step-by-step (DISCOVERY FIRST — do not assume the environment) 1. Discover: GPU model, driver + CUDA version (`nvidia-smi`), total VRAM, and VRAM currently used by Ollama with a model loaded. Discover whether server-01 runs models bare-metal (python venv) or in Docker, and where Ollama lives. Record all of this. 2. Choose model size from measured free VRAM: prefer `large-v3` or `distil-large-v3` if it fits alongside Ollama with margin; else fall back to `medium`. State the choice + why. 3. Install faster-whisper in the environment matching step 1 (venv or a container). Prefer the least-invasive path; do NOT disturb Ollama. 4. Smoke test: transcribe a short built-in/test audio clip (generate one with `say`/espeak or a tiny recorded WAV) and confirm the returned text is correct. 5. Report the exact invocation (command or endpoint) that main-session code will call, plus measured VRAM free with Ollama + Whisper both resident. ## Output (in addition to the mandatory wrap-up) Report: GPU/CUDA, environment type (venv/docker + path), chosen Whisper model + rationale, the exact transcribe invocation, and VRAM: total / Ollama-resident / Whisper-resident / free. ## MANDATORY WRAP-UP (required regardless of success or failure) Before stopping for ANY reason — task complete, error, or approaching turn limit — output this JSON as your final message. Do not stop without it. { "status": "succeeded|partially_succeeded|failed", "actions_taken": ["action 1 — outcome", "action 2 — outcome"], "actions_failed": ["action — reason it failed"], "notes": "anything relevant for the next session" } If you hit --max-turns before finishing, set status="partially_succeeded" and list what remains in actions_failed. --max-turns 15 If you issue the same tool call or command twice with identical arguments, STOP immediately and output the mandatory wrap-up with status=partially_succeeded. ``` --- ## AGENT A2 — Install XTTS-v2 + clone the Jarvis voice on server-01 ``` You are a setup agent operating on server-01 through sudo-bridge-server01 (http://192.168.1.90:8082 — VERIFY reachable first). You will install Coqui XTTS-v2 and produce a test WAV in a cloned voice. You will NOT touch primary, the laptop, or any hook. RUN AFTER A1 so VRAM numbers are coherent. ## Task Stand up XTTS-v2 for GPU voice-cloning TTS, clone from the provided reference sample, and emit a test WAV. Measure real-time factor (RTF) and VRAM with Ollama (+ Whisper if resident). ## Context / pre-fetched data - Reference voice sample (Jarvis), 6–30s clean mono WAV, will be at: <> - Engine: Coqui XTTS-v2 (model `tts_models/multilingual/multi-dataset/xtts_v2`). NOTE its license is non-commercial (Coqui Public Model License) — fine for personal use. ## Credentials you will need - sudo-bridge-server01 auth: Vault secret/data/claude/infrastructure-auth-map (field: sudo-bridge-server01). RE-VERIFY. - AppRole bootstrap as in A1 if a Vault read is needed; revoke-self after. ## Step-by-step (DISCOVERY FIRST) 1. Confirm the environment discovered in A1 (venv vs docker, CUDA). Reuse it; do not rebuild. 2. Install Coqui XTTS-v2 (the maintained `coqui-tts` fork) in that environment. Accept the model license non-interactively (set the env var the library expects). 3. Confirm the reference WAV exists at the provided path and is mono, 16k+ Hz. If missing, STOP and report — do not fabricate a voice. 4. Synthesize one test sentence (e.g. "All systems are online, sir.") cloning the reference, write it to a test WAV, and report the path. 5. Measure RTF (synthesis time / audio duration) and VRAM: total / free with Ollama (+ Whisper if A1's model is still resident) / XTTS-resident. 6. Report the exact synthesis invocation main-session code will call, and whether XTTS + Whisper + Ollama all fit in VRAM simultaneously (the key coexistence question). ## Output (in addition to the mandatory wrap-up) Report: install path, test WAV path, RTF, VRAM breakdown, coexistence verdict (all three fit? y/n), and the exact synthesis invocation. ## MANDATORY WRAP-UP (required regardless of success or failure) Before stopping for ANY reason — task complete, error, or approaching turn limit — output this JSON as your final message. Do not stop without it. { "status": "succeeded|partially_succeeded|failed", "actions_taken": ["action 1 — outcome", "action 2 — outcome"], "actions_failed": ["action — reason it failed"], "notes": "anything relevant for the next session" } If you hit --max-turns before finishing, set status="partially_succeeded" and list what remains in actions_failed. --max-turns 15 If you issue the same tool call or command twice with identical arguments, STOP immediately and output the mandatory wrap-up with status=partially_succeeded. ``` --- ## AGENT A3 — Write + validate the fail-safe Stop-hook TTS script (primary) ``` You are a script-authoring agent on the primary server (192.168.1.88). You will WRITE and VALIDATE a fail-safe Stop-hook script that sends Claude's response text to server-01's XTTS endpoint. You will NOT register it in settings.json (that is a human main-session step) and you will NOT touch server-01 or the laptop. ## Task Create /opt/appdata/docker/.claude/hooks/voice_tts_stop.py implementing a Stop hook that: 1. Reads the hook JSON from stdin. 2. GATE: if env var CLAUDE_VOICE != "1", exit 0 immediately (silent no-op). This is what keeps every non-voice session and background-agent context silent. 3. Extracts Claude's final response text from the hook payload. 4. Fires the text to the server-01 XTTS endpoint as FIRE-AND-FORGET with a SHORT timeout (<= 1.5s connect). Do not wait for audio; the laptop daemon handles playback + queueing. 5. FAIL-SAFE: wrap everything so ANY exception, timeout, or bad payload → exit 0 silently. The hook must NEVER raise, never block, never non-zero. Blast radius = zero. 6. Tag the request with the session/source id so the laptop playback queue can label it. ## Context / pre-fetched data - The XTTS endpoint URL is not final yet (A2 produces it). Read it from an env var VOICE_TTS_URL (default to a placeholder) so main-session can wire the real value without editing the script. - Hook JSON shape: confirm the current Stop-hook stdin schema against an existing hook in /opt/appdata/docker/.claude/hooks/ before parsing (do NOT guess field names). ## Step-by-step 1. Read one existing hook in /opt/appdata/docker/.claude/hooks/ to confirm stdin JSON shape and the house style. 2. Write voice_tts_stop.py per the Task spec above. Use only stdlib (urllib, json, os). 3. Validate syntax: python3 -c "import ast; ast.parse(open('/opt/appdata/docker/.claude/hooks/voice_tts_stop.py').read())" 4. Test the gate: run the script with CLAUDE_VOICE unset piping a sample payload → must exit 0 with no network call. Then with CLAUDE_VOICE=1 and VOICE_TTS_URL pointed at a dead port → must STILL exit 0 (proving fail-safe), just no audio. 5. Do NOT edit settings.json. Report the exact settings.json Stop-hook registration snippet for the human to paste later. ## Output (in addition to the mandatory wrap-up) Report: script path, ast.parse result, both gate-test outcomes, and the settings.json snippet. ## MANDATORY WRAP-UP (required regardless of success or failure) Before stopping for ANY reason — task complete, error, or approaching turn limit — output this JSON as your final message. Do not stop without it. { "status": "succeeded|partially_succeeded|failed", "actions_taken": ["action 1 — outcome", "action 2 — outcome"], "actions_failed": ["action — reason it failed"], "notes": "anything relevant for the next session" } If you hit --max-turns before finishing, set status="partially_succeeded" and list what remains in actions_failed. --max-turns 15 If you issue the same tool call or command twice with identical arguments, STOP immediately and output the mandatory wrap-up with status=partially_succeeded. ``` --- ## MAIN-SESSION CHECKLIST (NOT agent-izable — you + Claude in the build session) These cannot be background agents: they need design decisions and/or live on the LAPTOP, which agents (running on primary) cannot reach. 1. **Laptop audio daemon** — one local service that owns: - Mic capture from the WorkTunes headset + PTT hotkey (hold-to-talk). - STT client: POST captured audio to server-01 Whisper (A1's invocation) → get text. - Injection: SSH to primary and `tmux send-keys` the text into the target Claude pane. - Playback queue: single FIFO, consecutive/non-overlapping, source-tagged (main + each agent). - Control-command parser: intercept `pause ` / `resume ` (and future verbs) as commands, act on the agent's tmux instance instead of sending to Claude. 2. **tmux orchestration** — main Claude pane + one pane/window per background agent (up to 5), named so the control parser can target them. Decide the pause mechanism (Ctrl-C interrupt vs. cooperative signal) here — do NOT let an agent guess it. 3. **Global hook registration** — paste A3's Stop-hook snippet into settings.json; set the voice env gate in the interactive launch wrapper (so voice = default for interactive, silent for background/agent contexts). End-to-end test the gate. 4. **Streaming TTS** — wire sentence-by-sentence synthesis so Jarvis starts talking on sentence 1. 5. **WPM listening test** — build + run once XTTS is up (A2): synthesize passages at 150/175/200/225/250 WPM in the Jarvis voice, comprehension check after each, record max accurate WPM → set as default playback tempo. 6. **Training-data logging** — voice-session TEXT logged like typed sessions; audio never persisted. ```