Files
claude-projects/claude-config/config/prompts/voice/V3_laptop_client_build.md
T
2026-10-02 00:19:54 -05:00

5.2 KiB

V3 — Voice Chat laptop client (push-to-talk → Whisper → Claude Code; replies → Chatterbox) (spawned 2026-10-01, main session 771b2744)

Bounded background agent, personal_projects #214. Budget: --max-turns 60. Same tool call twice with identical arguments → STOP with a partial wrap-up.

READ FIRST (design is decided — do not re-design): /opt/appdata/docker/agent-builder/voice_chat_build_plan.md, then /opt/appdata/docker/research/voice_stt_v1_2026-09-28.md and voice_tts_engine_2026-09-28.md, and the #214 notes: PG=$(docker ps --format '{{.Names}}' | grep '^postgres-'); docker exec "$PG" psql -U postgres -d personal_projects -At -c "SELECT notes FROM projects WHERE id=214" If the build plan's laptop section conflicts with the live servers below, the live servers win; report the conflict.

LIVE SERVERS (verified 2026-10-01): STT voice-stt on server-01 http://192.168.1.90:8300 (read its API from the research doc / curl -s http://192.168.1.90:8300/docs or openapi.json); TTS voice-tts (Chatterbox-Turbo) on server-01 http://192.168.1.90:8004 — streaming POST /tts (stream:true), OpenAI POST /v1/audio/speech, reference voice jarvis.wav (see /data/docker-compose/voice-tts/measure.py on server-01 for working request bodies). Measured: first audio 0.65-0.76 s streaming; first call after start 10 s warm-up. Both are LAN-only, no auth. LAPTOP: ssh administrator@192.168.1.89 (key auth works, BatchMode). Hostname omarchy, Arch-based (Omarchy, Hyprland), Python 3.14, PipeWire (pw-record/pw-play), tmux, Claude Code installed (/.claude exists).

BUILD (on the laptop, all user-level, no sudo):

  1. ~/voice-chat/ project: a Python venv (pip into the venv is fine), the push-to-talk client:
    • toggle mode via a small CLI voice-chat ptt (start/stop recording on each invocation, state in $XDG_RUNTIME_DIR) — the keybind will call it; record 16 kHz mono WAV with pw-record from the default source;
    • on stop: POST to voice-stt, get text; type it into the tmux pane running Claude Code: target configurable (default: the pane whose current command is claude; find with tmux list-panes -a -F ...), using tmux send-keys -l for the text, then Enter only if config auto_submit=true (default false → owner reviews);
    • a desktop notification (notify-send if present) for recording on/off and errors.
  2. TTS side: a Claude Code Stop hook script (~/voice-chat/hooks/speak_last_reply.py) that reads the hook JSON from stdin, extracts the last assistant message text (check the current Claude Code hook input schema in the docs — https://docs.claude.com/en/docs/claude-code/hooks — do not guess field names), strips code blocks/tables to a speakable summary (first ~2 sentences or ≤ 400 chars), streams it from Chatterbox and plays with pw-play. Non-blocking (spawn + return immediately), never fails the hook (exit 0), mute file switch ($XDG_RUNTIME_DIR/voice-chat.mute), and skips when the reply is empty.
  3. A user systemd unit ONLY if the design needs a daemon; otherwise none. Prefer no daemon.
  4. STAGE, do NOT activate: (a) the Hyprland keybind line (find the user's Hyprland config include file used by Omarchy for user bindings; write the proposed line to ~/voice-chat/STAGED_hyprland_bind.conf and the exact include/edit instruction), (b) the Claude Code hook JSON snippet to ~/voice-chat/STAGED_claude_settings_hook.json with the exact place it goes in ~/.claude/settings.json. Do NOT edit ~/.claude/settings.json or Hyprland config. SELF-TEST (no human voice needed):
  • STT round trip: generate a WAV of a known sentence with Chatterbox (server-01), resample to 16 kHz mono, run the client's stop-path on that file → assert transcript ≈ sentence (word error rate ≤ 10 %).
  • tmux injection: start a throwaway tmux session running cat as a stand-in (NOT claude), point the client at it, assert the text arrives; kill the session.
  • Hook: feed a sample hook JSON (built from the documented schema) to the hook script with the mute file set (assert it exits 0 fast and logs what it would say) and once unmuted with output to a file instead of pw-play (assert a valid WAV of plausible length). Do not play audio aloud.
  • Latency: measure STT time for a 5 s clip and TTS time-to-first-chunk from the laptop. HARD RULES: no sudo (if a system package is truly required → ⏸ OWNER-COMMAND-REQUEST block: host, command, why, paste_back, resume_at); no edits outside ~/voice-chat on the laptop; never touch the running Claude Code session or ~/.claude/settings.json; no audio playback; no server-side changes (server-01 containers untouched); no git; mask secrets (there should be none). DELIVERABLES: ~/voice-chat/ with README.md (install/activate steps for the owner: keybind + hook, how to mute, how to change the tmux target, troubleshooting), tests, staged files. Copy the README to /opt/appdata/docker/docker-compose/voice-tts/LAPTOP_CLIENT_README.md on primary (scp from the laptop). FINAL message = wrap-up JSON only: status, project "voice-chat", files_built[], staged[] (path, what, where it goes), self_tests[] (name, result, numbers), latency {stt_5s_s, tts_first_chunk_s}, owner_steps[] (exact commands/edits to activate), design_conflicts[], actions_taken[], actions_failed[], unverified_claims[], next_step, notes.