Files
Backtalk6858 1a44c84fa7 docs(voice-chat): lock preflight decisions + decomposed build prompts
Voice Chat for Claude Code preflight/design session (personal_projects id=214).
Fully-local PTT voice: faster-whisper + Coqui XTTS-v2 on server-01 GPU, Jarvis
voice. Phase A includes background-agent audio-arbitration queue + pause/resume
voice control + per-agent tmux. Hard deadline ~2026-08-14 (Max upgrade).

- voice_chat_build_plan.md: locked decisions, topology, blast-radius/rollback
- voice_chat_agent_prompts.md: 3 bounded server-01 agent prompts (A1 Whisper,
  A2 XTTS+Jarvis, A3 fail-safe Stop-hook) + main-session checklist
- context.md: refresh CURRENT STATE block (was stale since 2026-06-04)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-29 16:58:30 -05:00

127 lines
7.6 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Voice Chat for Claude Code — Build Plan & Locked Decisions
# Preflight session: 2026-07-29 (claude-config conversation)
# HARD DEADLINE: functional by 2026-08-14 (Max subscription upgrade → up to 5 background agents run regularly)
# Status: DECISIONS LOCKED. Build not yet started.
---
## 0. One-paragraph summary
A fully-local, push-to-talk voice layer over Claude Code. You speak into a Bluetooth
headset on your laptop; a laptop-side daemon ships audio to server-01's GPU for
speech-to-text (Whisper), injects the transcript into Claude Code (running in tmux on
primary) via `tmux send-keys`; when Claude finishes, a Stop hook ships the response text
to server-01 for text-to-speech (XTTS-v2, Jarvis voice), and the audio is played back on
the headset. Every interactive session is a voice session. Up to 5 background agents each
run in their own tmux instance and speak through a single, non-overlapping playback queue;
you can say "pause <agent name>" to interrupt one, talk to it, then resume.
---
## 1. Topology (three machines)
| Machine | Role | Owns |
|---|---|---|
| **Laptop** (you) | Client you sit at; SSH into primary | WorkTunes Connect BT headset (mic + playback), the **audio daemon**, PTT hotkey, the **playback queue** |
| **Primary** (192.168.1.88) | Where Claude Code runs (you SSH in) | Claude Code in **tmux**, the **Stop hook**, `tmux send-keys` injection |
| **Server-01** (GPU) | Inference | **faster-whisper** (STT), **Coqui XTTS-v2** (TTS), Ollama (already there) |
Audio makes two network hops: laptop→server-01 (STT), server-01→laptop (playback).
---
## 2. LOCKED DECISIONS
### Interaction model — phased
- **Phase A (build now, due Aug 14):** dictation + read-back wrapper, single session, PTT, Jarvis voice, PLUS the full background-agent control layer (see §3). *The control layer is Phase A, not later — confirmed 2026-07-29.*
- **Phase B (later):** hands-free conversational loop (wake word / continuous listen), streaming-TTS polish.
- **Phase C (later):** custom character voices (e.g. Setsuna / Redo of Healer) + heavy speech-pattern / dictation styling layer.
### Engines — fully local, zero cloud, zero API cost
- **STT:** `faster-whisper` on server-01 GPU. (Aligns with `feedback_no_claude_api`, $20/mo cap, ambient-audio privacy.)
- **TTS:** **Coqui XTTS-v2** on server-01 GPU — voice-cloning capable, so Jarvis now and Setsuna later are just reference-sample swaps. (Piper rejected: fixed voices, can't clone.)
### Voice
- **Now:** Jarvis (Iron Man). Needs a 630s clean reference audio sample — **USER TO PROVIDE**.
- **Later (Phase C):** anime character voices + speech-pattern styling as a text-preprocessing layer before TTS.
### Activation
- **Push-to-talk** (hold key on laptop). Chosen for zero false triggers and because it mitigates the Bluetooth HFP/A2DP problem (mic only opens while held → playback stays high quality otherwise).
### Default behavior
- **Every interactive session is a voice session** (voice is the default / opt-OUT, not opt-in).
- Background/agent/non-voice contexts are gated separately (see §3, §4).
### Training data (`feedback_always_include_training_loop`)
- **Log the TEXT** of voice sessions (transcribed prompt + response) exactly as typed sessions are logged.
- **Never persist audio** — audio buffers deleted the moment STT/TTS completes. No conversation recordings.
---
## 3. Background-agent control layer (PHASE A)
This is the part that grew the scope and is required for the Aug-14 5-agent workflow.
1. **All agent outputs are spoken** — you want to hear every agent, not just the main session.
2. **Single arbitrated playback queue** on the laptop daemon: audio plays **consecutively, never overlapping**; each chunk is **source-tagged** (which agent / main session it came from) so you know who's talking.
3. **Voice control commands** are intercepted by the daemon as commands (NOT passed to Claude):
- `pause <agent name>` → interrupt that agent's tmux instance so you can speak to/about it.
- `resume <agent name>` → resume it.
- (grammar extensible: `stop`, `skip`, `mute`, etc.)
4. **Per-agent tmux instances:** each background agent runs as its own Claude Code process in its own tmux window/pane. This is what makes independent pause/resume and clean per-source audio possible.
Open design question for build session: how "pause" is delivered to an agent's tmux pane
(Ctrl-C interrupt vs. a cooperative pause signal) without corrupting its run — resolve in
main session, do NOT let an agent guess.
---
## 4. Safety, blast radius, rollback
- **Blast radius concern:** hooks are GLOBAL (`/home/administrator/.claude/settings.json`) — the Stop hook fires for EVERY session and background agent. A naïve synchronous hook that waits on server-01 could freeze every conversation.
- **Containment:** the TTS Stop hook must be (a) gated so it only speaks in voice contexts, (b) **async / fire-and-forget with a short timeout**, (c) **fail-safe: any error → silent `exit 0`, never blocks Claude.**
- **Rollback:** the layer is purely additive. Disable = unset the voice env gate / remove the hook registration. tmux and Claude Code are unaffected. No Vault/N8N/Coolify schema changes.
- **Blast radius after containment:** voice sessions only; worst case = no audio, never a blocked Claude.
---
## 5. Hardware / latency notes
- **WorkTunes Connect (Bluetooth):** HFP/A2DP mutual exclusion — mic + high-quality playback can't coexist; PTT mitigates. Internal mic fine for commands, NOT good enough for recording a reference voice sample.
- **Headset upgrade (DEFERRED):** wired USB headset sidesteps HFP/A2DP entirely (clean mic + full-quality playback simultaneously). Recommend when ready; not a build blocker.
- **Latency budget (LAN + GPU):** ~1s before Claude thinks (BT capture + STT) + ~13s after (XTTS), dominated by Claude's own generation. Mitigate with **sentence-by-sentence streaming TTS** (Jarvis starts talking on sentence 1).
---
## 6. Deliverables of the preflight session
- [x] Locked grill-me decisions (this document).
- [ ] **WPM listening test** — script that synthesizes passages at 150/175/200/225/250 WPM in the Jarvis voice, asks a comprehension check after each, and reports your max accurate listening WPM → becomes the default TTS playback tempo. Doubles as first end-to-end server-01→laptop audio check. (Build + run once XTTS stands up.)
- [ ] **Build-plan agent prompts** (see §8) — bounded server-01 prompts + main-session checklist.
## 7. Open items / user tasks
- **Jarvis reference audio** (630s clean) — user provides.
- **server-01 GPU/VRAM audit** — can XTTS-v2 + Whisper + Ollama coexist in VRAM? (First build step / first agent task.)
- **Headset upgrade** — deferred decision.
---
## 8. Build execution shape (NOT one agent)
Per `playbook_background_agent_prompts.md` (agents = bounded ≤15 tool calls, defined
completion, no design) and the topology (agents run on primary, CANNOT touch the laptop),
the build splits into:
- **Bounded background-agent prompts (server-01 only):**
- A1 — install + smoke-test `faster-whisper` on server-01 GPU; report VRAM headroom.
- A2 — install + smoke-test XTTS-v2; clone from a provided Jarvis sample; emit a test WAV.
- A3 — write + validate the fail-safe async Stop-hook TTS script (does NOT register it globally without confirmation).
- **Main-session work (you + Claude, cannot be agent-ized):**
- Laptop audio daemon (mic, playback queue, PTT, control-command parser).
- tmux injection wiring + per-agent tmux orchestration.
- Global hook registration + gating + end-to-end integration test.
Agent prompts to be written into `voice_chat_agent_prompts.md` (alongside `agent_prompts.md`).