Files
claude-projects/claude-config/config/prompts/voice/V0_tts_engine_research.md
T
Backtalk6858 d9419109a5 docs(claude-config): 2026-09-27 preflight prompts, owner decisions, redactions
Saved prompts: W2, OLLAMA-1, BOOT-1, S1, AS0, AS1, JH-1, V0, V1, VS-1.
DECISIONS.md 2026-09-27 entry (sudo-bridge retired, Jenkins deploys via
agent-sudo deploy_service, Chatterbox-Turbo, vault-sandbox auto-unseal).
Voice A1/A2 superseded. Redacted two plaintext secrets in agent-builder
context (still in history; rotation tracked under #192).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 12:46:31 -05:00

4.5 KiB
Raw Blame History

V0 — re-confirm the TTS engine for Voice Chat (#214) before V2 — RESEARCH ONLY

Written 2026-09-27 (infrastructure general questions conversation). Owner approved re-checking the July choice (Coqui XTTS-v2) because Coqui the company shut down (early 2024; community fork coqui-tts by idiap) and the XTTS-v2 model licence (CPML) is non-commercial. Research only — no installs. Safe to run alongside anything.


You are a bounded RESEARCH agent. Budget: --max-turns 25. If you issue the same tool call twice with identical arguments, STOP and output the wrap-up block with status=partially_succeeded.

HARD RULES

  • Research only: WebSearch / WebFetch + reading local files. No installs, no pulls, no containers, no SSH, no sudo, no Vault, no git, no Postgres writes. Only writes: the report, the INDEX row, the embed run.
  • Every claim about a project (licence, last release date, VRAM, cloning ability, Docker image) must cite a URL you actually fetched in this run. Anything you could not verify = mark UNVERIFIED. Today is 2026-09-27; prefer sources from 2026.

REQUIREMENTS (locked in the July build plan /home/administrator/Desktop/claude/agent-builder/voice_chat_build_plan.md — read §2, §5, §7 — plus 2026-09-27 facts)

  • R1 Fully local, zero cloud, zero API cost. English.
  • R2 Voice cloning from a 6–30 s reference clip (Jarvis now; anime character voices later = reference swap). Engines that cannot clone are baselines only.
  • R3 Fits the GPU budget: server-01 RTX 2060 SUPER 8 GB (Turing, compute capability 7.5 — no bf16, no FlashAttention-2; check each engine's minimum), shared with Ollama llama3.1:8b (~5 GB when loaded, unloads after 5 min idle), nomic-embed-text (~0.3 GB) and a faster-whisper STT (~1–2 GB, V1). Target: TTS ≤ ~3 GB VRAM, or state what must be unloaded. Driver 550 (CUDA 12.4).
  • R4 Low latency: sentence-by-sentence streaming; first audio ≤ ~1.5 s after a sentence arrives; real-time factor < 1 on that GPU class.
  • R5 Maintained: a release or meaningful commit in the last ~6 months; a usable Docker image or a simple containerisable server with an HTTP API.
  • R6 Licence: personal use is fine; ALSO state whether commercial use is allowed (the owner's "virtual IT department" idea #29 may become a product — a non-commercial licence is a flag, not a disqualifier).

FIXED CANDIDATE LIST (evaluate exactly these; add at most 2 others only if they appear in ≥2 independent 2026 comparisons)

  1. XTTS-v2 via the idiap coqui-tts fork (the July choice)
  2. F5-TTS
  3. Chatterbox (Resemble AI)
  4. Fish Speech / OpenAudio
  5. CosyVoice 2 (or its current successor)
  6. IndexTTS 2
  7. Zonos (Zyphra)
  8. Kokoro — baseline (no cloning)
  9. Piper — baseline (rejected in July: fixed voices)

DELIVERABLE Report /opt/appdata/docker/research/voice_tts_engine_2026-09-28.md:

  1. Verdict (≤5 lines): the recommended engine + runner-up, and whether it changes the July decision.
  2. Scoring table: rows = candidates; columns = R1–R6 (pass/fail/partial + one-line evidence + source URL), VRAM (measured-by-someone vs claimed), cloning clip length needed, streaming support, Docker image (exact name) + API style, last release date, licence (code / weights separately).
  3. For the top 2: exact install path we would use on server-01 (image or build), config for Turing/fp16, the HTTP call the Stop hook would make, and the VRAM-sharing plan with Ollama + STT (options, not a decision).
  4. Risks + UNVERIFIED list. Then add one row to /home/administrator/Desktop/claude/claude-config/research/INDEX.md (read it first; follow its format).

PERSIST BEFORE YOU FINISH

  • Append a dated "## V0 TTS engine research — 2026-09-28 (background agent)" block (verdict + report path) to /home/administrator/Desktop/claude/agent-builder/.claude/context.md — append only, one cat >> heredoc; do NOT read or print other parts of that file beyond what you need (it contains a plaintext secret scheduled for redaction).
  • Run: python3 /opt/appdata/docker/.claude/scripts/embed_memory_dir.py --only-recent 5
  • FINAL message = wrap-up JSON only, always: {"status":"succeeded|partially_succeeded|failed","project":"voice chat #214","phase":"V0 TTS research","actions_taken":[],"actions_failed":[],"files_touched":[],"containers_restarted":[],"recommended":"","runner_up":"","changes_july_decision":true,"top2":[{"engine":"","vram_gb":0,"clone_clip_s":"","streaming":true,"image":"","licence_code":"","licence_weights":"","commercial_ok":false,"turing_ok":true}],"unverified_claims":[],"next_step":"main session writes V2 prompt for the chosen engine","notes":""}