Files
claude-projects/claude-config/config/prompts/voice/V1_stt_server01_gpu.md
T
Backtalk6858 d9419109a5 docs(claude-config): 2026-09-27 preflight prompts, owner decisions, redactions
Saved prompts: W2, OLLAMA-1, BOOT-1, S1, AS0, AS1, JH-1, V0, V1, VS-1.
DECISIONS.md 2026-09-27 entry (sudo-bridge retired, Jenkins deploys via
agent-sudo deploy_service, Chatterbox-Turbo, vault-sandbox auto-unseal).
Voice A1/A2 superseded. Redacted two plaintext secrets in agent-builder
context (still in history; rotation tracked under #192).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 12:46:31 -05:00

8.7 KiB
Raw Blame History

V1 — speech-to-text server on server-01 GPU (personal_projects #214, Voice Chat for Claude Code)

Written 2026-09-27 (infrastructure general questions conversation). SUPERSEDES agent A1 in /home/administrator/Desktop/claude/agent-builder/voice_chat_agent_prompts.md — A1/A2 were written to go through "sudo-bridge-server01 :8082", which is dead AND is being retired (agent-sudo replaces sudo-bridge). None of this needs sudo: server-01 Docker already has the nvidia runtime (nvidia-container-toolkit 1.20.1, driver 550.163.01) and the administrator user is in the docker group, so the STT server is just a GPU container. OWNER GATE: spawn only after the owner has OK'd pulling a new image onto server-01 (shared host). Run it ALONE among server-01-mutating agents (W2/OLLAMA-1/S1/AS0/JH-1/BOOT-1 are fine alongside — they don't mutate server-01). Plan of record: /home/administrator/Desktop/claude/agent-builder/voice_chat_build_plan.md (read §1–§5, §7). Role boundary note (owner-locked 2026-06-25): new services deploy via a Jenkins pipeline and are monitored by Hermes. Jenkins is not usable yet (#181, JH-1 audits it), so V1 deploys by hand as a prototype on the test host; the compose is written so it can be dropped into a Gitea repo + Jenkinsfile.server01 later (list what that pipeline needs in the README), and the README names the health endpoint Hermes should watch.


You are a bounded background BUILD agent for personal_projects #214, phase V1 (STT only). Budget: --max-turns 35. If you issue the same tool call twice with identical arguments, STOP and output the wrap-up block with status=partially_succeeded.

HARD RULES

  • All work on server-01 via ssh -o BatchMode=yes administrator@192.168.1.90 '<cmd>'. No sudo anywhere. No changes on primary except the context.md append + embed run. No git add/commit/push, no Postgres writes, no Vault.
  • Container allowlist (rule 9): starts EMPTY. You may add exactly ONE container to it — the new voice-stt — and log that elevation in actions_taken. You may create/start/stop/recreate ONLY voice-stt. NEVER touch ollama, hermes, jenkins, agent-sudo, n8n-*, -sandbox, media-pipeline-sandbox-, incus. You MAY send read-only API requests to server-01's Ollama (/api/tags, /api/ps, and one /api/generate with "keep_alive":"5m" to load llama3.1:8b for the VRAM measurement — that is a normal client request, not a container action).
  • Write any wait loop as a small script FILE on server-01 run with timeout (no nested-quoted until one-liners); confirm no leftover waiters before the wrap-up (ps -eo pid,args | grep "[t]imeout").
  • Audio privacy (plan §2): never persist audio beyond the test clip; delete it at the end.
  • If a hook blocks something, stop with a partial wrap-up.

FACTS (verified 2026-09-27)

  • server-01 GPU: NVIDIA GeForce RTX 2060 SUPER, 8192 MiB, driver 550.163.01 (→ CUDA 12.4 driver; CUDA 12.x minor-version compatibility applies — if the image's CUDA runtime is newer and fails to init, record the exact error; that decides the fallback). Idle VRAM used ≈ 1 MiB. Ollama container ollama holds llama3.1:8b (~4.9 GB file) + nomic-embed-text (~0.3 GB), loads on demand, default keep-alive 5 min.
  • Fast disk: /data (1TB SSD, 854 GB free). Existing compose dirs there: /data/docker-compose/ollama/. Models dir /data/models/.
  • Image candidates verified to exist 2026-09-27: PRIMARY ghcr.io/speaches-ai/speaches:latest-cuda (faster-whisper, OpenAI-compatible /v1/audio/transcriptions; latest release v0.9.0-rc.3, 2025-12-27) — FALLBACK onerahmet/openai-whisper-asr-webservice:latest-gpu (with its faster_whisper engine). Do not pick anything else.
  • Primary already has an unrelated batch transcriber (/opt/appdata/docker/docker-compose/whisper/, meeting recordings) — do NOT touch or reuse it.

LOCKED DECISIONS

  1. Container name voice-stt, compose at /data/docker-compose/voice-stt/docker-compose.yml on server-01, runtime/gpus per the image docs, restart unless-stopped, published 192.168.1.90:8300:<container port> (LAN IP bind, NOT 0.0.0.0), model cache bind /data/models/voice-stt. Pin the image to an exact tag + digest you resolve (record both). Follow the compose template /opt/appdata/docker/non-docker-python-scripts/Docker Template/docker-compose.yml (read it on primary first).
  2. Read the chosen image's official docs ONCE (WebFetch speaches.ai docs: installation + configuration + model management) and inline the real env var names in the compose — do not guess (e.g. how to set device=cuda, compute type, model TTL/unload, and how models get downloaded in this version).
  3. Model shortlist — test exactly these two, compute type int8_float16: (a) large-v3-turbo faster-whisper conversion, (b) distil-large-v3 faster-whisper conversion (use the IDs the image's registry/docs name). Choose the default by: correct transcript of the test clip, then lowest VRAM, then speed. English only.
  4. Test clip = whisper.cpp's public samples/jfk.wav (download to /tmp/voice-stt-test/ on server-01; known text: "And so my fellow Americans, ask not what your country can do for you, ask what you can do for your country."). Delete the directory at the end.

STEPS

  1. Pre-flight snapshot: docker ps --format '{{.Names}}' (record the set — restore target), nvidia-smi --query-gpu=memory.used,memory.total --format=csv, ss -ltn | grep 8300 (must be free).
  2. Write compose; docker compose config -q; pull; up. Health: poll the image's health/models endpoint via a timeout'd script.
  3. For each model: transcribe the clip 3× (first = cold load), record transcript, wall time, real-time factor, and VRAM (nvidia-smi --query-compute-apps=pid,process_name,used_memory --format=csv + total).
  4. Coexistence test (plan §7): load llama3.1:8b via one Ollama /api/generate (tiny prompt, keep_alive 5m) and nomic via one /api/embeddings, keep the chosen whisper model loaded, then transcribe again. Record total VRAM, whether anything OOM'd or fell back to CPU (check docker logs ollama --since 5m | grep -i -E "cuda|offload|memory" — read-only), and latency change. Leave XTTS out (V2 — not built; budget it as "XTTS-v2 ≈ 2–3 GB, UNVERIFIED").
  5. Write a VRAM budget section (if /opt/appdata/docker/research/jenkins_hermes_state_2026-09-28.md exists, read which LLM backend Hermes uses — if it is server-01 Ollama, Hermes is a fourth GPU tenant; include it): measured numbers for Ollama-8B + nomic + STT, the remaining headroom for TTS, and 2–3 OPTIONS the main session will choose from (e.g. shorter Ollama keep-alive / OLLAMA_MAX_LOADED_MODELS, STT on CPU fallback, smaller LLM) — options only, do not apply any.
  6. Leave voice-stt RUNNING with the chosen default model configured; verify from PRIMARY: curl -s -F file=@<a local copy of jfk.wav> -F model=<id> http://192.168.1.90:8300/v1/audio/transcriptions (or the fallback image's endpoint) returns the transcript. Delete the local copy after.
  7. Restore check: the container set equals the step-1 snapshot + voice-stt.

DELIVERABLES

  1. /data/docker-compose/voice-stt/docker-compose.yml + README.md on server-01 (what it is, endpoint, model, how to change model, rollback = docker compose down, the exact client call).
  2. A copy of both files + the report into the primary repo: /opt/appdata/docker/docker-compose/voice-stt/ (compose + README) and /opt/appdata/docker/research/voice_stt_v1_2026-09-28.md (results table, VRAM budget, options, UNVERIFIED list, the exact invocation the laptop daemon will call).

SCOPE ALLOWLIST: server-01 /data/docker-compose/voice-stt/, /data/models/voice-stt/, /tmp/voice-stt-test/ (deleted), container voice-stt; primary /opt/appdata/docker/docker-compose/voice-stt/, the report, /opt/appdata/docker/docker-compose/voice-stt/.claude/context.md. Nothing else.

PERSIST BEFORE YOU FINISH

  • Create /opt/appdata/docker/docker-compose/voice-stt/.claude/context.md (template /home/administrator/.claude/projects/-opt-appdata-docker/memory/playbook_project_context_template.md), append "## V1 STT — 2026-09-28 (background agent)".
  • Run: python3 /opt/appdata/docker/.claude/scripts/embed_memory_dir.py --only-recent 5
  • FINAL message = wrap-up JSON only, always: {"status":"succeeded|partially_succeeded|failed","project":"voice chat #214","phase":"V1 STT","actions_taken":[],"actions_failed":[],"files_touched":[],"containers_restarted":[{"name":"voice-stt","action":"","health":""}],"image":{"ref":"","tag":"","digest":"","fallback_used":false},"models":[{"id":"","transcript_ok":true,"rtf":0.0,"vram_mib":0}],"chosen_model":"","coexistence":{"total_vram_mib":0,"oom":false,"ollama_cpu_fallback":false},"endpoint":"http://192.168.1.90:8300/...","unverified_claims":[],"next_step":"","notes":""}