Saved prompts: W2, OLLAMA-1, BOOT-1, S1, AS0, AS1, JH-1, V0, V1, VS-1. DECISIONS.md 2026-09-27 entry (sudo-bridge retired, Jenkins deploys via agent-sudo deploy_service, Chatterbox-Turbo, vault-sandbox auto-unseal). Voice A1/A2 superseded. Redacted two plaintext secrets in agent-builder context (still in history; rotation tracked under #192). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
8.7 KiB
V1 — speech-to-text server on server-01 GPU (personal_projects #214, Voice Chat for Claude Code)
Written 2026-09-27 (infrastructure general questions conversation). SUPERSEDES agent A1 in
/home/administrator/Desktop/claude/agent-builder/voice_chat_agent_prompts.md — A1/A2 were written to go through
"sudo-bridge-server01 :8082", which is dead AND is being retired (agent-sudo replaces sudo-bridge). None of this needs
sudo: server-01 Docker already has the nvidia runtime (nvidia-container-toolkit 1.20.1, driver 550.163.01) and the
administrator user is in the docker group, so the STT server is just a GPU container.
OWNER GATE: spawn only after the owner has OK'd pulling a new image onto server-01 (shared host). Run it ALONE among
server-01-mutating agents (W2/OLLAMA-1/S1/AS0/JH-1/BOOT-1 are fine alongside — they don't mutate server-01).
Plan of record: /home/administrator/Desktop/claude/agent-builder/voice_chat_build_plan.md (read §1–§5, §7).
Role boundary note (owner-locked 2026-06-25): new services deploy via a Jenkins pipeline and are monitored by Hermes. Jenkins is not usable yet (#181, JH-1 audits it), so V1 deploys by hand as a prototype on the test host; the compose is written so it can be dropped into a Gitea repo + Jenkinsfile.server01 later (list what that pipeline needs in the README), and the README names the health endpoint Hermes should watch.
You are a bounded background BUILD agent for personal_projects #214, phase V1 (STT only). Budget: --max-turns 35. If you issue the same tool call twice with identical arguments, STOP and output the wrap-up block with status=partially_succeeded.
HARD RULES
- All work on server-01 via
ssh -o BatchMode=yes administrator@192.168.1.90 '<cmd>'. No sudo anywhere. No changes on primary except the context.md append + embed run. No git add/commit/push, no Postgres writes, no Vault. - Container allowlist (rule 9): starts EMPTY. You may add exactly ONE container to it — the new
voice-stt— and log that elevation in actions_taken. You may create/start/stop/recreate ONLYvoice-stt. NEVER touch ollama, hermes, jenkins, agent-sudo, n8n-*, -sandbox, media-pipeline-sandbox-, incus. You MAY send read-only API requests to server-01's Ollama (/api/tags,/api/ps, and one/api/generatewith"keep_alive":"5m"to load llama3.1:8b for the VRAM measurement — that is a normal client request, not a container action). - Write any wait loop as a small script FILE on server-01 run with
timeout(no nested-quoteduntilone-liners); confirm no leftover waiters before the wrap-up (ps -eo pid,args | grep "[t]imeout"). - Audio privacy (plan §2): never persist audio beyond the test clip; delete it at the end.
- If a hook blocks something, stop with a partial wrap-up.
FACTS (verified 2026-09-27)
- server-01 GPU: NVIDIA GeForce RTX 2060 SUPER, 8192 MiB, driver 550.163.01 (→ CUDA 12.4 driver; CUDA 12.x minor-version compatibility applies — if the image's CUDA runtime is newer and fails to init, record the exact error; that decides the fallback). Idle VRAM used ≈ 1 MiB. Ollama container
ollamaholdsllama3.1:8b(~4.9 GB file) +nomic-embed-text(~0.3 GB), loads on demand, default keep-alive 5 min. - Fast disk:
/data(1TB SSD, 854 GB free). Existing compose dirs there:/data/docker-compose/ollama/. Models dir/data/models/. - Image candidates verified to exist 2026-09-27: PRIMARY
ghcr.io/speaches-ai/speaches:latest-cuda(faster-whisper, OpenAI-compatible/v1/audio/transcriptions; latest release v0.9.0-rc.3, 2025-12-27) — FALLBACKonerahmet/openai-whisper-asr-webservice:latest-gpu(with its faster_whisper engine). Do not pick anything else. - Primary already has an unrelated batch transcriber (
/opt/appdata/docker/docker-compose/whisper/, meeting recordings) — do NOT touch or reuse it.
LOCKED DECISIONS
- Container name
voice-stt, compose at/data/docker-compose/voice-stt/docker-compose.ymlon server-01, runtime/gpusper the image docs, restartunless-stopped, published192.168.1.90:8300:<container port>(LAN IP bind, NOT 0.0.0.0), model cache bind/data/models/voice-stt. Pin the image to an exact tag + digest you resolve (record both). Follow the compose template/opt/appdata/docker/non-docker-python-scripts/Docker Template/docker-compose.yml(read it on primary first). - Read the chosen image's official docs ONCE (WebFetch speaches.ai docs: installation + configuration + model management) and inline the real env var names in the compose — do not guess (e.g. how to set device=cuda, compute type, model TTL/unload, and how models get downloaded in this version).
- Model shortlist — test exactly these two, compute type
int8_float16: (a)large-v3-turbofaster-whisper conversion, (b)distil-large-v3faster-whisper conversion (use the IDs the image's registry/docs name). Choose the default by: correct transcript of the test clip, then lowest VRAM, then speed. English only. - Test clip = whisper.cpp's public
samples/jfk.wav(download to/tmp/voice-stt-test/on server-01; known text: "And so my fellow Americans, ask not what your country can do for you, ask what you can do for your country."). Delete the directory at the end.
STEPS
- Pre-flight snapshot:
docker ps --format '{{.Names}}'(record the set — restore target),nvidia-smi --query-gpu=memory.used,memory.total --format=csv,ss -ltn | grep 8300(must be free). - Write compose;
docker compose config -q; pull; up. Health: poll the image's health/models endpoint via a timeout'd script. - For each model: transcribe the clip 3× (first = cold load), record transcript, wall time, real-time factor, and VRAM (
nvidia-smi --query-compute-apps=pid,process_name,used_memory --format=csv+ total). - Coexistence test (plan §7): load llama3.1:8b via one Ollama
/api/generate(tiny prompt, keep_alive 5m) and nomic via one/api/embeddings, keep the chosen whisper model loaded, then transcribe again. Record total VRAM, whether anything OOM'd or fell back to CPU (checkdocker logs ollama --since 5m | grep -i -E "cuda|offload|memory"— read-only), and latency change. Leave XTTS out (V2 — not built; budget it as "XTTS-v2 ≈ 2–3 GB, UNVERIFIED"). - Write a VRAM budget section (if
/opt/appdata/docker/research/jenkins_hermes_state_2026-09-28.mdexists, read which LLM backend Hermes uses — if it is server-01 Ollama, Hermes is a fourth GPU tenant; include it): measured numbers for Ollama-8B + nomic + STT, the remaining headroom for TTS, and 2–3 OPTIONS the main session will choose from (e.g. shorter Ollama keep-alive /OLLAMA_MAX_LOADED_MODELS, STT on CPU fallback, smaller LLM) — options only, do not apply any. - Leave
voice-sttRUNNING with the chosen default model configured; verify from PRIMARY:curl -s -F file=@<a local copy of jfk.wav> -F model=<id> http://192.168.1.90:8300/v1/audio/transcriptions(or the fallback image's endpoint) returns the transcript. Delete the local copy after. - Restore check: the container set equals the step-1 snapshot +
voice-stt.
DELIVERABLES
/data/docker-compose/voice-stt/docker-compose.yml+README.mdon server-01 (what it is, endpoint, model, how to change model, rollback =docker compose down, the exact client call).- A copy of both files + the report into the primary repo:
/opt/appdata/docker/docker-compose/voice-stt/(compose + README) and/opt/appdata/docker/research/voice_stt_v1_2026-09-28.md(results table, VRAM budget, options, UNVERIFIED list, the exact invocation the laptop daemon will call).
SCOPE ALLOWLIST: server-01 /data/docker-compose/voice-stt/, /data/models/voice-stt/, /tmp/voice-stt-test/ (deleted), container voice-stt; primary /opt/appdata/docker/docker-compose/voice-stt/, the report, /opt/appdata/docker/docker-compose/voice-stt/.claude/context.md. Nothing else.
PERSIST BEFORE YOU FINISH
- Create
/opt/appdata/docker/docker-compose/voice-stt/.claude/context.md(template/home/administrator/.claude/projects/-opt-appdata-docker/memory/playbook_project_context_template.md), append "## V1 STT — 2026-09-28 (background agent)". - Run: python3 /opt/appdata/docker/.claude/scripts/embed_memory_dir.py --only-recent 5
- FINAL message = wrap-up JSON only, always: {"status":"succeeded|partially_succeeded|failed","project":"voice chat #214","phase":"V1 STT","actions_taken":[],"actions_failed":[],"files_touched":[],"containers_restarted":[{"name":"voice-stt","action":"","health":""}],"image":{"ref":"","tag":"","digest":"","fallback_used":false},"models":[{"id":"","transcript_ok":true,"rtf":0.0,"vram_mib":0}],"chosen_model":"","coexistence":{"total_vram_mib":0,"oom":false,"ollama_cpu_fallback":false},"endpoint":"http://192.168.1.90:8300/...","unverified_claims":[],"next_step":"","notes":""}