Saved prompts: W2, OLLAMA-1, BOOT-1, S1, AS0, AS1, JH-1, V0, V1, VS-1. DECISIONS.md 2026-09-27 entry (sudo-bridge retired, Jenkins deploys via agent-sudo deploy_service, Chatterbox-Turbo, vault-sandbox auto-unseal). Voice A1/A2 superseded. Redacted two plaintext secrets in agent-builder context (still in history; rotation tracked under #192). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
7.0 KiB
OLLAMA-1 — make server-01 (GPU) the ONLY Ollama; repoint primary callers; write the retirement steps
Written 2026-09-27 (infrastructure general questions conversation). Owner decision 2026-09-27: run ONE Ollama, on server-01's GPU; retire primary's CPU-only copy. Safe to run in parallel with W2, S1, AS0, JH-1, BOOT-1.
You are a bounded background agent. Task: repoint every primary-side Ollama caller to server-01, prove it works, and write (not run) the owner steps that retire primary's Ollama. Budget: --max-turns 25. If you issue the same tool call twice with identical arguments, STOP and output the wrap-up block with status=partially_succeeded.
HARD RULES
- Do NOT stop/start/restart/remove any container, do NOT touch systemd units, no sudo, no git add/commit/push, no Postgres writes except what
embed_memory_dir.pyitself does when you run it. The owner retires the primary container. - Never print secret values. If a hook blocks something, stop with a partial wrap-up.
FACTS (verified by the main session 2026-09-27 14:00 — re-check only the ones marked RE-CHECK)
- Primary (192.168.1.88) has no NVIDIA GPU (AMD Cezanne iGPU only). Its Ollama = container
ollama-ollama-1, CPU-only, published127.0.0.1:11434, compose at/opt/appdata/docker/docker-compose/ollama/(secretspec-managed:ollama-secretspec-resolver.service+ an ENABLEDollama-secretspec-resolver.timerthat re-ups it periodically — so stopping the container alone is not enough; the timer must be disabled too). - server-01 (192.168.1.90) Ollama = container
ollama(imageollama-fixed:1.0.0, runtime nvidia, RTX 2060 SUPER 8 GB, logs show 13/13 layers offloaded to CUDA0), published0.0.0.0:11434(LAN). Compose/data/docker-compose/ollama/docker-compose.yml. - Both hosts hold the same models with identical digests:
nomic-embed-text:latest0a109f422b47e3a3…,llama3.1:8b46e0c10c039e0191… → embeddings from either host are the same vectors; NO re-index needed. - Known callers:
/opt/appdata/docker/.claude/scripts/embed_memory_dir.pyline ~24OLLAMA_URL = "http://localhost:11434/api/embeddings"(→ primary CPU — THE one to repoint);.claude/scripts/semantic_recall.pyalready useshttp://192.168.1.90:11434;non-docker-python-scripts/youtube-pipeline/script_gen.pyalready uses 192.168.1.90.legacy-docker-compose/open web uipoints at .88 but is legacy/not running — leave it, just list it. - RE-CHECK: grep for any other caller the main session missed:
grep -rn -E "11434|OLLAMA_(HOST|URL|BASE)" /opt/appdata/docker/.claude /home/administrator/.claude/hooks /home/administrator/.claude/scripts /home/administrator/.claude/skills /home/administrator/.claude/settings.json /opt/appdata/docker/non-docker-python-scripts /opt/appdata/docker/docker-compose --include=*.py --include=*.sh --include=*.json --include=*.yml --include=*.yaml --include=*.toml(excludepin-bumper/.cache,logs/,*.bak*). Also check Docker networks: is any running primary container configured to reachollama-ollama-1by name (docker inspectConfig.Image/NetworkSettings only — grep compose files forollama:hostnames)? Also check the Stop hook's embed path (the hook that embeds[[MEMORY_EMBED]]tags — find which script it calls and what URL that script uses).
DECISIONS (locked — do not revisit)
- Every repointed caller reads
OLLAMA_URLfrom the environment with defaulthttp://192.168.1.90:11434/api/embeddings(or the matching base URL + path style the file already uses). One-line change per file + a one-line comment "single Ollama lives on server-01 GPU (2026-09-28)". Keep the file's existing style. - server-01's Ollama stays as it is (LAN-exposed, no auth — router does not forward 11434). Do not change it; note the exposure in the report as a known accepted risk.
- Failure behaviour: if server-01 is unreachable,
embed_memory_dir.pymust fail LOUDLY (non-zero exit,[EMBED] FAILEDline) — it already does since #183; confirm, do not add a silent fallback to primary.
STEPS
- RE-CHECK callers (above). Record each file + line.
- Back up each file you will edit to
<file>.bak-20260928-ollama(cp), then edit. python3 -c "import ast; ast.parse(open('<file>').read())"for each edited .py.- Prove it: run
python3 /opt/appdata/docker/.claude/scripts/embed_memory_dir.py --only-recent 1and confirm exit 0 +[EMBED] Done. Then prove the TRAFFIC went to server-01, not primary:curl -s 192.168.1.90:11434/api/psright after (nomic-embed-text should be loaded there with size_vram > 0) andcurl -s localhost:11434/api/pson primary (should NOT show it newly loaded — note expires_at). - Run one semantic recall query end-to-end (
python3 /opt/appdata/docker/.claude/scripts/semantic_recall.py "mediashare drive re-enumeration"or its documented CLI — read its argparse first) and confirm results come back. - Write the OWNER retirement runbook into
/opt/appdata/docker/docker-compose/ollama/RETIRE.md: (a)sudo systemctl disable --now ollama-secretspec-resolver.timer, (b)docker stop ollama-ollama-1(keep the container + volume 7 days as rollback), (c) 24-hour watch list (next end-session embed run succeeds; /recall works), (d) after 7 days:docker compose -f /opt/appdata/docker/docker-compose/ollama/docker-compose.yml down(NO-vuntil the owner decides), how much RAM/disk it frees (measure the volume size withdu -shon the bind/volume path viadocker inspectMounts, and the image size), (e) rollback =sudo systemctl enable --now ollama-secretspec-resolver.timer. - Update docs that state where Ollama lives: grep the memory dir
/home/administrator/.claude/projects/-opt-appdata-docker/memory/for "localhost:11434" / "ollama" claims about primary (skipfile_snapshot_*); list them in the report; edit ONLY lines that are now factually wrong (append "(2026-09-28: single Ollama = server-01 GPU)" rather than rewriting paragraphs).
SCOPE ALLOWLIST: the caller files you repoint (+ their .bak copies), /opt/appdata/docker/docker-compose/ollama/RETIRE.md, /opt/appdata/docker/docker-compose/ollama/.claude/context.md (create from /home/administrator/.claude/projects/-opt-appdata-docker/memory/playbook_project_context_template.md if missing), the memory-dir lines in step 7. Nothing else.
PERSIST BEFORE YOU FINISH
- Append "## OLLAMA-1 — 2026-09-28 (background agent)" to
/opt/appdata/docker/docker-compose/ollama/.claude/context.md(What was done / Current state / Owner steps / Next step). - Run: python3 /opt/appdata/docker/.claude/scripts/embed_memory_dir.py --only-recent 5
- FINAL message = wrap-up JSON only, always: {"status":"succeeded|partially_succeeded|failed","project":"ollama single-instance","phase":"OLLAMA-1","actions_taken":[],"actions_failed":[],"files_touched":[],"containers_restarted":[],"callers_found":[{"file":"","line":0,"old":"","new":"","status":"repointed|already_server01|legacy_left"}],"proof":{"embed_exit":0,"server01_ps_shows_model":true,"primary_ps_idle":true,"recall_ok":true},"primary_ollama_frees":{"ram":"","disk":""},"unverified_claims":[],"next_step":"owner runs RETIRE.md","notes":""}