Files
claude-projects/claude-config/config/prompts/infra/OLLAMA-1_single_instance_server01.md
T
Backtalk6858 d9419109a5 docs(claude-config): 2026-09-27 preflight prompts, owner decisions, redactions
Saved prompts: W2, OLLAMA-1, BOOT-1, S1, AS0, AS1, JH-1, V0, V1, VS-1.
DECISIONS.md 2026-09-27 entry (sudo-bridge retired, Jenkins deploys via
agent-sudo deploy_service, Chatterbox-Turbo, vault-sandbox auto-unseal).
Voice A1/A2 superseded. Redacted two plaintext secrets in agent-builder
context (still in history; rotation tracked under #192).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 12:46:31 -05:00

7.0 KiB

OLLAMA-1 — make server-01 (GPU) the ONLY Ollama; repoint primary callers; write the retirement steps

Written 2026-09-27 (infrastructure general questions conversation). Owner decision 2026-09-27: run ONE Ollama, on server-01's GPU; retire primary's CPU-only copy. Safe to run in parallel with W2, S1, AS0, JH-1, BOOT-1.


You are a bounded background agent. Task: repoint every primary-side Ollama caller to server-01, prove it works, and write (not run) the owner steps that retire primary's Ollama. Budget: --max-turns 25. If you issue the same tool call twice with identical arguments, STOP and output the wrap-up block with status=partially_succeeded.

HARD RULES

  • Do NOT stop/start/restart/remove any container, do NOT touch systemd units, no sudo, no git add/commit/push, no Postgres writes except what embed_memory_dir.py itself does when you run it. The owner retires the primary container.
  • Never print secret values. If a hook blocks something, stop with a partial wrap-up.

FACTS (verified by the main session 2026-09-27 14:00 — re-check only the ones marked RE-CHECK)

  • Primary (192.168.1.88) has no NVIDIA GPU (AMD Cezanne iGPU only). Its Ollama = container ollama-ollama-1, CPU-only, published 127.0.0.1:11434, compose at /opt/appdata/docker/docker-compose/ollama/ (secretspec-managed: ollama-secretspec-resolver.service + an ENABLED ollama-secretspec-resolver.timer that re-ups it periodically — so stopping the container alone is not enough; the timer must be disabled too).
  • server-01 (192.168.1.90) Ollama = container ollama (image ollama-fixed:1.0.0, runtime nvidia, RTX 2060 SUPER 8 GB, logs show 13/13 layers offloaded to CUDA0), published 0.0.0.0:11434 (LAN). Compose /data/docker-compose/ollama/docker-compose.yml.
  • Both hosts hold the same models with identical digests: nomic-embed-text:latest 0a109f422b47e3a3…, llama3.1:8b 46e0c10c039e0191… → embeddings from either host are the same vectors; NO re-index needed.
  • Known callers: /opt/appdata/docker/.claude/scripts/embed_memory_dir.py line ~24 OLLAMA_URL = "http://localhost:11434/api/embeddings" (→ primary CPU — THE one to repoint); .claude/scripts/semantic_recall.py already uses http://192.168.1.90:11434; non-docker-python-scripts/youtube-pipeline/script_gen.py already uses 192.168.1.90. legacy-docker-compose/open web ui points at .88 but is legacy/not running — leave it, just list it.
  • RE-CHECK: grep for any other caller the main session missed: grep -rn -E "11434|OLLAMA_(HOST|URL|BASE)" /opt/appdata/docker/.claude /home/administrator/.claude/hooks /home/administrator/.claude/scripts /home/administrator/.claude/skills /home/administrator/.claude/settings.json /opt/appdata/docker/non-docker-python-scripts /opt/appdata/docker/docker-compose --include=*.py --include=*.sh --include=*.json --include=*.yml --include=*.yaml --include=*.toml (exclude pin-bumper/.cache, logs/, *.bak*). Also check Docker networks: is any running primary container configured to reach ollama-ollama-1 by name (docker inspect Config.Image/NetworkSettings only — grep compose files for ollama: hostnames)? Also check the Stop hook's embed path (the hook that embeds [[MEMORY_EMBED]] tags — find which script it calls and what URL that script uses).

DECISIONS (locked — do not revisit)

  1. Every repointed caller reads OLLAMA_URL from the environment with default http://192.168.1.90:11434/api/embeddings (or the matching base URL + path style the file already uses). One-line change per file + a one-line comment "single Ollama lives on server-01 GPU (2026-09-28)". Keep the file's existing style.
  2. server-01's Ollama stays as it is (LAN-exposed, no auth — router does not forward 11434). Do not change it; note the exposure in the report as a known accepted risk.
  3. Failure behaviour: if server-01 is unreachable, embed_memory_dir.py must fail LOUDLY (non-zero exit, [EMBED] FAILED line) — it already does since #183; confirm, do not add a silent fallback to primary.

STEPS

  1. RE-CHECK callers (above). Record each file + line.
  2. Back up each file you will edit to <file>.bak-20260928-ollama (cp), then edit.
  3. python3 -c "import ast; ast.parse(open('<file>').read())" for each edited .py.
  4. Prove it: run python3 /opt/appdata/docker/.claude/scripts/embed_memory_dir.py --only-recent 1 and confirm exit 0 + [EMBED] Done. Then prove the TRAFFIC went to server-01, not primary: curl -s 192.168.1.90:11434/api/ps right after (nomic-embed-text should be loaded there with size_vram > 0) and curl -s localhost:11434/api/ps on primary (should NOT show it newly loaded — note expires_at).
  5. Run one semantic recall query end-to-end (python3 /opt/appdata/docker/.claude/scripts/semantic_recall.py "mediashare drive re-enumeration" or its documented CLI — read its argparse first) and confirm results come back.
  6. Write the OWNER retirement runbook into /opt/appdata/docker/docker-compose/ollama/RETIRE.md: (a) sudo systemctl disable --now ollama-secretspec-resolver.timer, (b) docker stop ollama-ollama-1 (keep the container + volume 7 days as rollback), (c) 24-hour watch list (next end-session embed run succeeds; /recall works), (d) after 7 days: docker compose -f /opt/appdata/docker/docker-compose/ollama/docker-compose.yml down (NO -v until the owner decides), how much RAM/disk it frees (measure the volume size with du -sh on the bind/volume path via docker inspect Mounts, and the image size), (e) rollback = sudo systemctl enable --now ollama-secretspec-resolver.timer.
  7. Update docs that state where Ollama lives: grep the memory dir /home/administrator/.claude/projects/-opt-appdata-docker/memory/ for "localhost:11434" / "ollama" claims about primary (skip file_snapshot_*); list them in the report; edit ONLY lines that are now factually wrong (append "(2026-09-28: single Ollama = server-01 GPU)" rather than rewriting paragraphs).

SCOPE ALLOWLIST: the caller files you repoint (+ their .bak copies), /opt/appdata/docker/docker-compose/ollama/RETIRE.md, /opt/appdata/docker/docker-compose/ollama/.claude/context.md (create from /home/administrator/.claude/projects/-opt-appdata-docker/memory/playbook_project_context_template.md if missing), the memory-dir lines in step 7. Nothing else.

PERSIST BEFORE YOU FINISH

  • Append "## OLLAMA-1 — 2026-09-28 (background agent)" to /opt/appdata/docker/docker-compose/ollama/.claude/context.md (What was done / Current state / Owner steps / Next step).
  • Run: python3 /opt/appdata/docker/.claude/scripts/embed_memory_dir.py --only-recent 5
  • FINAL message = wrap-up JSON only, always: {"status":"succeeded|partially_succeeded|failed","project":"ollama single-instance","phase":"OLLAMA-1","actions_taken":[],"actions_failed":[],"files_touched":[],"containers_restarted":[],"callers_found":[{"file":"","line":0,"old":"","new":"","status":"repointed|already_server01|legacy_left"}],"proof":{"embed_exit":0,"server01_ps_shows_model":true,"primary_ps_idle":true,"recall_ok":true},"primary_ollama_frees":{"ram":"","disk":""},"unverified_claims":[],"next_step":"owner runs RETIRE.md","notes":""}