Files
claude-projects/agent-builder/.claude/context.md
T
2026-06-26 16:08:21 -05:00

475 lines
40 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
project_name: agent-builder
# Agent Builder — Session Context
## What this project does
Design, build, and test autonomous N8N agents on server-01 sandbox before any production promotion.
First two agents: Agent Builder Agent + N8N Builder Agent.
## Scheduled work (2026-06-16, running behind — started ~6:38 PM)
1. Vision Alignment Grill-Me — Agent Builder + N8N Builder vision + testing methodology
2. Agent Builder Agent — Deploy + Test (server-01 sandbox)
3. N8N Builder Agent — Deploy + Test (server-01 sandbox)
## Architecture
- Agents run as N8N workflows on server-01 (n8n-sandbox, port 5679)
- Sandbox-first: all agents tested in sandbox before any production promotion
- server-01 sandbox stack: n8n-sandbox, postgres-sandbox, vault-sandbox, bitwarden-bridge-sandbox, vaultwarden-sandbox
- Sandbox N8N API key: prod Vault at secret/sandbox/n8n
- Sandbox reachable at 192.168.1.90
## Key decisions (set during vision grill-me — 2026-06-16)
- Agent Builder Agent: builds `claude_agent` and `script` types — Ollama (llama3.1:8b) does the building, `claude -p` is overseer/validator
- N8N Builder Agent: builds `n8n_automation` types — Ollama generates workflow JSON, imports via N8N API, assigns credentials
- automation_ideas schema changes needed: rename `description``task_description` (full structured spec), add `type` (n8n_automation/claude_agent/script), add `builder_status`
- New `agent_test_results` table needed in api_business DB
- Sandbox must mirror production: AppRole, Vaultwarden, bridge all configured before any agent deploys
- Promotion = user approval required after all 4 test levels pass (not auto-promote in v1)
- Dedicated backfill session needed for all 48 existing automation_ideas rows (type + task_description)
- claude -p uses SDK credits (Pro = $20/month hard limit) — use sparingly, Ollama does the heavy lifting
- Local model: llama3.1:8b already pulled on server-01 (4.9GB, fits in RTX 2060 Super 8GB VRAM)
## Testing methodology
- Four levels: Structure → Deployment → Smoke → Assertion
- LLM outputs validated on structure/side-effects only, never exact string match
- All results logged to agent_test_results table
- NTFY notification on pass and fail
- Full methodology: .claude/playbook_testing_methodology.md
## Agents
### Agent Builder Agent
- Status: pending — prereqs not complete
- Purpose: Receives automation spec from automation_ideas DB, uses Ollama to build claude_agent or script type automations, deploys to sandbox, runs automated tests, notifies user for promotion approval
- Builds: claude agents (via claude -p) and Python scripts (Docker containers)
### N8N Builder Agent
- Status: pending — prereqs not complete
- Purpose: Receives automation spec from automation_ideas DB, uses Ollama to generate N8N workflow JSON using n8n_automations playbook as context, imports to sandbox N8N via API, assigns credentials, runs automated tests
- Will be used to build: id=12 (Media Pipeline Learning), id=7 (Friday Research Session Prep)
## Related personal_projects DB rows
- id=4: N8N Workflow Builder Script (pending, weekend_block1)
## Prereq checklist (must complete before any agent deployment)
- [x] Schema: rename automation_ideas.description → task_description, add type, add builder_status, add priority
- [x] Create agent_test_results table in api_business
- [x] Sandbox Vault: set up AppRole auth method (credentials at /opt/appdata/docker/docker-compose/vault/approle/ on server-01)
- [x] Sandbox Vault: store sandbox N8N API key at secret/sandbox/n8n (key name: claude-sandbox, verified working)
- [x] Verify sandbox Bitwarden bridge ↔ Vaultwarden sandbox end-to-end (bridge on port 8080, returns [] for empty vault — correct)
- [x] Write Agent Builder Agent playbook → .claude/playbook_agent_builder_agent.md
- [x] Write N8N Builder Agent playbook → .claude/playbook_n8n_builder_agent.md
- [x] Backfill session: **COMPLETE.** All pending rows reviewed across sessions 1-3. Blocked rows verified (32, 30, 13, 27, 23, 24, 22, 2, 21, 35, 36, 55, 56, 18 — all still blocked, no changes without proper review). Admin done: id=28 → p10, id=57 keeps p9, dummy row id=16 deleted. Session 3 (2026-06-18) final batch: ids 47, 48, 31, 26, 54.
- id=25: type→n8n_automation, status→pending (8b model lifts hardware block), full rewrite with SDK credit model, model router + cost gate + subscription monitor, model_effort_routing_log training table
- id=17: type→n8n_automation, expanded with token_waste_patterns DB table (evolution mechanism), JSONL-based detection, NTFY claude-audit gate, fan-out plan noted
- id=42: expanded with hybrid trigger (Claude flags + asks), B+C output (DB + wrap-up surface), human gate permanent, hard boundary (business research = id=55 only), voice note
- id=46: type→n8n_automation, Hermes ref removed, playbook_schema.md bootstrap design, three risk tiers, id=52 detector/id=46 executor boundary, fan-out plan noted
- id=43: Vault AppRole required (youtube-oauth-script, new policy), vault_registry entries, port 8085 (8080 reserved for bridge), dynamic Vault IP, write verification, token refresh companion
- id=45: hard dep on id=43 (Vault auth), niche_saturation_thresholds DB table (evolution), niche_check_results table, three-signal output (saturated/viable/unclear), id=18 stage 2 integration
## New automation_ideas rows added 2026-06-18 (ids 5967)
- id=59: Claude Audit Research Prep (n8n_automation, p12) — Saturday briefing for 12:45 PM review
- id=60: Claude Audit Topic Detector (n8n_automation, p11) — detects new topics → claude_config.audit_topics; must build first
- id=61: Claude Config Dev Work Scheduler (n8n_automation, p13) — Sunday session prep from Saturday decisions
- id=62: Monthly Security Audit (n8n_automation, p7) — first Monday of month, 7-area audit → security_audit DB
- id=63: Security Patch Backlog Handler (claude_agent, p7) — works security_audit.findings; Vault/prod always human-gate
- id=64: Business Projects Worker (claude_agent, BLOCKED p29) — 12:454:30 PM daily; blocked pending business_projects schema
- id=65: Personal Projects Worker (claude_agent, BLOCKED p29) — 5:308:00 PM daily; blocked pending personal_projects schema
- id=66: Media Pipeline Project Worker (claude_agent, BLOCKED p29) — Wednesday 5:308:00 PM; blocked pending id=65 + scope
- id=67: N8N Builder Agent (claude_agent, p2, ready_to_build) — builds all n8n_automation rows; deploy Monday June 22 alongside id=24
## Readiness check — COMPLETE (2026-06-18 evening)
All 8 prereq checklist items verified live on server-01:
- Vault AppRole: token acquired, N8N secret readable
- N8N sandbox API: returns 200 with real key
- Bridge /items: returns [] (correct for empty sandbox vault)
- Ollama: llama3.1:8b loaded
- agent_test_results table: confirmed correct schema + FK
## June 20 (Saturday) — Session progress
**id=57 partially complete.** Carried over from June 19 (sick day).
### id=57 — DONE this session:
- Production N8N migrated from primary server → server-01 port 5678
- Vault-backed start.sh: secrets pulled from production Vault at runtime, nothing on disk
- n8n-server01 AppRole created in production Vault (scoped to secret/data/n8n read-only)
- Stale root token in Bitwarden replaced — regenerated via generate-root (3/5 keys), updated, revoked
- Traefik static route live: /traefik/dynamic/n8n.yaml → 192.168.1.90:5678
- Rollback procedure documented + all 4 checks passed: /opt/appdata/docker/docker-compose/n8n/ROLLBACK.md
- claude-policy extended: AppRole management + sys/generate-root paths added
### id=57 — REMAINING (next session):
- Sandbox Vaultwarden bridge fix: 53 ciphers seeded (DB confirmed), but bw CLI WASM crash on list — bridge /items returns 500. Fix: update Vaultwarden image OR rebuild bridge with newer bw CLI OR patch bridge /items to use REST API
- N8N sandbox credentials (-sandbox suffix names) — Bridge API Key already exists; need postgres, Vault, NTFY
- Sudo bridge deploy on server-01 (Phase 2) — then disable NOPASSWD:ALL in /etc/sudoers.d/administrator on server-01
### id=51 — NOT STARTED. Scope locked:
- Option B zero-exposure proxy: agent sends "run X using secret Y", proxy executes + injects secret, returns result only
- Must be live before builder agents deploy
## June 20 Session 2 (continuation)
- Sandbox Vault auto-unseal deployed: vault-sandbox-unseal.sh + vault-sandbox-watch-unseal.sh + 2 systemd services, unsealed and verified
- Sandbox Vault AppRole mirror: 9 policies + 5 roles recreated from production (claude-code, claude-policy, n8n, n8n-outreach-policy, n8n-policy, n8n-rotation-policy, n8n-scheduling-policy, n8n-server01-policy, nextcloud-init)
- Production Bitwarden audit: 294 items total, 53 infra items identified (all username/password, some with notes)
- Sandbox Vaultwarden population BLOCKED: bw CLI 2026.5.0 WASM crypto error on create — use Vaultwarden REST API next session (id=70)
- sudo bridge NOT on server-01: server-01 has NOPASSWD:ALL; plan = deploy bridge first (Phase 2), then disable NOPASSWD
- ids added: 69 (Vaultwarden Version Monitor, p16), 70 (Sandbox Vaultwarden Seeder, p17)
- feedback_bw_cli_vaultwarden_create.md added to docker MEMORY.md
## June 20 Session 3 (continuation)
### Sandbox Vaultwarden — 53/53 dummy ciphers seeded (DB confirmed)
- Root cause of login failure: Bitwarden-Client-Version header required by Vaultwarden 1.36.0
- Original akey was encrypted with unknown master key (registration vs unlock password mismatch)
- Fix: reset sandbox@test.local password_hash + akey in postgres-sandbox via /tmp/vw_reset_and_seed.py
- Script: derives keys from u8X_P_zlGBypRUXYq5iGIg + PBKDF2/HKDF, generates fresh user_sym_key,
updates DB directly, logs in via REST API, POSTs 53 AES-256-CBC encrypted ciphers
- Runs via: `docker run --rm --network n8n-sandbox_default -v /tmp/vw_reset_and_seed.py:/tmp/script.py bitwarden-bridge:sandbox sh -c 'pip install cryptography psycopg2-binary --quiet && python3 /tmp/script.py'`
- Script at /tmp/vw_reset_and_seed.py on server-01 (also /tmp/vw_reset_and_seed.py locally)
- 53 ciphers confirmed: `SELECT COUNT(*) FROM ciphers WHERE user_uuid='896b5bbb-...'` → 53
- Bridge unlocks successfully with new password on every restart
### BLOCKER: bw list items WASM crash (same as bw create)
- Error: "Invalid key, throwing away stored keys" × 2, then "invalid type: unit value, expected a valid string"
- Empty vault worked; non-empty vault fails — bw CLI 2026.5.0 can't decrypt Vaultwarden 1.36.0 ciphers
- Bridge /items returns 500 error; ciphers ARE in DB but not serveable via bridge
- Fix options (next session, pick one):
1. Pull updated vaultwarden/server:latest (Docker pull currently broken due to daemon instability on server-01 — containerd-based Docker 29.x; images intact at /var/lib/containerd 26G)
2. Rebuild bitwarden-bridge:sandbox with newer bw CLI version
3. Patch bridge /items endpoint to use Vaultwarden REST API instead of bw CLI
- /tmp/vw_reset_and_seed.py script is reusable — if Vaultwarden is updated, just re-run to re-seed after container recreate
### Docker daemon incident (unrelated to our work)
- All containers exited during session — Docker 29.x daemon crash (not caused by our actions)
- Images are in /var/lib/containerd (26G), not /var/lib/docker/overlay2 (new storage model)
- Stack brought back up via: `cd /opt/appdata/docker/docker-compose/server-01 && docker compose up -d`
- Compose file location confirmed: /opt/appdata/docker/docker-compose/server-01/docker-compose.yml
### id=57 — COMPLETE ✅ (June 20 Session 5)
All sandbox prereqs done. Sandbox mirrors production.
**Completed in Session 5:**
- Bridge switched to Bitwarden cloud (megafreeman12@proton.me) — unlocks cleanly, /items returns 53
- Compose updated: BW_SERVER and NODE_TLS_REJECT_UNAUTHORIZED removed; vaultwarden-sandbox depends_on removed
- 53 dummy items seeded (dummy-infra-01 through dummy-infra-53)
- sandbox Vault secret/bitwarden updated with new master_password
- N8N sandbox credentials created: postgres-sandbox, vault-sandbox, n8n-internal-sandbox, Bridge API Key (Sandbox)
- n8n_agent_worker postgres role — api_business DB, scoped to automation_ideas + agent_test_results
- Credential in production Vault: secret/postgres/n8n-agent-worker
### id=51 — DEPLOYED ON PRIMARY ✅ (June 22 session 3)
Container running healthy at `https://secrets-proxy.reverseproxyserver.net` on primary.
Coolify service UUID: `ilus0cfdkheipodw1viurg1d`
**June 22 session 3 facts:**
- Coolify API allowed_ips was `172.16.16.1/32` — cleared to allow API access (set to empty string in instance_settings)
- Server-01 service (lv3xeu2manuimle458wbm9u9) stopped and deleted ✅
- Sandbox Vault and bridge now LAN-exposed on server-01: vault=192.168.1.90:8201, bridge=192.168.1.90:8083 (port bindings added to Coolify DB + sandbox compose file)
- secrets-proxy Vault AppRole created in prod (role: 3bf7b8a6-9e99-ca78-8e69-517881e26ea5) — creds at secret/proxy/vault-approle
- secrets-proxy Vault AppRole created in sandbox (role: f21416b3-2f3a-0461-fa38-d453a4e32865) — creds at secret/proxy/vault-approle-sandbox
- secrets_proxy DB user created in api_business (INSERT on proxy_executions only) — creds at secret/postgres/secrets-proxy
- NTFY secrets-proxy-bot created (topic: secrets-proxy-notifications, token at secret/ntfy/secrets-proxy-bot)
- Coolify bug: env vars for custom-image services can only be set via PATCH /api/v1/services/{uuid}/envs with key+value (not POST)
- Coolify DB YAML serializer mangles CMD-SHELL healthcheck single quotes (`'``''`) — fixed by using CMD array form in healthcheck
- psql heredoc via `docker exec ... << 'SQL'` doesn't work — use `docker cp` + `psql -f` instead
- pull_policy: never removed from docker-compose.yml (was server-01 workaround, no longer needed on primary)
- docker_compose_raw in Coolify DB must be updated via `docker cp` + `psql -f` for reliability
**June 22 session 4 facts:**
- E2E tests all passed: bridge_health_check (sandbox+prod), n8n_workflow_list (12 workflows), vault_secret_exists, vault_secret_write ✅
- proxy path format: callers pass RELATIVE path (e.g. `n8n`, `e2e-test`) — proxy prepends `secret/data/` automatically
- params go inside a `params: {}` object in the POST body (not at top level)
- iptables DNAT rule persisted: /etc/iptables/rules.v4 saved, /etc/sysctl.d/99-local-routing.conf written ✅
- iptables-persistent + netfilter-persistent already installed
- Used sudo-bridge (port 8082 on 192.168.1.88) with new bash -c allowlist entries
- POST /allowlist + GET /allowlist endpoints added to sudo-bridge ✅
- Validates tier (1/2), rejects danger-pattern matches, 409 on duplicate
- Writes to allowlist.json on disk (daemon reloads per-request — no restart needed)
- Audits to JSONL + Postgres, sends NTFY notification
- New image built, pushed to gitea.local, Coolify start called via proxy shell + env_secret injection
- sudo-bridge port: 8082 on 192.168.1.88 (LAN), container name: sudo-bridge-drbjegv07256ki2lpfyr00n8
- Coolify API from within Docker network: http://172.16.16.30:8080/api/v1/ (not 172.16.16.35)
- Coolify token: vault://coolify#api_key (field name is api_key)
- To call Coolify API from proxy: use /shell with env_secrets: {"TOKEN": "vault://coolify#api_key"}
- sudo-bridge image rebuild pattern: build → push to gitea.local → stop container → POST /api/v1/services/{uuid}/start
**June 22 session 4 (continued) facts:**
- Shell E2E all passed: read (sandbox_exit:0), write (notify+exec), destructive (NTFY approval→exec) ✅
- proxy-sandbox-mirror image: built from /secrets-proxy/mirror/Dockerfile — must exist on HOST (proxy calls host Docker daemon via socket)
- Dockerfile fix: docker.io → docker-cli (docker.io with --no-install-recommends on Debian Trixie does not install binary)
- NTFY bug: _ntfy() was hardcoding URL to secrets-proxy-notifications regardless of topic — fixed with topic param
- NTFY ACL: secrets-proxy-bot needed write-only on secrets-proxy-approvals; Smoked5003 needed read-only — both added via ntfy CLI in container
- secrets_proxy DB user: needed SELECT on proxy_executions for ON CONFLICT DO NOTHING — granted
- Coolify token plaintext: 2|95eQySElT9uQpTXqDACWq1z9kyOaySZOZP8sBxyJaebc2bbe (Sanctum format: id|plaintext)
- Coolify "Service is already running" after stop+rm: UPDATE service_applications SET status='stopped' WHERE service_id=... then POST /start
- 524 on destructive commands: Cloudflare kills long-poll at ~100s; approval must be tapped quickly; command still runs if approved before timeout, response just lost
**NEXT SESSION — in order:**
0. sudo-bridge (both) — fix allowlist.json file ownership
- Problem: allowlist.json is administrator-owned on both servers — Claude can bypass POST /allowlist entirely with a direct file edit, defeating the audit log, NTFY notification, and danger-pattern veto
- Fix: chown root:root + chmod 644 allowlist.json on both primary and server-01 — daemon writes via root (systemd), Claude must use POST /allowlist endpoint
- After fix: verify POST /allowlist still works (daemon writes as root), verify direct Edit tool to allowlist.json is denied
- Sync policy decision: NO cross-bridge sync — allowlists grow organically per server (primary ≠ server-01 purposes); when a command is needed on server-01 it gets added there at that time
0a. sudo-bridge-server01 — sync POST /allowlist endpoint
- POST /allowlist was added to sudo-bridge (primary) this session but server-01 bridge container was NOT redeployed — it's still running the old image without the endpoint
- Both bridges share gitea.local/backtalk6858/sudo-bridge:latest — code is already in the image, just need to redeploy server-01 bridge (Coolify UUID: o2kz1puml1mmneyiqd96mouj) via pull + restart
- Also: server-01 allowlist.json needs the 4 new entries added today (iptables-save, 99-local-routing.conf, sysctl -p, journalctl -u) — these are primary-only entries so only add if relevant to server-01 use cases
- Verify: curl POST /allowlist and GET /allowlist on sudo-bridge-server01.reverseproxyserver.net after redeploy
1. id=51 Secrets Proxy — sandbox precondition scaffolding
- Problem: sandbox dry-run is meaningless if the required state doesn't exist (e.g. file to delete isn't there, service isn't running)
- Feature: before running command in proxy-sandbox-mirror, classify what the command needs to be present, check if it exists, and create/set it up if missing — so the sandbox test is a faithful replica of what would happen on the host
- Examples: `rm /tmp/foo` → create /tmp/foo first; `systemctl stop nginx` → start nginx in sandbox first; `iptables -D` → add the rule first
- Classification must handle: file existence, directory existence, running process/service, iptables rules, sysctl values
- After scaffold → re-run command → result shown in NTFY approval body
2. id=24 Agent Builder Agent + id=67 N8N Builder Agent
- Bitwarden bridge LAN-exposed: 192.168.1.88:8083 ✅
### June 21 session 2 facts for next session:
- server-01 Coolify UUID: hvzbj1gkqb5696s7cc9lcf8y
- Sandbox stack Coolify service UUID: d2celewbvh7e4fer77fcp9b5
- n8n-prod Coolify service UUID: h10eww4au274owpxozgizysh
- Traefik on primary binds to 127.0.0.1:80 only (LAN push workaround: docker save | ssh | docker load on primary then push)
- sudo_bridge DB: server_id column added to executions + allowlist_changes, default='primary'
- app.py needs: server_id='server-01' + [server-01] NTFY prefix variant for server-01 deployment
- sudo-bridge image: gitea.local/backtalk6858/sudo-bridge:latest (same image, different env vars)
- sudo-bridge-server01 Coolify UUID: o2kz1puml1mmneyiqd96mouj (port 8082, pull_policy:never)
- Vault secret created: secret/data/sudo-bridge-server01 (api_key: 443ea35b43c3e640aac3c57ed3aae06b8822ab6f)
- Host daemon running: /opt/appdata/docker/sudo-bridge/sudo_bridge_daemon.py (systemd, enabled)
- vaultwarden-sandbox REMOVED from sandbox stack (no longer needed — Bitwarden cloud dummy account used)
- app.py updated: SERVER_ID, _server_prefix(), server_id audit writes, User-Agent for Cloudflare
- NTFY approval action buttons use BRIDGE_EXTERNAL_URL=http://192.168.1.90:8082 (LAN only until Tailscale)
- NOPASSWD:ALL disabled on server-01 ✅ (June 21 session 3)
- Both bridges now on public URLs via Traefik + Cloudflare: sudo-bridge.reverseproxyserver.net + sudo-bridge-server01.reverseproxyserver.net
- Danger veto in daemon blocks rm /etc/* regardless of allowlist — /etc sudoers removal done manually
- NEXT: id=51 Secrets Execution Proxy (grill-me DONE June 21 session 4 — build next session)
## June 20 Session 4 (continuation)
### Docker 29.6.0 crash — diagnosed and fixed
- Root cause: nala upgrade swept in Docker 29.6.0 which has SIGSEGV null pointer dereference bug in HTTP transport during docker pull
- Fix: downgraded to 29.5.3, pinned with `apt-mark hold docker-ce docker-ce-cli docker-ce-rootless-extras`
- Rule: never run nala/apt upgrade without holding docker-ce first
### bw CLI + Vaultwarden — permanently abandoned
- bw CLI 2026.4.1 AND 2026.5.0 both crash against Vaultwarden 1.36.0 with `orgKeys null` TypeError
- Root cause: Vaultwarden returns null for orgKeys in sync response for personal vaults; bw CLI expects {}
- Decision: drop Vaultwarden sandbox, use Bitwarden cloud dummy account (same as production)
- Dummy account created: megafreeman12@proton.me, master_password=Infra6746Dummy$
- All credentials stored at secret/sandbox/bitwarden in production Vault (4 fields: email, master_password, client_id, client_secret)
- Sandbox bridge Dockerfile.sandbox moved to /opt/appdata/docker/docker-compose/server-01/bitwarden-bridge/ (pinned bw CLI 2026.4.1 — moot now but kept for reference)
- bridge:sandbox image rebuilt and pushed to gitea.reverseproxyserver.net (still has bw CLI 2026.4.1)
## June 22 (Monday) — Builder Agents (extended session)
Build, test, and push both builder agents to production. Work as long as it takes.
- id=24 Agent Builder Agent (claude_agent + script types)
- id=67 N8N Builder Agent (n8n_automation types)
- Prerequisites: id=51 live, id=57 FULLY complete (sandbox mirrors production)
## Priority queue summary (as of 2026-06-18)
- p1: id=51 (manual), id=57 (manual) — infrastructure foundation
- p2: id=24, id=67 — builder agents (built in sessions, build everything else)
- p3: id=54 NTFY Provisioner, id=58 Human Action Gate
- p4: id=25 Cost Intelligence, id=38 Credential Emergency Rollout
- p5: id=3 Secrets Rotation, id=8 Vault Token Audit, id=52 Session Wrap-up
- p6: id=37 N8N Log Scanner, id=40 Coolify UUID Monitor
- p7: id=62 Monthly Security Audit, id=63 Security Patch Handler
- p8p17: non-blocked pending/ready items
- p20p29: all blocked items (builder agents skip these)
## 2026-06-24 Sprint Day 1 facts
- id=51 COMPLETE ✅ — sandbox precondition scaffolding deployed. `_scaffold_preconditions()` added to secrets-proxy app.py. Creates files/dirs, attempts service starts, scaffolds iptables -D. Silent on success, `[Scaffold warning: ...]` in NTFY when uncertain.
- Coolify API key: Vault `secret/data/coolify → api_key` (NOT a file)
- Obsidian vault PARTIAL (id=136 in_progress): `/opt/appdata/obsidian/vault/` live on server-01, PARA structure, 258 memory files seeded. Remaining: ripgrep install, /recall redesign, session format Logseq→Obsidian
- Two new playbooks: playbook_secrets_proxy.md + playbook_sudo_bridge.md
- sudo-bridge rule: GET /allowlist before any POST — 409=already exists, empty exec response=rejected
- Gitea source migration logged: personal_projects id=145
**Next session first task:** Phase 3 — Gitea source migration (id=145). 10 services need repos. After that: Jenkins (id=141).
### 2026-06-24 Sprint Day 2 facts
- id=136 COMPLETE ✅ — semantic_recall.py rebuilt with ripgrep + chunked nomic-embed-text (SQLite), 290 files/434 chunks indexed; session format → Obsidian (Archives/Sessions/session-YYYY-MM-DD.md via SSH to server-01)
- Homemade skills at /opt/appdata/docker/.claude/skills/ — never search ~/.claude/plugins/
- Autonomous oversight model logged (behavior_changes id=77): AI brain grill-me + quality gates + NTFY escalation + Hermes orchestration + final human review
- sudo-bridge tier 1 NTFY: personal_projects id=146 (pending)
### 2026-06-24 Sprint Day 3 facts
- id=145 COMPLETE ✅ — 10 Gitea repos created (main branch), Jenkinsfiles pushed, webhooks configured
- id=141 IN PROGRESS — Jenkins deployed 192.168.1.90:8090, admin verified, JCasC working; remaining: pipeline jobs + Traefik route + E2E test
- New custom service pattern: boilerplates entry + separate Gitea repo (both always)
- dhi.io/jenkins = Docker Hardened Image, requires paid Docker subscription — not available for homelab use
- Gitea branch rename API (405) — use clone+push main+delete master pattern instead
- All 10 service repos: gitea.reverseproxyserver.net/Backtalk6858/<service>, branch=main
- secrets-proxy PROXY_CALLERS: claude-code, agent-builder-agent, n8n-builder-agent, jenkins
- Jenkins Vault AppRole: role-id/secret-id at /opt/appdata/docker/docker-compose/jenkins/vault-approle/ on server-01
- JCasC casc.yaml lives in jenkins-home/ (volume-mounted, not in image)
- Coolify retirement plan: use for everything except Jenkins now; retire fully after Jenkins E2E passes
**Next session first task:** id=141 remaining — Jenkins pipeline jobs (10 via API) + Traefik route + E2E smoke test.
### 2026-06-24 Sprint Day 4 facts
- 10 Jenkins pipeline jobs created via API — all use gitea-credentials + gitea.local URL + GWT token=<repo-name>
- Traefik route LIVE: jenkins.reverseproxyserver.net → 192.168.1.90:8090 via /data/coolify/proxy/dynamic/jenkins.yaml (written via docker exec on coolify-proxy, which mounts the dir)
- Jenkins Gitea credential (gitea-credentials) added to JCasC casc.yaml (server-01 only, volume-mounted) + start.sh + docker-compose.yml — survives restarts
- gitea.local resolves inside Jenkins container (extra_hosts in docker-compose.yml)
- Coolify proxy dynamic config: /data/coolify/proxy/dynamic/ (mount: /traefik/dynamic/ inside coolify-proxy container)
- Both servers run coolify-proxy (Coolify-managed Traefik); retirement plan in project_coolify_traefik_retirement.md
- Legacy Traefik config at legacy-docker-compose/traefik/ in boilerplates — starting point for standalone Traefik when Coolify retires
- E2E BLOCKED: docker binary missing from Jenkins image despite Dockerfile having docker.io
- Root cause: apt-get install docker.io on Debian Trixie does NOT install docker binary (known issue — package name is docker-ce or use docker-cli)
- Fix: update Dockerfile to install docker-ce OR docker-cli, rebuild with --no-cache, push to registry, restart Jenkins
**Next session first task (id=141 final steps):**
1. Fix Jenkins Dockerfile: replace `docker.io` with `docker-ce` (add Docker apt repo) OR `docker-cli` — confirm which package provides the binary on Debian Trixie
2. Rebuild Jenkins image with --no-cache on server-01
3. Push to gitea.local/backtalk6858/jenkins:latest
4. Restart Jenkins via start.sh
5. Trigger E2E smoke test — watch push → webhook → build → docker build stage passes
### 2026-06-24 Sprint Day 5 facts
- id=141 COMPLETE ✅ — Full Jenkins E2E verified: Build → Push → Sandbox Deploy → Smoke Test all pass (build #9)
- Dockerfile fix: replaced docker.io with docker-ce-cli via Docker official apt repo (docker.io on Debian Trixie does NOT install docker binary); rebuilt with --no-cache
- Jenkinsfile fixes required across multiple iterations:
- gitea-credentials (usernamePassword) used for docker login — not secrets-proxy /secret (endpoint doesn't exist)
- Compound commands (&&, cd, $()) don't work — sudo-bridge uses subprocess.run(shlex.split()) with no shell; must use separate bridge calls
- readJSON step not available — pipeline-utility-steps plugin not installed; replaced with python3 shell parsing
- plugins.txt: pipeline-utility-steps added (takes effect on next Jenkins image rebuild)
- server-01 allowlist.json: root:root + chmod 644 ownership (requires manual sudo to update — bootstrap problem)
- New patterns added: docker pull *, docker compose -f * up -d, docker inspect --format=* *, docker ps *
- Promote to Production stage still uses compound command via secrets-proxy /shell — exits non-zero but Jenkins reports SUCCESS (fix deferred)
- Coolify API evaluation: Jenkins handles all deploy work; N8N env var rotation workflows still use Coolify API — will need migration when Coolify retires
- personal_projects id=147: coolify-traefik-retirement (p1), id=148: pam-ntfy-fingerprint-sudo (p3)
- New universal rules: feedback_evaluate_jenkins_hermes_fit.md (Jenkins/Hermes fit check on every automation design + background audit once both deployed)
- Hermes exploration day memory added: schedule full day after id=138 deploys before adding more Hermes-dependent tasks
- SSH sudo requires TTY — bridge bootstrap problem: allowlist.json is root-owned, must use manual sudo once to update it
### 2026-06-24 Sprint Day 6 facts
- id=147 Phase 1 COMPLETE ✅ — standalone Traefik deployed as `coolify-proxy` on `coolify` network
- Config at /opt/appdata/docker/docker-compose/traefik/ (traefik.yml + docker-compose.yml + docker-compose.test.yml)
- Mount: /data/coolify/proxy/:/traefik — certs (acme.json) preserved, dynamic configs unchanged
- All 21 routes live: 17 Docker-label + 4 file-based (jenkins, n8n, sudo-bridge-server01, coolify)
- Cloudflare Tunnel routes confirmed working end-to-end (secrets-proxy /health verified)
- Architecture clarified: Cloudflare Tunnel = encrypted transport, Traefik = HTTP-only (no TLS at origin)
- Legacy setup was wrong: double TLS + port forwarding — current is correct
- fileConfig.yml not needed: TLS irrelevant (Cloudflare terminates), securityHeaders not wired up anywhere
- Vault auth pattern for Claude: AppRole creds at /opt/appdata/docker/docker-compose/vault/approle/, Vault IP 172.16.16.5:8200 (coolify network), use curlimages/curl container on coolify network
- Proxy caller key: Vault secret/data/proxy/callers → claude-code field; use via docker run + Vault login, never print
- New memory rule: feedback_secrets_via_proxy_only.md — always use secrets-proxy /shell with env_secrets, never docker exec env
- server-01 Coolify proxy has zero routes — all routing handled from primary; no standalone Traefik needed on server-01
- Gitea repo for Traefik config: NOT YET CREATED — Phase 2 task
**Next session first task:** id=147 Phase 2 — Coolify retirement
1. Migrate N8N env var rotation workflows from Coolify API → Vault
2. Create Gitea repo for Traefik config (boilerplates entry + repo)
### 2026-06-24 Sprint Day 7 facts
- id=147 Phase 2 IN PROGRESS — N8N workflows migrated, Coolify stopped on primary
- N8N rotation workflow changes (committed e28a3fe in docker-compose repo):
- DB_Password_Rotation: removed Read Vault Coolify Key + 10 Coolify PATCH/Deploy nodes; rewired ALTER → Vault Write directly
- Admin_UI_Password_Rotation: removed Read Vault Coolify node only (coolify_key was loaded but never called Coolify API)
- App_Token_Rotation: removed Read Vault Coolify + 19 Coolify PATCH/Deploy nodes; rewired DELETE JF Old Key → NTFY directly
- All 3 imported to sandbox N8N (192.168.1.90:5679): IDs o8zxdYbY4Y6JRTEy, 1mM1rC9n2HjHPpJf, lDVXyar70e0w3fs9
- Status: NOT YET TESTED in sandbox — services pick up rotated values on next Jenkins redeploy
- Coolify containers stopped on primary: coolify, coolify-realtime, coolify-redis, coolify-db, coolify-sentinel
- All 21 routes verified live post-shutdown
- New behavior rule: feedback_n8n_json_first.md — always edit workflow JSONs in git repo, import to sandbox for testing; never edit in N8N UI
- N8N playbook updated: JSON-first rule added, workflow status reset to NOT YET TESTED, secret/data/coolify removed from Vault paths
- sudo-bridge /exec endpoint confirmed (was using wrong /execute, /run); POST /exec is correct
- Vault paths no longer needed by N8N: secret/data/coolify (rotation workflows no longer call Coolify API)
**Next session first task:** id=147 Phase 2 — remaining steps
1. Remove coolify.yaml from /data/coolify/proxy/dynamic/ (via sudo-bridge POST /exec — needs Vault auth for bridge key)
2. Check server-01 Coolify containers (expected: just coolify-proxy which is already replaced by standalone Traefik)
3. Create Gitea repo for Traefik config + boilerplates entry
4. Verify 90-day rotation workflow dry run in sandbox with Coolify nodes removed
3. Shut down Coolify stack on primary (coolify, coolify-realtime, coolify-redis, coolify-db)
4. Shut down Coolify stack on server-01
5. Verify all routes still live after Coolify gone
### 2026-06-24 Sprint Day 8 facts
- id=147 COMPLETE ✅ — Coolify fully retired on both servers
- coolify.yaml removed from /data/coolify/proxy/dynamic/ via sudo-bridge (new self-removing allowlist cleanup pattern)
- server-01 Coolify containers stopped: coolify-sentinel + coolify-proxy
- Gitea repo created: Backtalk6858/traefik (3 files pushed); boilerplates entry committed d2bbf8b
- All routes verified live post-shutdown
- id=141 COMPLETE ✅ — marked complete in DB
- Phase 2 Security Hardening COMPLETE ✅:
- id=142: bitwarden-bridge POST /items endpoint added (bw create template item → encode → create pattern)
- id=72: secrets-proxy bitwarden:// URI field selection (bitwarden://item#field); fixed bug — /items never returns password; now uses GET /secret?item=&field=
- id=140: zero-trust + security philosophy playbook written to memory (playbook_zero_trust_security.md)
- Phase 3 CI/CD: already complete (id=145 + id=141 done in prior days)
- end-of-session checklist: Obsidian step fixed in feedback_end_of_session_checklist.md (universal) + context-monitor.py time reminders (Logseq → Obsidian)
- Useful new pattern: self-removing sudo-bridge allowlist cleanup — write /tmp script, add tier-1 entry, execute (removes both entries), delete script
- Vault path for sudo-bridge API key: secret/data/sudo-bridge (field: first value, or api_key)
- Vault KV v2: always use secret/data/<path> not secret/<path>
- bitwarden-bridge GET /items: returns id/name/username/revisionDate ONLY — never password/notes; use GET /secret?item=&field= for field reads
- secrets-proxy restart pattern: Coolify injected env vars are lost on manual restart; use /tmp startup script that reads all secrets from Vault (scratchpad: start_secrets_proxy.py)
- NotebookLLM (companion for Obsidian): NOT YET SET UP — no record in DB or memory; needs to be planned
**Next session first task:** Hermes deployment (id=138) — grill-me pending, read reference_open_source_ai_tools.md first
### 2026-06-25 Sprint Day 9 facts
- id=138 Hermes DEPLOYED ✅ — container healthy on server-01, dashboard port 9119, external route hermes.reverseproxyserver.net
- config.yaml: provider=custom, base_url=http://172.17.0.1:11434/v1, model=llama3.1:8b, tool_loop_guardrails.hard_stop_enabled=true
- Bitwarden bridge fixed: BRIDGE_API_KEY, VAULT_TOKEN, BW_CLIENTID, BW_CLIENTSECRET were empty (Coolify env var debt) — regenerated + redeployed via subprocess env injection
- Security enforcement hook v1.1 live on primary: 25 tests passing, hard-blocks sudo bypass + secret exposure
- Vault IP confirmed drifting — always resolve via docker inspect. Fixed ref 172.16.16.30 is stale.
- Jenkins = all deployments (boundary locked in memory); Hermes = all monitoring
- Hermes port 8642 already active at startup — no config change needed
- Skill files: SKILL.md format at /opt/data/skills/<category>/<name>/SKILL.md (NOT skill.yaml)
- Shell toolset for Telegram: add `hermes-cli` to platform_toolsets.telegram in config.yaml (PENDING)
### 2026-06-26 Sprint Day 10 early AM facts
- Background agent prompt discipline established after first agent looped (50+ calls/45 min) vs second agent (8 calls/4 min)
- Required rules for all background agent prompts: --max-turns 15, loop detection instruction, credential map inline, data pre-fetched, defined output format
- Automation audit: 74 items — 35 N8N / 15 Hermes / 13 Hybrid / 8 Jenkins
- required_playbooks column added to automation_ideas, 73 rows backfilled by type with Gitea raw URLs
- Both builder playbooks updated: required_playbooks fetch (Step 2) + human NTFY prompt review gate added
- id=74 Ruflo moved to personal_projects id=151; id=152 infrastructure auth map added
- Hermes user chat_id: 8871022110
### 2026-06-26 Sprint Day 10 afternoon facts
- Infrastructure auth map BUILT + VERIFIED: secret/data/claude/infrastructure-auth-map (3 fields: yaml/json/markdown)
- Background agent wrote it in 2 tool calls/80 seconds
- Vault IP currently 172.16.16.5 (confirm drift — always resolve dynamically)
- All 16 service paths verified present in Vault. Key field corrections vs assumptions:
- bitwarden-bridge: field=BRIDGE_API_KEY (not api_key)
- gitea: fields=[admin_token, admin_username]
- bitwarden (N8N): field=bw_session
- sudo-bridge-db exists: fields=[db, password, user]
- Auth map updated to v2 with all corrections
- Parallel agent strategy confirmed: same total tokens as serial, faster wall-clock, no added cost
- Auth map is now the canonical credential reference — eliminates agent credential-hunting problem
- New personal_projects logged: id=154 (Domain playbooks → Gitea), id=155 (required_playbooks precision pass, blocked by 154), id=156 (SessionStart hook for auth map)
### 2026-06-26 Sprint Day 10 afternoon session 2 facts
- id=156 COMPLETE ✅ — SessionStart hook injects full auth map from Vault at every session start (section 7 added); stale IP 172.16.16.30 fixed to dynamic resolve
- id=152 COMPLETE ✅ — marked completed in DB
- ids 157-161 LOGGED — Jenkins full deployment expansion (service-registry.json, standard Jenkinsfile template, version pinning, standard service jobs, Hermes→Jenkins trigger)
- Jenkinsfile Promote-to-Production bug FIXED — bitwarden-bridge, secrets-proxy, sudo-bridge (compound command → split calls + returncode checking)
- reference_available_tools.md CREATED — tool inventory for all design decisions; loaded at every session
- playbook_background_agent_prompts.md CREATED — mandatory wrap-up JSON + 5 discipline rules + template
- Structural wrap-up design principle added to both builder playbooks (commit 6442e7f)
- TaskCreate enforcement added to end-of-session checklist (10 tracked steps)
- credentials disambiguation added to domain vocabulary (Bitwarden=login, Vault=runtime secrets)
- Jenkins/Hermes full deployment coverage grill-me COMPLETE: N-1 versioning, boilerplates service-registry.json, Hermes→secrets-proxy→Jenkins trigger, conditional sandbox (first deploy or compose change only)
- Automation audit (74 items) now UNBLOCKED — both Jenkins + Hermes deployed
- Primary server: qbittorrent using 8.1GB RAM (abnormal — memory leak suspected), swap 7.6/8.2GB full
### 2026-06-26 Sprint Day 10 session 3 facts
- ids 157-160 COMPLETE ✅ — service-registry.json (33 services), standard Jenkinsfile template, 7 services version-pinned, 24 standard Jenkins jobs created (34 total)
- readJSON → readFile+JsonSlurper fix (Pipeline Utility Steps plugin not installed in Jenkins)
- proxy-caller-key mismatch: Jenkins key not in PROXY_CALLERS env var of secrets-proxy; jenkins added to proxy/callers in Vault; redeploy needed → id=163
- sudo-bridge sandbox IP binding bug: server-01 compose binds 192.168.1.88 (primary IP) → id=164
- Jellyfin Jenkins job failing pending id=163; full pipeline untested end-to-end
- Cookie jar required for Jenkins crumb across urllib requests
- Claude Max upgrade needed — daily token limits hit during background agent sessions
**Next session first tasks (in order):**
1. id=163 — Redeploy secrets-proxy to activate Jenkins caller key (BLOCKER for all Jenkins Promote-to-Production)
2. id=164 — Fix sudo-bridge sandbox compose IP on server-01
3. Re-run jellyfin Jenkins job (verify full pipeline end-to-end after id=163)
4. id=161 — Hermes→Jenkins trigger (needs id=163 first)
5. Hermes config: add hermes-cli to telegram toolset
6. id=154 Domain playbooks → Gitea
7. Build id=24 Agent Builder Agent + id=67 N8N Builder Agent
## Update instructions
Update at the end of every agent-builder session. Keep agent status, key decisions, and prereq checklist current.