@@ -7,6 +7,7 @@ re-verify versions, repos, and Linux support before acting on anything.
| Topic | File | Status | Key open decision |
|---|---|---|---|
| **Infrastructure synthesis** — the whole landscape + reconciliation of the new research against decisions already made | [infrastructure-synthesis.md](infrastructure-synthesis.md) | Synthesized 2026-09-08 | **Confirm the July→Sept reality** (Max upgrade / Voice-Chat / Tailscale-Twingate / Agent-Sudo) before acting on anything. |
| **Autonomy + isolation evaluation** — does the plan work, gaps, and the 2026 Docker/Anthropic isolation landscape (auto mode, Bash sandbox, sandbox-runtime, Docker Sandboxes microVMs) | [autonomy-isolation-evaluation.md](autonomy-isolation-evaluation.md) | Evaluated 2026-09-08 (Fable 5.1) | **Adopt the VM-boundary + brokers shape?** (Incus KVM VM on server-01 running Claude Code; Agent-Sudo/secrets-proxy/Jenkins as the only host reach). Also: what killed the control plane at 16:07 on 2026-09-07. |
| Local AI coding stack (inference engine, coding harness, LifeOS, voice, skill porting) | [local-ai-coding-stack-research.md](local-ai-coding-stack-research.md) | Surface-level, in progress | **Goal framing:** RESOLVED by the existing vision = cost-reduction + tooling-independence, Claude stays the brain (NOT fully-local). See synthesis Part 3. |
### Where the prior infrastructure research lives (memory corpus)
@@ -20,6 +21,8 @@ Not duplicated here — cited in [infrastructure-synthesis.md](infrastructure-sy
## Cross-cutting open decisions
Pulled up from the individual briefs so they don't get buried:
0.**Isolation boundary** — where does Claude Code itself run? Proposed: Incus KVM VM on server-01, non-root, egress allowlist, brokers only. Decides whether `--dangerously-skip-permissions`/auto mode is safe. (autonomy-isolation-evaluation.md §4)
1.**The goal** — cost / independence / fully-local. Governs every other choice in the local
**Daemon crash-loop — root cause found (journal):** on boot the daemon's SUDO.md signature gate does an
AppRole login; Vault wasn't up → `REFUSE TO LOAD … AppRole login failed` → exit 3 (fail-closed, correct) →
systemd restarts in 5 s → repeat. 17.6k restarts × 5 s ≈ 24 h = the whole window Vault was down. Fail-closed is
right; **restarting every 5 s with no backoff and no `After=`/wait-for-Vault is the bug.** It self-healed the
moment Vault returned (17:00 today), which is why it looks "active" now while its HTTP front-end is still dead.
**What this proves about the plan:** the security control plane has **no post-reboot recovery path**, and the
thing designed to notice (Hermes) died in the same event. That is Gap #1 below, and it is not a design flaw
of any single component — it's a missing *boot-order + recovery* layer.
---
## 2. Will the plan work? Verdict by layer
| Layer | Design verdict | Deployment verdict |
|---|---|---|
| **Doctrine** — constrain capability, don't gate access; EXECUTE vs EXPAND-allowlist; tier-4 human-only by design | ✅ Sound. Matches what Anthropic's own auto-mode team concluded ("prompts you rubber-stamp aren't safety") and what Docker's sandbox model assumes. | n/a |
| **Agent-Sudo** (broker for root: tiers, sandbox-test, rollback, breaker, signed SUDO.md, audit → training data) | ✅ Sound and *more* rigorous than anything upstream offers (transit-signed policy, replayed breaker log, tier-4 self-cannibalization guard). | ⚠️ HTTP layer down both hosts; tier-1 undo still broken (A1); tiers 2/3 = 503 stubs; primary cutover #146 undone by the outage anyway. |
| **secrets-proxy** (broker for secrets) | ✅ Right idea — exactly the "credential injection outside the boundary" pattern Docker Sandboxes and Claude Code on the web use. | ❌ Down since July 9; scope never investigated (S1). Secrets rule has no proxy → file-passing workarounds. |
| **Constrained Autonomy hook** (CA-P1a/b: regex fast-path + catastrophic deny + secret-exposure blocks) | ✅ for the *deny* half. ⚠️ the *allow/classify* half is now **superseded by auto mode** (see §3). Don't build CA-P1 full / D2 structural classifier. | ✅ Live (hook v2.1). Still the only deterministic deny on the box. |
| **Jenkins = deploy / Hermes = monitor** split | ✅ Clean boundary. | ❌ Both down. Hermes-Phase-2 recovery ladder (CA-D6) can't exist while Hermes shares the failure domain. |
| **Networking** (Tailscale me / Twingate others, host-terminated, outbound-only) | ✅ Locked decisions still correct. IPs survived the new router (.88/.90). | ⏸ Not started; cloudflared still the inbound path. Ethernet ~09-09 is the natural predecessor. |
| **Memory/recall** | ✅ pull-not-push is right. | ⚠️ `/recall` embeds via server-01's Ollama over flaky WiFi; primary's Ollama (up 29 h) is the obvious local fallback. |
| **Model/billing** | ✅ Opus 4.8/medium settled; Fable via promo credits; no API key. | ✅ |
**Bottom line:** nothing needs to be un-decided. Three things need to change: (1) add a boot/recovery layer,
(2) re-scope Constrained Autonomy around auto mode instead of a homemade classifier, (3) add the isolation
boundary the July design assumed but never named — and *that* is what makes "just do it" safe.
---
## 3. The upstream changes since July that change our build list
### 3a. Claude Code **auto mode** (default on Pro/Max/Team — we are in it right now)
- A second model (a Sonnet-class classifier) reviews each action instead of the human. Two stages: a fast
single-token filter tuned to over-block, then chain-of-thought only on flagged actions. It deliberately
strips assistant text and tool results so the agent can't talk it into approvals. ~20 default rules in four
- The VM is the **isolation boundary** (own kernel — the strongest tier in Anthropic's table). Everything the
brokers refuse is physically impossible from inside it. Now `--dangerously-skip-permissions` is exactly the
case Anthropic says it's for, and auto mode becomes a bonus layer rather than the last line.
- **We already have the hypervisor:** Incus 6.0.4 on server-01 reports `driver: lxc | qemu (10.0.11)`,
`/dev/kvm` present, VM images available, 24 cores / 27 GB free. **Zero new software.** Primary has KVM too but
no Incus (the documented FORWARD-DROP pain, A3) — run the VM on server-01 and reach primary over the brokers,
which is what the brokers are for. This also *dissolves* gameplan A3 ("what does tier 2 mean on primary"): the
question was hard because Claude ran *on* primary; if Claude runs in a VM on server-01, primary is just another
broker target and can stay tier-2-less (option c) with no loss.
- **Why not Docker Sandboxes here:** unsupported on LMDE/Debian, and it would add a second VMM beside Incus for
the same result. Revisit if we ever standardize an Ubuntu host. Steal its ideas instead: **credential
injection at the boundary (= secrets-proxy)** and **egress policy at the VM** (Incus network ACLs / nftables
on the bridge — allow api.anthropic.com, claude.ai, platform.claude.com, gitea.local, the three brokers; deny
the rest).
- **What it costs:** Claude Code's *own* reach shrinks to the repos + brokers, so anything we do today by
"just running docker/ssh from the primary shell" has to have a broker path. That is the entire point — and it's
the forcing function that finally finishes Agent-Sudo tiers 1–3 and secrets-proxy, because the VM makes their
absence *felt* instead of worked around.
---
## 5. Gaps (ranked) and what closes each
| # | Gap | Severity | Closes it |
|---|---|---|---|
| 1 | **No reboot survivability / recovery layer.** Control plane died 25 min after boot on both hosts and stayed dead; daemon busy-looped 24 h; Hermes died with it. | 🔴 | A `boot-recovery` step: systemd `After=`/`Wants=` on Vault reachability + `RestartSec=30`/`StartLimitBurst` for the daemon; a post-boot health job (Jenkins or a systemd timer) that `compose up`s the control plane in order (Vault → bridges → Agent-Sudo → secrets-proxy → Hermes) and NTFYs; Hermes must run *outside* the failure domain it watches (host unit or separate host). **Also: find out what killed 15 containers at 16:07:53 on both hosts simultaneously** (exit 128 in the same second across two machines smells like a network/IPAM event when the WiFi came up, or a `docker` daemon restart). |
| 2 | **No isolation boundary around Claude Code**; agent shell = root-equivalent on the production host. | 🔴 | §4 VM on server-01 via Incus. |
| 3 | secrets-proxy down since July; scope unknown (S1 never done). | 🔴 | S1 investigate → S2 finish. Becomes mandatory once the VM exists (only path to secrets). |
| 5 | CA plan over-scoped vs auto mode (D2 classifier, learned allowlist, CA-P4 promotion). | 🟡 | Re-scope: hook = deny-only + training log; drop CA-P1 full/CA-P4 promotion; keep CA-P3 (secrets path) and CA-P5 (Hermes recovery). |
| 6 | Auto mode's 17% miss on overeager actions has no compensating control on **irreversible prod actions** (git force-push, DB migrations, `compose down` on prod). | 🟡 | Hook: add explicit tier-3/4 denies for those verbs on prod paths (deny rules apply in every mode); Jenkins is the only path to deploy. |
| 7 | Tailscale/Twingate not started; cloudflared still inbound. | 🟡 | Resume preflight after Ethernet; unchanged decisions. |
| 8 | `/recall` single-homed on server-01 Ollama over WiFi. | 🟢 | Fall back to primary's `localhost:11434` (both have `nomic-embed-text`). |
| 9 | Coolify-era N8N workflows still call a dead API (~7). | 🟢 | Jenkins migration, when N8N is back. |
| 10 | Docker socket exposure inside any future Bash-sandbox config (`allowUnixSockets`). | 🟢 | Never allow it in the VM; brokers only. |
---
## 6. Proposed next actions (in order — pending user decisions in DECISIONS.md)
1.**Tonight/tomorrow (Opus is fine):** bring the control plane back up in order and capture *why* it died
(the 16:07 event). Add `RestartSec`/backoff to `agent-sudo-daemon.service` on both hosts. Wire Ethernet.
2.**Decide** (grill-me): adopt the *VM-boundary + brokers* shape (§4)? If yes → a CA design amendment
(CA-D11 "runner boundary") and re-scope CA-P1/P4 as above.
3.**Build order (revised, freeing-capability first):** boot-recovery layer → `claude-runner` VM on server-01
(Incus, non-root, egress allowlist, repos mounted) → run this very workflow from inside it in auto mode →
let the brokers' gaps surface → secrets-proxy S1/S2 → Agent-Sudo A1/A2/A5 → Hermes outside the failure domain.
4. Optional second layer inside the VM: install `socat`, enable `sandbox.enabled` with `denyRead` on secret
paths and a tight `allowedDomains`.
---
## Sources
- Anthropic — [How we built Claude Code auto mode](https://www.anthropic.com/engineering/claude-code-auto-mode)
- Claude Code docs — [Permission modes](https://code.claude.com/docs/en/permission-modes) ·
[Blog: run Claude Code unsupervised but safely](https://www.docker.com/blog/docker-sandboxes-run-claude-code-and-other-coding-agents-unsupervised-but-safely/)
- Background — [What's new in Docker 2026 (Sandboxes, Hardened Images, MCP Gateway)](https://collabnix.com/whats-new-in-docker-in-2026-sandboxes-hardened-images-and-the-ai-native-container-platform/) ·
[Your container is not a sandbox: microVM isolation in 2026](https://emirb.github.io/blog/microvm-2026/)
Obsidian) as Track 2; **Obsidian = source of truth, Hermes = curator.** Companion tools: Graphify,
Claudian. The load-bearing lesson: the **lint/curation pass** decides whether the system compounds
or rots. **All of this is BLOCKED on the Max upgrade.**
- ⚙️ **Operational finding today:**`/recall`'s semantic layer errors with *"No route to host"* on its
Ollama embed endpoint, even though **localhost:11434 is UP**. The recall script points at an
unreachable host → misconfig. **In-scope fix for this claude-config directory.**
- ⚙️ ~~**Operational finding today:** `/recall` errors "No route to host" … misconfig.~~**RESOLVED same day: NOT a misconfig** — server-01 (which serves the embed endpoint) was offline after the move; `/recall` worked once it was back on WiFi. Remaining softness: single-homed on server-01 over WiFi; primary's Ollama is the natural fallback. See [autonomy-isolation-evaluation.md](autonomy-isolation-evaluation.md) §5 #8.
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.