docs(claude-config): autonomy/isolation evaluation + live-state findings + proposed decisions

- research/autonomy-isolation-evaluation.md: plan verdict by layer, ranked gaps,
  2026 Docker/Anthropic isolation landscape, proposed VM-boundary + brokers shape
- decisions: two PROPOSED entries (Incus KVM VM boundary; CA re-scope around auto mode)
- synthesis: /recall "misconfig" line corrected (was server-01 offline)
- context.md: live-state alert (control plane down both hosts since 09-07 16:07,
  daemon crash-loop root cause) + next-session plan

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X9iCzmxK2zbb8H1Ld3f8AN
This commit is contained in:
Backtalk6858
2026-09-08 20:37:34 -05:00
parent 97df616f9a
commit d8ec3001a4
5 changed files with 264 additions and 3 deletions
+25
View File
@@ -73,3 +73,28 @@ the LifeOS-vs-Gbrain context-layer overlap it forces.
## Update instructions
Update this file during every session debrief that touches this project — keep "Current state",
"Directory layout", and "Known issues" current after every session.
## NEXT SESSION PLAN (set 2026-09-08 20:40 — user's explicit order, stay on FABLE 5.1 for 1–3; credits expire 2026-09-19, $99.95 left)
1. **Security-infra GAME PLAN** — detailed, buildable plan to get ALL security infrastructure deployed: boot/recovery
layer → `claude-runner` Incus KVM VM on server-01 → secrets-proxy S1/S2 → Agent-Sudo A1/A2/A5 → Hermes outside the
failure domain → CA re-scope (deny-only hook + prod-verb denies). Inputs: `research/autonomy-isolation-evaluation.md`
§5–§6, `agent-builder/GAMEPLAN_agent-sudo_secrets-proxy.md`. Apply feedback_build_the_freeing_capability_first.
2. **Grill-me to FINALIZE the two PROPOSED decisions** in `decisions/DECISIONS.md` (VM boundary + brokers; CA re-scope)
plus gameplan A3 (primary tier-2 → lean (c)). Log the outcomes as DECIDED entries.
3. **If credits remain: AI-harness / LifeOS track** — grill-me FIRST on "what do I want out of an AI harness" (user's
words), then compare LifeOS (costs money; voice feature; fragile Linux) vs free/open-source harnesses (OpenCode,
Aider, Goose, Cline, OpenHands, Hermes-agent…) and decide voice: port LifeOS VoiceServer vs alternative.
Inputs: `research/local-ai-coding-stack-research.md` §3–§6, `reference_open_source_ai_tools`, synthesis Part 3 rows 3/6/7.
4. Then switch to Opus 4.8 for execution (bring control plane up in order, RestartSec on the daemon, Ethernet).
- Not done this session (context limit): personal_projects DB follow-up rows for items 1–3 — create them next session
(read playbook first, dedup check).
## LIVE-STATE ALERT (measured 2026-09-08 20:20 — re-verify before trusting)
- **Control plane DOWN on both hosts since 2026-09-07 16:07** (exit 128, same second, both machines):
Agent-Sudo HTTP, sudo-bridge, bitwarden-bridge, Hermes, Jenkins, N8N prod+sandbox, vault-sandbox.
Media/personal stack is up. `agent-sudo-daemon` busy-looped ~24 h (NRestarts 17.6k/18.9k) because Vault was
down at boot → AppRole login fail → fail-closed exit 3 → 5 s restart; self-healed when Vault returned 17:00.
**First job next session: bring it up in order + find what killed 15 containers simultaneously.**
- Full evaluation + proposed decisions: `research/autonomy-isolation-evaluation.md` (Fable 5.1 pass) and
`decisions/DECISIONS.md` (two PROPOSED entries: VM-boundary + brokers; re-scope CA around auto mode).
+25
View File
@@ -15,6 +15,31 @@ entry links back to the research that drove it, so the reasoning survives.
---
## 2026-09-08 — PROPOSED (not decided): run Claude Code inside an isolation boundary + brokers
- **Decision:** PROPOSED adopt — Claude Code runs in an **Incus KVM VM on server-01** (non-root, no sudo, no
docker.sock, no Vault creds, egress allowlist); all host reach goes through the existing brokers
(Agent-Sudo, secrets-proxy, Jenkins). Inside that boundary, auto mode / `--dangerously-skip-permissions`
is the Anthropic-sanctioned case. Docker Sandboxes rejected for now (unsupported on LMDE 7 / Debian 13;
duplicate VMM beside Incus) — steal its credential-injection + egress-policy ideas.
- **Why:** today the agent shell is root-equivalent on the production host; the July design brokered
actions but never said where the agent runs. The hypervisor already exists (Incus qemu driver, KVM).
- **Research:** [../research/autonomy-isolation-evaluation.md](../research/autonomy-isolation-evaluation.md) §3–§4
- **Implementation:** pending — needs grill-me + a CA design amendment (CA-D11) first.
- **Revisit when:** a host moves to Ubuntu 24.04+ (Docker Sandboxes becomes supported), or Incus VM
overhead proves too high (fallback: `@anthropic-ai/sandbox-runtime`).
## 2026-09-08 — PROPOSED (not decided): re-scope Constrained Autonomy around auto mode
- **Decision:** PROPOSED — keep `security-enforcement.py` as a **deny-only + training-log** layer; **drop**
CA-P1-full / D2 structural classifier / CA-P4 learned-allowlist promotion (auto mode's classifier now does
the judgment half, default on Pro); keep CA-P3 (secrets path) and CA-P5 (Hermes recovery ladder). Add
explicit deterministic denies for irreversible prod verbs (force-push, prod `compose down`, DB migrations)
to cover auto mode's measured 17% miss on overeager actions.
- **Why:** the prompt-theater tax CA-P1 attacked is solved upstream; regex should do what a model classifier
can't (unbypassable denies), not compete with it.
- **Research:** [../research/autonomy-isolation-evaluation.md](../research/autonomy-isolation-evaluation.md) §3a, §5 #5–6
- **Implementation:** pending.
- **Revisit when:** auto mode is removed from Pro, or its false-negative rate is published as materially better.
## 2026-09-08 — Establish claude-config as the research + config home
- **Decision:** adopt — this directory is where Claude/Claude-Code research, decisions, and
config/implementation work now live.
+3
View File
@@ -7,6 +7,7 @@ re-verify versions, repos, and Linux support before acting on anything.
| Topic | File | Status | Key open decision |
|---|---|---|---|
| **Infrastructure synthesis** — the whole landscape + reconciliation of the new research against decisions already made | [infrastructure-synthesis.md](infrastructure-synthesis.md) | Synthesized 2026-09-08 | **Confirm the July→Sept reality** (Max upgrade / Voice-Chat / Tailscale-Twingate / Agent-Sudo) before acting on anything. |
| **Autonomy + isolation evaluation** — does the plan work, gaps, and the 2026 Docker/Anthropic isolation landscape (auto mode, Bash sandbox, sandbox-runtime, Docker Sandboxes microVMs) | [autonomy-isolation-evaluation.md](autonomy-isolation-evaluation.md) | Evaluated 2026-09-08 (Fable 5.1) | **Adopt the VM-boundary + brokers shape?** (Incus KVM VM on server-01 running Claude Code; Agent-Sudo/secrets-proxy/Jenkins as the only host reach). Also: what killed the control plane at 16:07 on 2026-09-07. |
| Local AI coding stack (inference engine, coding harness, LifeOS, voice, skill porting) | [local-ai-coding-stack-research.md](local-ai-coding-stack-research.md) | Surface-level, in progress | **Goal framing:** RESOLVED by the existing vision = cost-reduction + tooling-independence, Claude stays the brain (NOT fully-local). See synthesis Part 3. |
### Where the prior infrastructure research lives (memory corpus)
@@ -20,6 +21,8 @@ Not duplicated here — cited in [infrastructure-synthesis.md](infrastructure-sy
## Cross-cutting open decisions
Pulled up from the individual briefs so they don't get buried:
0. **Isolation boundary** — where does Claude Code itself run? Proposed: Incus KVM VM on server-01, non-root, egress allowlist, brokers only. Decides whether `--dangerously-skip-permissions`/auto mode is safe. (autonomy-isolation-evaluation.md §4)
1. **The goal** — cost / independence / fully-local. Governs every other choice in the local
stack. (from local-ai-coding-stack-research.md §0)
2. **8 GB VRAM ceiling** — RTX 2060 Super caps a fully-local setup at ~7–8B @ Q4. Which coding
@@ -0,0 +1,210 @@
# Autonomy + Isolation — evaluation of the infra plan, and the 2026 Docker/Anthropic isolation landscape
**Written:** 2026-09-08 (Fable 5.1 pass) · **Inputs:** live checks on both hosts tonight, the July designs
(`agent-builder/constrained_autonomy_design_decisions.md`, `agent_sudo_design_decisions.md`,
`GAMEPLAN_agent-sudo_secrets-proxy.md`), [infrastructure-synthesis.md](infrastructure-synthesis.md), and
current Docker + Claude Code docs (cited at the bottom). **Re-verify anything version-shaped before acting.**
> **User's question:** "Does the infrastructure plan look like it will work? What are the gaps? And can
> Docker's new AI-agent isolation (agents in isolated VMs) make it safe to run Claude Code privileged, so
> I never have to grant permission?"
>
> **Short answer:** the *design* is sound and — this is the notable part — it independently arrived at the
> architecture Anthropic and Docker now recommend (broker consequential actions, don't gate them; put the
> agent behind a real boundary). But the *deployment* is currently down and does not survive a reboot, and
> the July Constrained-Autonomy plan predates two upstream changes (Claude Code **auto mode**, now the default
> on Pro, and **microVM sandboxes**) that let us **delete ~half of what we planned to build**. Yes, "Claude
> Code that just does it" is achievable safely — but the safety comes from an **isolation boundary + broker**,
> not from any permission flag. The boundary we need already exists on server-01 (Incus + KVM).
---
## 1. Live state tonight (measured, not remembered)
Both hosts rebooted during the move (primary 2026-09-07 15:42, server-01 15:55). Then at **16:07–16:08 on
2026-09-07, every control-plane container on BOTH hosts exited (code 128) within 20 seconds of each other
and never came back**, despite `restart: unless-stopped`. The media/personal stack (Jellyfin, Nextcloud,
Grafana, Ollama, Vault, Postgres, Gitea, cloudflared…) is up.
| Component | primary (192.168.1.88) | server-01 (192.168.1.90) |
|---|---|---|
| Agent-Sudo HTTP (`/health`) | **DOWN** (container exited 128) | **DOWN** (container exited 128) |
| Agent-Sudo host daemon (unit) | active, **NRestarts=17,598** | active, **NRestarts=18,882** |
| sudo-bridge | exited 128 | never started (`created`) |
| bitwarden-bridge | exited 128 | sandbox copy exited 128 |
| secrets-proxy | exited 0 — **since 2026-07-09** (deliberate) | — |
| Hermes (the monitor) | — | **exited 128** |
| Jenkins (all deploys) | — | exited 143 |
| N8N prod + sandbox | — | **both exited 128** |
| Vault | up (restarted ~17:00 today) | sandbox copy exited 128 |
| Incus sandboxes | n/a (no Incus) | `sandbox-primary` + `sandbox-server01` RUNNING |
**Daemon crash-loop — root cause found (journal):** on boot the daemon's SUDO.md signature gate does an
AppRole login; Vault wasn't up → `REFUSE TO LOAD … AppRole login failed` → exit 3 (fail-closed, correct) →
systemd restarts in 5 s → repeat. 17.6k restarts × 5 s ≈ 24 h = the whole window Vault was down. Fail-closed is
right; **restarting every 5 s with no backoff and no `After=`/wait-for-Vault is the bug.** It self-healed the
moment Vault returned (17:00 today), which is why it looks "active" now while its HTTP front-end is still dead.
**What this proves about the plan:** the security control plane has **no post-reboot recovery path**, and the
thing designed to notice (Hermes) died in the same event. That is Gap #1 below, and it is not a design flaw
of any single component — it's a missing *boot-order + recovery* layer.
---
## 2. Will the plan work? Verdict by layer
| Layer | Design verdict | Deployment verdict |
|---|---|---|
| **Doctrine** — constrain capability, don't gate access; EXECUTE vs EXPAND-allowlist; tier-4 human-only by design | ✅ Sound. Matches what Anthropic's own auto-mode team concluded ("prompts you rubber-stamp aren't safety") and what Docker's sandbox model assumes. | n/a |
| **Agent-Sudo** (broker for root: tiers, sandbox-test, rollback, breaker, signed SUDO.md, audit → training data) | ✅ Sound and *more* rigorous than anything upstream offers (transit-signed policy, replayed breaker log, tier-4 self-cannibalization guard). | ⚠️ HTTP layer down both hosts; tier-1 undo still broken (A1); tiers 2/3 = 503 stubs; primary cutover #146 undone by the outage anyway. |
| **secrets-proxy** (broker for secrets) | ✅ Right idea — exactly the "credential injection outside the boundary" pattern Docker Sandboxes and Claude Code on the web use. | ❌ Down since July 9; scope never investigated (S1). Secrets rule has no proxy → file-passing workarounds. |
| **Constrained Autonomy hook** (CA-P1a/b: regex fast-path + catastrophic deny + secret-exposure blocks) | ✅ for the *deny* half. ⚠️ the *allow/classify* half is now **superseded by auto mode** (see §3). Don't build CA-P1 full / D2 structural classifier. | ✅ Live (hook v2.1). Still the only deterministic deny on the box. |
| **Jenkins = deploy / Hermes = monitor** split | ✅ Clean boundary. | ❌ Both down. Hermes-Phase-2 recovery ladder (CA-D6) can't exist while Hermes shares the failure domain. |
| **Networking** (Tailscale me / Twingate others, host-terminated, outbound-only) | ✅ Locked decisions still correct. IPs survived the new router (.88/.90). | ⏸ Not started; cloudflared still the inbound path. Ethernet ~09-09 is the natural predecessor. |
| **Memory/recall** | ✅ pull-not-push is right. | ⚠️ `/recall` embeds via server-01's Ollama over flaky WiFi; primary's Ollama (up 29 h) is the obvious local fallback. |
| **Model/billing** | ✅ Opus 4.8/medium settled; Fable via promo credits; no API key. | ✅ |
**Bottom line:** nothing needs to be un-decided. Three things need to change: (1) add a boot/recovery layer,
(2) re-scope Constrained Autonomy around auto mode instead of a homemade classifier, (3) add the isolation
boundary the July design assumed but never named — and *that* is what makes "just do it" safe.
---
## 3. The upstream changes since July that change our build list
### 3a. Claude Code **auto mode** (default on Pro/Max/Team — we are in it right now)
- A second model (a Sonnet-class classifier) reviews each action instead of the human. Two stages: a fast
single-token filter tuned to over-block, then chain-of-thought only on flagged actions. It deliberately
strips assistant text and tool results so the agent can't talk it into approvals. ~20 default rules in four
groups: destructive/exfiltration, security-degradation (disabling logging, persistence, editing permission
config), trust-boundary (untrusted code, credential-store scans), and bypass-safeguards (prod pushes/deploys).
- Anthropic's own numbers: **0.4% false-positive**, but **17% false-negative on confirmed overeager actions**
(the classifier "finds approval-shaped evidence" and can't judge blast radius from vague consent). They say
plainly it is **a per-action control, not an isolation boundary**, and "not a replacement for careful review
on critical infrastructure."
- **Consequence for CA:** the prompt-theater tax that CA-P1a/P1b attacked is now solved upstream. Keep the hook
for what a model classifier is bad at — **deterministic, unbypassable denies** (tier-4 verbs, secret-path
reads, `$VAR` echo, token patterns) — and **retire the plan to grow it into a full effect-classifier (D2,
CA-P1 full, learned allowlist promotion)**. Hooks and deny rules still fire in every mode, including
`bypassPermissions`, so the hook stays load-bearing.
### 3b. Claude Code **sandboxed Bash tool** (bubblewrap + socat on Linux; `sandbox.enabled`)
- OS-enforced filesystem (write only cwd + `$TMPDIR`, `denyRead` for secrets) and network (domain allowlist via
proxy) isolation for Bash *and its children*; "auto-allow" mode runs sandboxed commands with no prompt.
- **Does not cover** file tools, MCP servers, or hooks (they run on the host), and the docs flag the exact
thing that makes it a poor fit for us: **allowing `/var/run/docker.sock` "effectively grants access to the
host"** — and our whole workflow on primary is docker/ssh/vault from Bash. Also `socat` is not installed on
primary. Verdict: **useful inside a VM as a second layer; not our primary boundary.**
### 3c. **sandbox-runtime** (`@anthropic-ai/sandbox-runtime`, beta) — wraps the *whole* Claude Code process
(hooks + MCP included) in the same bubblewrap boundary. No Docker needed. Weaker than a VM (shared kernel) and
the Linux deny-list is built once at launch. A credible middle step if a VM is ever too heavy.
### 3d. **Docker Sandboxes (`sbx`)** — microVM per agent
- Free, standalone, **does not require Docker Desktop or Docker Engine**. Each sandbox = its own VM, own Docker
daemon, own filesystem/network; egress goes through a host-side proxy that enforces network policy and does
**credential injection** (secrets never live inside the VM); workspace via virtiofs mount or read-only clone.
`sbx run claude` launches Claude Code inside it. This is Docker's answer to "run agents unsupervised."
- **Platform support is the catch:** Linux = **Ubuntu 24.04+ only** (KVM + `kvm` group), and Docker explicitly
says it does **not test or support Ubuntu derivatives (Mint, Pop!_OS)**. Primary is **LMDE 7** (Debian-based,
not even an Ubuntu derivative) and server-01 is **Debian 13**. It may run; it is **unsupported on both hosts.**
- Also relevant: Docker **Hardened Images** are now free/Apache-2.0 (harden our base images), and the **MCP
Gateway** runs each MCP server in its own restricted container with credential injection — a fit for the
Hermes/N8N tool surface later. ECI (Sysbox user-namespaces) is Docker-Desktop-only — not applicable here.
### 3e. Anthropic's explicit guidance (docs, 2026)
- `--dangerously-skip-permissions`: **always inside a container, VM, or sandbox-runtime, as a non-root user**
(Claude Code refuses to start with it as root). The devcontainer with default-deny iptables is the reference.
- Auto mode: no boundary *required*, but "an isolation boundary still adds defense in depth for unattended runs."
- "Sandbox isolation reduces the impact of a breach, it does not eliminate risk. Any approach that allows
network egress can still leak data the agent can read." (Domain fronting through broad allowlists is called
out; TLS is not inspected by default.)
---
## 4. The answer to "can it just do it, safely?" — yes, with this shape
The July design already split the world correctly: Claude Code proposes; **brokers** (Agent-Sudo for root,
secrets-proxy for secrets, Jenkins for deploys) execute under constraint. What it never stated is *where Claude
Code itself runs*. Today it runs as `administrator` on the **production host** with the docker group (= root via
the socket) and sudo — so the only things between a prompt-injected agent and the box are a regex hook and a
17%-miss classifier. That, not the permission mode, is the real gap.
**Target shape (proposed, not decided):**
```
┌─ server-01 ───────────────────────────────────────────────────────────┐
│ Incus KVM VM "claude-runner" (Debian 13, non-root user, no sudo, │
│ no docker.sock, no Vault creds, egress allowlist at the VM NIC) │
│ ├─ Claude Code in auto mode (or --dangerously-skip-permissions) │
│ ├─ security-enforcement.py hook (deterministic denies, unchanged) │
│ ├─ optional: built-in Bash sandbox as a 2nd layer │
│ └─ repos mounted (virtiofs) from the host │
│ │
│ consequential host actions ──HTTP──▶ Agent-Sudo :8082/:8084 (broker)│
│ secrets ────────────────────HTTP──▶ secrets-proxy /shell (broker) │
│ deploys ────────────────────HTTP──▶ Jenkins (broker) │
└───────────────────────────────────────────────────────────────────────┘
```
- The VM is the **isolation boundary** (own kernel — the strongest tier in Anthropic's table). Everything the
brokers refuse is physically impossible from inside it. Now `--dangerously-skip-permissions` is exactly the
case Anthropic says it's for, and auto mode becomes a bonus layer rather than the last line.
- **We already have the hypervisor:** Incus 6.0.4 on server-01 reports `driver: lxc | qemu (10.0.11)`,
`/dev/kvm` present, VM images available, 24 cores / 27 GB free. **Zero new software.** Primary has KVM too but
no Incus (the documented FORWARD-DROP pain, A3) — run the VM on server-01 and reach primary over the brokers,
which is what the brokers are for. This also *dissolves* gameplan A3 ("what does tier 2 mean on primary"): the
question was hard because Claude ran *on* primary; if Claude runs in a VM on server-01, primary is just another
broker target and can stay tier-2-less (option c) with no loss.
- **Why not Docker Sandboxes here:** unsupported on LMDE/Debian, and it would add a second VMM beside Incus for
the same result. Revisit if we ever standardize an Ubuntu host. Steal its ideas instead: **credential
injection at the boundary (= secrets-proxy)** and **egress policy at the VM** (Incus network ACLs / nftables
on the bridge — allow api.anthropic.com, claude.ai, platform.claude.com, gitea.local, the three brokers; deny
the rest).
- **What it costs:** Claude Code's *own* reach shrinks to the repos + brokers, so anything we do today by
"just running docker/ssh from the primary shell" has to have a broker path. That is the entire point — and it's
the forcing function that finally finishes Agent-Sudo tiers 1–3 and secrets-proxy, because the VM makes their
absence *felt* instead of worked around.
---
## 5. Gaps (ranked) and what closes each
| # | Gap | Severity | Closes it |
|---|---|---|---|
| 1 | **No reboot survivability / recovery layer.** Control plane died 25 min after boot on both hosts and stayed dead; daemon busy-looped 24 h; Hermes died with it. | 🔴 | A `boot-recovery` step: systemd `After=`/`Wants=` on Vault reachability + `RestartSec=30`/`StartLimitBurst` for the daemon; a post-boot health job (Jenkins or a systemd timer) that `compose up`s the control plane in order (Vault → bridges → Agent-Sudo → secrets-proxy → Hermes) and NTFYs; Hermes must run *outside* the failure domain it watches (host unit or separate host). **Also: find out what killed 15 containers at 16:07:53 on both hosts simultaneously** (exit 128 in the same second across two machines smells like a network/IPAM event when the WiFi came up, or a `docker` daemon restart). |
| 2 | **No isolation boundary around Claude Code**; agent shell = root-equivalent on the production host. | 🔴 | §4 VM on server-01 via Incus. |
| 3 | secrets-proxy down since July; scope unknown (S1 never done). | 🔴 | S1 investigate → S2 finish. Becomes mandatory once the VM exists (only path to secrets). |
| 4 | Agent-Sudo tier-1 undo broken (A1); tiers 2/3 stubs; #146 cutover. | 🟡 | Gameplan A1 → A2 → A5, unchanged. |
| 5 | CA plan over-scoped vs auto mode (D2 classifier, learned allowlist, CA-P4 promotion). | 🟡 | Re-scope: hook = deny-only + training log; drop CA-P1 full/CA-P4 promotion; keep CA-P3 (secrets path) and CA-P5 (Hermes recovery). |
| 6 | Auto mode's 17% miss on overeager actions has no compensating control on **irreversible prod actions** (git force-push, DB migrations, `compose down` on prod). | 🟡 | Hook: add explicit tier-3/4 denies for those verbs on prod paths (deny rules apply in every mode); Jenkins is the only path to deploy. |
| 7 | Tailscale/Twingate not started; cloudflared still inbound. | 🟡 | Resume preflight after Ethernet; unchanged decisions. |
| 8 | `/recall` single-homed on server-01 Ollama over WiFi. | 🟢 | Fall back to primary's `localhost:11434` (both have `nomic-embed-text`). |
| 9 | Coolify-era N8N workflows still call a dead API (~7). | 🟢 | Jenkins migration, when N8N is back. |
| 10 | Docker socket exposure inside any future Bash-sandbox config (`allowUnixSockets`). | 🟢 | Never allow it in the VM; brokers only. |
---
## 6. Proposed next actions (in order — pending user decisions in DECISIONS.md)
1. **Tonight/tomorrow (Opus is fine):** bring the control plane back up in order and capture *why* it died
(the 16:07 event). Add `RestartSec`/backoff to `agent-sudo-daemon.service` on both hosts. Wire Ethernet.
2. **Decide** (grill-me): adopt the *VM-boundary + brokers* shape (§4)? If yes → a CA design amendment
(CA-D11 "runner boundary") and re-scope CA-P1/P4 as above.
3. **Build order (revised, freeing-capability first):** boot-recovery layer → `claude-runner` VM on server-01
(Incus, non-root, egress allowlist, repos mounted) → run this very workflow from inside it in auto mode →
let the brokers' gaps surface → secrets-proxy S1/S2 → Agent-Sudo A1/A2/A5 → Hermes outside the failure domain.
4. Optional second layer inside the VM: install `socat`, enable `sandbox.enabled` with `denyRead` on secret
paths and a tight `allowedDomains`.
---
## Sources
- Anthropic — [How we built Claude Code auto mode](https://www.anthropic.com/engineering/claude-code-auto-mode)
- Claude Code docs — [Permission modes](https://code.claude.com/docs/en/permission-modes) ·
[Sandboxing (Bash sandbox)](https://code.claude.com/docs/en/sandboxing) ·
[Sandbox environments (comparison)](https://code.claude.com/docs/en/sandbox-environments)
- Docker — [Docker Sandboxes](https://docs.docker.com/ai/sandboxes/) · [Install/requirements](https://docs.docker.com/ai/sandboxes/install/) ·
[Architecture](https://docs.docker.com/ai/sandboxes/architecture/) ·
[Blog: run Claude Code unsupervised but safely](https://www.docker.com/blog/docker-sandboxes-run-claude-code-and-other-coding-agents-unsupervised-but-safely/)
- Background — [What's new in Docker 2026 (Sandboxes, Hardened Images, MCP Gateway)](https://collabnix.com/whats-new-in-docker-in-2026-sandboxes-hardened-images-and-the-ai-native-container-platform/) ·
[Your container is not a sandbox: microVM isolation in 2026](https://emirb.github.io/blog/microvm-2026/)
@@ -87,9 +87,7 @@
Obsidian) as Track 2; **Obsidian = source of truth, Hermes = curator.** Companion tools: Graphify,
Claudian. The load-bearing lesson: the **lint/curation pass** decides whether the system compounds
or rots. **All of this is BLOCKED on the Max upgrade.**
- ⚙️ **Operational finding today:** `/recall`'s semantic layer errors with *"No route to host"* on its
Ollama embed endpoint, even though **localhost:11434 is UP**. The recall script points at an
unreachable host → misconfig. **In-scope fix for this claude-config directory.**
- ⚙️ ~~**Operational finding today:** `/recall` errors "No route to host" … misconfig.~~ **RESOLVED same day: NOT a misconfig** — server-01 (which serves the embed endpoint) was offline after the move; `/recall` worked once it was back on WiFi. Remaining softness: single-homed on server-01 over WiFi; primary's Ollama is the natural fallback. See [autonomy-isolation-evaluation.md](autonomy-isolation-evaluation.md) §5 #8.
### Hardware roadmap (`project_ai_infrastructure_vision`)
- Now: LMDE 7 desktop + **RTX 2060 Super, 8GB (Turing SM 7.5)**.