feat(netbird): N1 prompt v2 (containerized) + sandbox run handoffs (#258)

N1 prompt rewritten to a fully-containerized two-peer NET_ADMIN relay test
(no host install; teardown-to-baseline) with the owner-command relay
(pause-for-sudo/browser) protocol wired in. Design doc §7 gains the v2 +
v2b live-run handoffs: control plane + OIDC verified, relay is
architecturally non-cloudflare (rel://relay:33080), Authelia OIDC solved
via CORS + management UseIDToken; remaining gap = NetBird instance-setup
(name-claim) hang.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
Backtalk6858
2026-09-18 18:12:12 -05:00
parent 1569cccae4
commit a01844ccc2
2 changed files with 185 additions and 66 deletions
@@ -1,16 +1,21 @@
# N1 — NetBird server-01 Sandbox Validation (background agent)
# N1 — NetBird server-01 Sandbox Validation (background agent) — v2 (containerized)
> Executes Phase N1 of `playbook_netbird_phases.md`. Authoritative design =
> `research/netbird-design_2026-09-17.md` (§4 runbook). This prompt is self-contained; do NOT
> assume you can read those — every procedure-critical fact is inlined below.
> **This is a THROWAWAY sandbox. Nothing here is production. Tear it all down at the end.**
>
> **v2 rewrite (2026-09-18):** the v1 prompt planned a HOST-level netbird/WireGuard install and a
> "back to Ollama + Obsidian only" teardown. BOTH were wrong and burned a spawn:
> server-01 is a **shared host** (it runs offloaded prod services — see baseline below), and host
> WireGuard needs root the agent can't get. v2 is **fully containerized** (two peer containers, no
> host install) and tears down **to the captured baseline**, never to a hard-coded list.
## How this is spawned
The main session spawns this via the Agent tool (general-purpose), max 15 turns, subject to the same
security hooks as the main session (no direct mutating sudo; privileged mutations route through
sudo-bridge with the announce-first pattern). Spawn ONLY after the main session confirms the N0 gate:
`secret/netbird/db`, `secret/netbird/oidc`, `secret/cloudflare/dns-api` all populated and Authelia
restarted with OIDC live. Do NOT self-check the gate by mutating anything.
Spawned via the Agent tool (general-purpose), `--max-turns 15`, subject to the same security hooks
as the main session (no direct mutating `sudo`). Spawn ONLY after the main session confirms the N0
gate: `secret/netbird/oidc`, `secret/cloudflare/dns-api` populated and Authelia live with OIDC. Do
NOT self-check the gate by mutating anything.
---
@@ -18,87 +23,149 @@ restarted with OIDC live. Do NOT self-check the gate by mutating anything.
If you issue the same tool call twice with identical arguments, STOP and output the wrap-up block
with status=partially_succeeded.
## Blocked = stop
If a security hook or permission blocks an action, report a partial wrap-up. Never attempt creative
workarounds around a block.
## Owner-command relay (pause-for-sudo) — READ FIRST, this is how you get privileged commands run
You are non-interactive and CANNOT run host `sudo`, interactive auth, or production/outward-facing
mutations (e.g. writing a real DNS-zone TXT record). When you hit one, you do **not** improvise,
self-escalate, or silently die. You **pause and ask the owner**, who is at the keyboard, to run it in
the main conversation and paste the output back. Mechanism:
1. End your run with, as your **final message**, a block headed exactly `⏸ OWNER-COMMAND-REQUEST`
(this sentinel, NOT the wrap-up JSON — a pause is not a finish), with exactly:
- `host:` where to run it (`server-01 — ssh administrator@192.168.1.90`, or `primary (this host)`).
- `command:` exact copy-pasteable command(s), one per line, **no placeholders, no secret values**
(name a Vault path instead of inlining a value).
- `why:` one line — what it unblocks.
- `paste_back:` exactly what you need returned (stdout, stderr, exit code).
- `resume_at:` the step number you'll continue from.
Then STOP. Do not run it yourself, do not wrap up, do not proceed past `resume_at`.
2. The owner runs it and the main session feeds the raw output back to you (context intact). Resume at
`resume_at`.
3. If the pasted output shows failure, you may issue **one** follow-up `⏸ OWNER-COMMAND-REQUEST` to
remediate; if still blocked, emit a partial wrap-up and stop.
4. **Cap: at most 3 owner-command relays this run.** Beyond that = partial wrap-up.
## Blocked = stop (unanticipated blocks only)
If a security hook or permission blocks something the owner-command relay does NOT cover, report a
partial wrap-up. Never attempt creative workarounds around a block.
## Objective (bounded)
On **server-01 (`ssh administrator@192.168.1.90`)** stand up a throwaway NetBird self-hosted stack and
validate the machine-testable parts of the design, then tear it down. You are NOT proving the
cross-NAT streaming case and you are NOT completing the interactive browser SSO login — those are
explicitly **owner-in-the-loop** and you only set them up + report readiness (see "Owner-verifies").
On **server-01 (`ssh administrator@192.168.1.90`)** stand up a throwaway NetBird self-hosted control
plane **plus two containerized peers**, validate the machine-testable parts of the design (OIDC wiring,
relay-transport-is-not-cloudflare), then tear it all down. You are NOT completing the interactive
browser SSO login and you are NOT proving the cross-NAT streaming case — those are **owner-in-the-loop**
(see "Owner-verifies").
## server-01 baseline (FRESHLY VERIFIED 2026-09-18 — do NOT treat this host as a clean sandbox)
`docker ps` on server-01 at prompt-writing time shows these **7 resident containers** (offloaded prod +
sandbox instances — leave every one of them UNTOUCHED):
```
agent-sudo-agent-sudo-1
bitwarden-bridge-sandbox-d2celewbvh7e4fer77fcp9b5
hermes
jenkins
n8n-prod-h10eww4au274owpxozgizysh
n8n-sandbox-d2celewbvh7e4fer77fcp9b5
vault-sandbox-d2celewbvh7e4fer77fcp9b5
```
There is **no Obsidian container and no host-level netbird** — the v1 "Ollama/GPU + Obsidian only"
assumption was false. Host facts: `administrator` is in the `docker` group (all docker ops need NO
sudo); `/dev/net/tun` exists world-rw; the `wireguard` **kernel module is NOT loaded** (so force
**userspace WireGuard** in the peers — see step 5; do NOT `modprobe` unless userspace genuinely fails,
and then only via the owner-command relay).
**Step 0 MUST re-snapshot** `docker ps --format '{{.Names}}' | sort` and treat THAT live set as the
teardown baseline; the list above is only a sanity anchor.
## Scope allowlist (touch NOTHING else)
- **server-01 only**, in a scratch dir `~/netbird-sandbox/` you create on server-01.
- Vault **reads** of the paths below (AppRole login->use->revoke-self).
- **No changes to the primary server. No prod DB reads/writes. No container restarts on primary.
No commits, no pushes.** server-01's only permanent residents are Ollama/GPU + Obsidian — do not
disturb them; deploy into your scratch dir and remove everything you add.
- **server-01 only**, in a scratch dir `~/netbird-sandbox/` you create, and a dedicated docker bridge
network `nbsbx` + the containers your compose defines (control plane + `peer-a` + `peer-b`).
- Vault **reads** of the paths below (AppRole login→use→revoke-self).
- **Do NOT** touch any of the 7 resident containers, the primary server, any prod DB, or the host
outside your scratch dir. **No host installs. No commits, no pushes.**
## Credential map (exact — never discover auth at runtime)
Vault AppRole (from THIS host, primary):
Vault AppRole (resolve from primary, then use over the network from server-01 — or fetch on primary and
pass the values in; you decide, but revoke-self when done):
- `VAULT_IP=$(docker inspect vault-iwaulpoi5hwirdlogshmul40 --format '{{.NetworkSettings.Networks.coolify.IPAddress}}')`
- role-id file: `/opt/appdata/docker/docker-compose/vault/approle/role-id`
- secret-id file: `/opt/appdata/docker/docker-compose/vault/approle/secret-id`
- login: `POST http://$VAULT_IP:8200/v1/auth/approle/login` -> `.auth.client_token`; **revoke-self when done**.
- login: `POST http://$VAULT_IP:8200/v1/auth/approle/login` → `.auth.client_token`; **revoke-self when done**.
Paths + fields:
- `secret/netbird/oidc` -> `client_id`, `client_secret`, `issuer` (=`https://auth.reverseproxyserver.net`), `token_endpoint_auth_method` (=`client_secret_post`)
- `secret/cloudflare/dns-api` -> `api_token` (scoped Zone:DNS:Edit for reverseproxyserver.net; for DNS-01)
- **Do NOT fetch `secret/netbird/db`** — the Postgres DSN dry-run is an N2 task on primary (server-01
need not route to the prod Postgres network). Skip it here.
- `secret/netbird/oidc` → `client_id`, `client_secret`, `issuer` (=`https://auth.reverseproxyserver.net`), `token_endpoint_auth_method` (=`client_secret_post`)
- `secret/cloudflare/dns-api` → `api_token` (scoped Zone:DNS:Edit; only needed IF the owner authorizes DNS-01 at step 2)
- **Do NOT fetch `secret/netbird/db`** — the Postgres DSN dry-run is an N2 task on primary. Skip it.
Fetch each secret via a `/tmp` python script reading Vault over http (NOT `curl -sf | python3` — `-sf`
hides the body on error). **Mask every secret value in all output and in the wrap-up JSON.**
## Pre-decided design (do NOT re-decide — you execute these)
1. **Store = SQLite** (sandbox only). Official `netbirdio` compose + `setup.env`.
2. **Dashboard hostname = `netbird-sandbox.reverseproxyserver.net`** (already registered as an Authelia
redirect URI). Resolve it to server-01 on the LAN by adding a hosts entry on server-01 and telling
the owner to add one on their test machine (`192.168.1.90 netbird-sandbox.reverseproxyserver.net`).
3. **TLS = DNS-01 wildcard** for `*.reverseproxyserver.net` using the CF token — this simultaneously
validates DNS-01 (design §4 step 6). Use lego or the NetBird proxy's built-in DNS-01 if supported;
otherwise issue with `lego` in a container and mount the cert into the NetBird reverse proxy.
NEVER http-01 (server-01 has no public :80). If DNS-01 issuance fails, record it and continue with
a self-signed cert so the rest of the validation can proceed — but mark `dns01_works=false`.
4. **OIDC** = generic OIDC to Authelia (issuer/client from Vault). Configure the dashboard + management.
2. **Everything on a dedicated docker bridge `nbsbx`.** Control plane (`management`, `signal`, `relay`,
`dashboard`) + two peer containers all attach to it. Dashboard internal hostname
`netbird-sandbox.reverseproxyserver.net` (already an Authelia redirect URI); resolve it inside the
compose network (compose service alias / container `/etc/hosts`), NOT via the real public DNS.
3. **TLS = self-signed by default** for the sandbox (internal-only hostname; the browser login is
owner-verified anyway). **DNS-01 wildcard issuance mutates the REAL `reverseproxyserver.net` zone**,
so it is an **owner-command-relay** step: prepare the exact `lego`/acme DNS-01 command (CF token
from Vault, referenced by path — never inline the token) and **pause** for the owner to run it, then
paste back the result. If the owner declines / it's not authorized this run, mark
`dns01_works="deferred-owner"` and continue with self-signed. NEVER http-01 (no public :80).
4. **OIDC** = generic OIDC to Authelia (issuer/client from Vault). Configure dashboard + management.
## Steps (issue mutating commands standalone; verify as a SEPARATE command)
1. SSH to server-01, create `~/netbird-sandbox/`, pull the official NetBird self-hosted compose +
`setup.env`. Fill `setup.env`: domain `netbird-sandbox.reverseproxyserver.net`, OIDC issuer/client
from Vault, SQLite store.
2. Issue DNS-01 wildcard cert with the CF token. Verify: cert file exists + CN/SAN covers the host.
3. Bring the stack up. Verify (separate command): `management`, `signal`, `dashboard`, `relay`
containers are Up/healthy (`docker ps` on server-01).
## Steps (issue each mutating command standalone; verify as a SEPARATE command)
0. SSH to server-01. **Snapshot the baseline:** `docker ps --format '{{.Names}}' | sort` → save it;
this is your teardown target. Create `~/netbird-sandbox/`.
1. Pull the official NetBird self-hosted compose + `setup.env` into the scratch dir. Add the `nbsbx`
bridge network and two peer services `peer-a`/`peer-b` (see step 5). Fill `setup.env`: domain
`netbird-sandbox.reverseproxyserver.net`, OIDC issuer/client from Vault, SQLite store, self-signed
TLS.
2. **(Optional, owner-authorized only)** DNS-01 wildcard cert — via the owner-command relay per design
note 3. Otherwise skip → self-signed.
3. Bring the control plane up (`docker compose up -d` — no sudo needed, you're in the docker group).
Verify (separate command): `management`, `signal`, `dashboard`, `relay` are Up/healthy
(`docker ps`), and none of the 7 resident containers changed state.
4. **OIDC wiring check (agent-testable, no browser):** fetch
`https://auth.reverseproxyserver.net/.well-known/openid-configuration` and confirm `issuer` +
endpoints resolve; fetch the dashboard's served `config.json`/env and confirm it carries the right
`authority`/`clientId`. Record both. (The actual browser login is Owner-verifies #1.)
5. **Relay transport check (the Q10 proof, agent-testable):** enroll two throwaway peers on server-01
using a setup key (device-flow is Owner-verifies #2). Confirm they connect. Then force relay
fallback (block the direct path) and confirm the relayed path uses the **UDP relay endpoint, NOT any
cloudflare hostname** — capture the relay endpoint host:port and the transport. Record
`relay_transport` = `udp-direct` or whatever it actually is. Record whether NetBird's built-in
`relay` sufficed or a separate `coturn` was needed (`coturn_needed`).
6. **Tear down:** stop + remove every container you created, delete `~/netbird-sandbox/` and any cert
material, remove the hosts entry you added on server-01. Verify server-01 is back to Ollama/GPU +
Obsidian only (`docker ps`). Leave NOTHING running.
endpoints resolve over https; fetch the dashboard's served `config.json`/env and confirm it carries
the right `authority`/`clientId`. Record both. (Actual browser login = Owner-verifies #1.)
5. **★ Containerized relay-transport proof (the Q10 proof — no host install):**
- `peer-a` and `peer-b` each run the netbird client from `netbirdio/netbird` with
`cap_add: [NET_ADMIN]` and `devices: ["/dev/net/tun:/dev/net/tun"]`, joined to `nbsbx`. The
kernel `wireguard` module is NOT loaded, so netbird uses **userspace WireGuard (wireguard-go)**
entirely inside each container — no host module, no host sudo. If netbird refuses to fall back to
userspace, that `modprobe wireguard` is an **owner-command-relay** request (host sudo) — do NOT
attempt it yourself.
- Enroll both peers with a **setup key** (generate one via the management API/CLI inside the control
plane; device-flow is Owner-verifies #2). Confirm both show connected in `netbird status`.
- **Force relay fallback:** inside each peer container (NET_ADMIN, container-scoped — NOT a host
rule) add an `iptables` DROP on the direct peer↔peer WireGuard UDP path (block the other peer's
container IP on the WG data port) while leaving the path to the `relay` container open. Re-check
`netbird status --detail`.
- **Capture + assert:** the connection is now `relayed`, and the relay endpoint is the **`relay`
container's private bridge IP:port (172.x/10.x docker), NOT any `*.cloudflare*`/CF-edge host**.
Record `relay_transport` (e.g. `udp-relay-container`), `relay_is_cloudflare=false`, and whether
NetBird's built-in `relay` sufficed or a separate `coturn` was needed (`coturn_needed`).
6. **Tear down to baseline:** `docker compose down -v` in the scratch dir (removes the control plane +
both peers + the `nbsbx` network), `rm -rf ~/netbird-sandbox`. Verify (separate command):
`docker ps --format '{{.Names}}' | sort` **equals the step-0 baseline** — the 7 resident containers,
nothing added, none restarted. Leave NOTHING of yours running. (If the owner loaded the wireguard
module via relay, note it in `notes`; a loaded module is harmless to leave but must be disclosed.)
## Owner-verifies (you SET UP + REPORT READINESS; you do NOT perform these — no browser, and cross-NAT
## needs the owner's off-LAN devices)
1. **Browser OIDC login** to the dashboard (auth-code + PKCE) -> owner confirms Authelia login lands in
the NetBird dashboard. Leave a one-line instruction in the wrap-up `notes`.
2. **Device-flow enrollment** (`netbird up` opening a browser / device code) vs setup-keys -> owner runs
`netbird up` on a real device against the sandbox and reports if device-flow works. State this.
## Owner-verifies (you SET UP + REPORT READINESS; you do NOT perform these)
1. **Browser OIDC login** to the dashboard (auth-code + PKCE) → owner confirms Authelia login lands in
the dashboard. Leave a one-line instruction in `notes`.
2. **Device-flow enrollment** (`netbird up` device code) vs setup-keys → owner runs `netbird up` on a
real device and reports if device-flow works. State this.
3. **Cross-NAT DIRECT streaming proof** (the star Jellyfin proof) — deferred to N2 with real off-LAN
nodes; a single-host sandbox cannot represent it. State this explicitly; do NOT claim P2P streaming
is proven from the sandbox.
## Verify before reporting success
Do not set `status=succeeded` unless: stack came up healthy, OIDC discovery + dashboard config verified,
relay-transport captured as non-cloudflare, DNS-01 result recorded, AND teardown confirmed clean.
Do not set `status=succeeded` unless: control plane came up healthy, OIDC discovery + dashboard config
verified, the two peers enrolled, relay-transport captured as **non-cloudflare**, DNS-01 result recorded
(issued OR `deferred-owner`), AND teardown confirmed clean (== step-0 baseline). Otherwise
`partially_succeeded`.
## Persistence (do these as final actions — your context is discarded on finish)
## Persistence (final actions — your context is discarded on finish)
- Append a dated handoff to `research/netbird-design_2026-09-17.md` §7 (What was done / results for
steps 2,4,5 / device-flow + coturn verdicts / Next step) — this is a research doc, appending is in
steps 3,4,5 / DNS-01 + device-flow + coturn verdicts / Next step) — appending to a research doc is in
scope; do NOT commit.
- Write per-agent run log `logs/agent_runs/N1.json` with the wrap-up JSON.
- Update the semantic index: `python3 /opt/appdata/docker/.claude/scripts/embed_memory_dir.py --only-recent 3`.
@@ -108,18 +175,21 @@ relay-transport captured as non-cloudflare, DNS-01 result recorded, AND teardown
```json
{
"status": "succeeded | partially_succeeded | failed",
"project": "netbird #258 / N1 sandbox",
"project": "netbird #258 / N1 sandbox (v2 containerized)",
"actions_taken": [],
"actions_failed": [],
"files_touched": [],
"containers_restarted": [],
"owner_command_relays": [],
"containers_touched": [],
"baseline_restored": true,
"stack_healthy": true,
"oidc_discovery_ok": true,
"dashboard_config_ok": true,
"relay_transport": "udp-direct | ... ",
"peers_enrolled": 2,
"relay_transport": "udp-relay-container | ...",
"relay_is_cloudflare": false,
"coturn_needed": true,
"dns01_works": true,
"coturn_needed": false,
"dns01_works": "true | false | deferred-owner",
"device_flow_status": "owner-to-verify",
"teardown_clean": true,
"next_step": "",
@@ -191,3 +191,52 @@ to a genuinely clean host); (2) authorize the production DNS-01 challenge for th
a path for the root-WireGuard peer/relay test — either sudo-bridge allowlist entries for `netbird`/`wg`/
firewall-block, or a fully containerized two-peer NET_ADMIN design. Then re-run to deploy + capture
`relay_transport`. Cross-NAT DIRECT streaming stays an N2 item (a single host cannot represent it).
## 7 (cont.) — N1 v2 (containerized) sandbox handoff (2026-09-18, background agent)
**Status: partially_succeeded.** (This v2 run used the corrected baseline — the live `docker ps` set,
not the "Ollama/Obsidian" assumption — and a fully containerized peer design, resolving the two v1 blockers.)
### What was done
- Re-snapshotted the true baseline (7 resident containers) as the teardown target; created scratch
`~/netbird-sandbox/`. Confirmed `/dev/net/tun` present, **wireguard kernel module NOT loaded** (userspace
WG path), `administrator` in `docker` group → zero sudo used all run.
- Vault AppRole: read `secret/netbird/oidc`, token revoked (HTTP 204). Secrets masked.
- Pulled official NetBird images; built a concrete stack from the current `docker-compose.yml.tmpl` +
`management.json.tmpl` schema: SQLite, single-account-mode, generic OIDC to Authelia,
`IdpManagerConfig.ManagerType=none` (Authelia has no NetBird IDP-manager type — JWT validation via
HttpConfig is the correct generic-OIDC path), dedicated `nbsbx` bridge.
- **TLS deviation:** internal data plane (management/signal/relay) run **plaintext over nbsbx**
(`http` / `rel://`) to avoid a multi-turn cross-container self-signed-trust rabbit hole; dashboard stays
browser-facing self-signed. Does NOT affect OIDC JWT validation. `dns01_works=deferred-owner`.
### Results — steps 3, 4, 5
- **Step 3 (control plane healthy): PASS.** management/signal/relay/dashboard all Up. Management log:
loaded OIDC discovery from Authelia, overrode issuer/jwks/token/authz endpoints from it, SQLite engine,
migrations + indexes created, `Relay addresses: [rel://relay:33080]`.
- **Step 4 (OIDC wiring): PASS.** Discovery fetched over https from server-01 (issuer + authz/token/
userinfo/jwks resolve; `end_session_endpoint` null in Authelia — logout note). Dashboard-served runtime
config carries authority `auth.reverseproxyserver.net` + `netbird` clientId. Management independently
boot-loaded the same discovery.
- **Step 5 (relay-transport proof): BLOCKED (not fixable non-interactively).** Peer enrollment needs a
setup key; minting one needs an authenticated admin; the first NetBird account/user is bootstrapped only
by the OIDC **browser** login (auth-code+PKCE) = Owner-verifies #1. The headless `client_credentials`
bootstrap was tried and **Authelia rejected it**: `unauthorized_client — client not allowed grant
'client_credentials'` (netbird is registered auth-code/PKCE only). No setup-key CLI in the mgmt image.
Peers were therefore never started; `modprobe wireguard` never reached / not attempted.
- Design-level relay assertion holds: management advertises `rel://relay:33080` = relay container's
private nbsbx IP → **relay_is_cloudflare=false**. Live `relayed`-state capture needs enrolled peers.
### Verdicts
- **device-flow:** owner-to-verify. Note the browser auth-code+PKCE login is also the account-bootstrap
prerequisite for setup keys.
- **coturn:** not deployed (skipped; not needed for the relay-container proof). built-in relay healthy.
### Next step
1. Owner does one browser OIDC login to the dashboard to bootstrap the first account/admin
(Owner-verifies #1) → unlocks setup-key creation.
2. Mint a setup key (dashboard/REST) and run the containerized 2-peer + iptables-forced-relay proof to
capture live `relayed` transport = non-cloudflare.
3. For headless CI: register an Authelia client permitted `client_credentials` (or a PAT/service-account)
so the account can be bootstrapped without a browser.
4. N2: proper TLS via DNS-01, Postgres store, cross-NAT DIRECT streaming (needs real off-LAN nodes).