docs(netbird): add grill-me question bank (ran 2026-09-17, #258)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,93 @@
|
||||
# NetBird Design — Grill-Me Question Bank (prepared 2026-09-17)
|
||||
|
||||
> Prepared per owner request, to RUN in a later session — **after Twingate is finished + tested.**
|
||||
> Sequence: finish/test Twingate → run this NetBird grill-me → NetBird design → deploy (#258).
|
||||
> Sources refamiliarized: `netbird-twingate-gameplan.md`, `netbird-selfhost-buildplan.md`,
|
||||
> `cloudflared-tunnel-inventory_2026-09-17.md`, gameplan/buildplan decisions, `security_principles.md`.
|
||||
> Locked already (do NOT re-litigate unless owner reopens): NetBird = owner's own device mesh;
|
||||
> self-hosted is the destination; server-01 = throwaway sandbox; production plane on the PRIMARY server;
|
||||
> Postgres on homelab instance for prod; Authelia via generic-OIDC is the intended IdP.
|
||||
> New owner input (2026-09-17): **NetBird fully replaces cloudflared**; owner reaches everything via the
|
||||
> mesh; partners reach nextcloud + jellyfin (+ others) via Twingate; owner wants the Nextcloud public
|
||||
> link-shares to "operate on NetBird if possible."
|
||||
|
||||
## Priority 1 — the potential showstoppers (ask these first)
|
||||
|
||||
**Q1 — NetBird's own public reachability behind CGNAT (the circular-dependency).**
|
||||
Cloudflared works today because it's outbound-only and your network is (almost certainly) behind CGNAT.
|
||||
NetBird's self-hosted control plane (management/signal/relay) must be reachable from the internet for a
|
||||
roaming phone to connect. If cloudflared is retired, **what fronts NetBird's 443/gRPC (and relay) to the
|
||||
outside?** Sub-questions:
|
||||
- Is the primary server's WAN actually CGNAT'd, or do you have a routable public IP / port-forward
|
||||
capability at home? (This single fact decides everything below.)
|
||||
- If CGNAT: are you OK with (a) keeping ONE thin tunnel just for NetBird's control-plane endpoints
|
||||
(Cloudflare tunnel or NetBird's own), or (b) renting a small public-IP VPS to host the control plane,
|
||||
or (c) — the pragmatic hybrid — using NetBird Cloud for the control plane and self-hosting only where
|
||||
it helps? Your "self-hosted, no third party" principle pushes against (c); which wins if CGNAT forces
|
||||
the issue?
|
||||
|
||||
**Q2 — the public link-shares.** Your Nextcloud `business-proposals` link-shares are login-less URLs for
|
||||
external people (prospects) who won't install any client. NetBird/Twingate both REQUIRE a client, so
|
||||
those can't literally "operate on NetBird." What's the real requirement?
|
||||
- Who receives these links, and must they stay truly anonymous/public, or is a one-time login / passcode
|
||||
acceptable? Would you accept keeping a single narrow public path (one cloudflared hostname, or an
|
||||
object-store/static publish) *only* for outbound shares, while everything else goes mesh-only?
|
||||
|
||||
**Q3 — bootstrap-on-cloud vs straight self-host.** The buildplan recommended bootstrapping on NetBird
|
||||
Cloud first (fast, low-risk) then migrating. Given your "security paramount / self-hosted" stance and
|
||||
that NetBird is now your *primary* access to everything — do you still want the temporary cloud
|
||||
bootstrap as a proving ground, or go straight to self-hosted (no third-party plane ever touches the
|
||||
mesh, even temporarily)? (Ties to Q1 — if CGNAT forces a cloud/VPS plane anyway, this answer shifts.)
|
||||
|
||||
## Priority 2 — scope & membership
|
||||
|
||||
**Q4 — exact node set.** Which devices join your mesh? Confirm: primary server, server-01, laptop,
|
||||
phone. Any others (work machine, tablet, a second phone)? Which is the **routing peer** that advertises
|
||||
the home LAN + docker `172.16.16.x` so the mesh reaches services without per-service exposure?
|
||||
|
||||
**Q5 — phone as a first-class peer.** Earlier you wanted off-LAN SSH from your phone to the primary to
|
||||
provide the Bitwarden master password (the JIT unlock). That makes the phone a required mesh node with
|
||||
SSH-to-primary access. Confirm — and should the phone also carry ntfy push reachability if ntfy goes
|
||||
mesh-only?
|
||||
> ⚠️ CONSTRAINT (2026-09-17): the owner's **current phone is too old for the Twingate Android app** and
|
||||
> almost certainly too old for the **NetBird Android client** too. So the phone-based off-LAN BW-unlock
|
||||
> is NOT available on the current hardware — until a newer phone, the off-LAN unlock peer must be the
|
||||
> **laptop**, not the phone. Owner rarely leaves the residence; Discord is the accepted comms fallback.
|
||||
> Revisit "phone as peer" when the phone is upgraded.
|
||||
|
||||
**Q6 — users/identities.** Is the mesh single-user (just you), or do you want distinct NetBird
|
||||
identities (e.g., an admin identity vs a day-to-day one)? Any posture rules you care about (block
|
||||
enrollment from an unknown OS/version)?
|
||||
|
||||
## Priority 3 — the cutover (16 endpoints)
|
||||
|
||||
**Q7 — migration order & rollback.** We keep cloudflared running until NetBird is proven. Proposed:
|
||||
migrate the low-risk admin surfaces first (cadvisor, prometheus, grafana), prove mesh access, then the
|
||||
sensitive ones (vault, sudo-bridge, traefik), then nextcloud/jellyfin last. Agree? What's your rollback
|
||||
appetite — how long do cloudflared + NetBird run in parallel before you pull the tunnel?
|
||||
|
||||
**Q8 — Authelia sequencing.** Several admin services likely sit behind Authelia forward-auth. If `auth.`
|
||||
becomes mesh-only, anything depending on it must also be mesh-only (and reach it) — we can't orphan a
|
||||
service from its login. Do you know which services currently enforce Authelia, or should the design
|
||||
phase audit that first?
|
||||
|
||||
**Q9 — ntfy.** ntfy drives sudo-bridge phone approvals + alerts. If it goes mesh-only, your phone must
|
||||
be a peer for pushes to arrive off-LAN. Keep ntfy public, or make it mesh-only and rely on the phone
|
||||
peer? (Note: sudo-bridge → ntfy dependency means ntfy can't be bounced through sudo-bridge.)
|
||||
|
||||
## Priority 4 — build specifics (mostly confirm the research)
|
||||
|
||||
**Q10 — TURN/relay & P5.** Classic coturn needs a public UDP 3478; NetBird's WSS relay on 443 may avoid
|
||||
it. This was to be tested on server-01. With cloudflared possibly gone (Q1), the relay-through-tunnel
|
||||
assumption changes — fold this into whatever Q1 decides. Are you OK opening a scoped UDP port if the WSS
|
||||
relay can't carry roaming peers?
|
||||
|
||||
**Q11 — Postgres.** Confirm: dedicated `netbird` DB + user on the homelab Postgres, DSN in Vault
|
||||
(`secret/netbird/db`). Sandbox uses throwaway SQLite. Agree?
|
||||
|
||||
**Q12 — timing.** Roadmap has NetBird after Twingate + secrets-proxy #174 + agent-sudo #176. Does the
|
||||
"NetBird replaces cloudflared" priority bump it ahead of secrets-proxy/agent-sudo, or hold the order?
|
||||
|
||||
## After the grill-me
|
||||
Fold answers into an updated `netbird-selfhost-buildplan.md`, resolve Q1/Q2 with a short research spike
|
||||
if needed, then write `playbook_netbird_phases.md` (the phase-prompt set) and start on server-01.
|
||||
Reference in New Issue
Block a user