docs(claude-config): infra synthesis + confirmed Sept status reset

- research/infrastructure-synthesis.md: reload infra corpus, fold in the
  local-AI-coding-stack research, reconcile vs prior decisions (no reversals;
  LifeOS is the one new open item; most of it is Horizon B)
- record confirmed July->Sept reality: still Pro (Max not upgraded), move done,
  server-01 offline, Tailscale/Voice unblocked, Boring Burger partnership reset,
  digital businesses start 2026-09-14
- INDEX + context.md updated to point at the synthesis and next-session deferred work

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X9iCzmxK2zbb8H1Ld3f8AN
This commit is contained in:
Backtalk6858
2026-09-08 19:00:21 -05:00
parent 293b8d51b7
commit 97df616f9a
3 changed files with 192 additions and 6 deletions
+13 -5
View File
@@ -56,11 +56,19 @@ happens here too.
- Do not migrate the memory directories — they are homelab-wide, not specific to this project.
## Current state
2026-09-08: Directory established as the Claude/Claude-Code research + config home. Moved the
existing `local-ai-coding-stack-research.md` into `research/`, created the catalog, decision log,
and config convention. Next: fold in findings from the earlier Claude-chatbot research session
and decide the goal framing (cost / independence / fully-local) that the local-stack brief hinges
on before going deeper on jcode/OpenCode/local-model choices.
2026-09-08: Directory established as the Claude/Claude-Code research + config home (committed +
pushed, `cead081`). Wrote `research/infrastructure-synthesis.md` — reloaded the infra corpus, folded
in the local-AI-stack research, reconciled against prior decisions (nothing reversed; LifeOS is the
one new open item; most of it is Horizon B). Confirmed the July→Sept reality reset: still **Pro**
(Max not upgraded), move done, **server-01 OFFLINE** (WiFi-only new house), Tailscale/Voice now
unblocked, Boring Burger partnership restructured, digital businesses start 2026-09-14.
**NEXT SESSION (server-01 must be online first):** finish the deferred end-of-session steps
(semantic index, Obsidian note, remaining memory stale-sweep) — see `session_last_summary.md`
DEFERRED list. Then the user wants to **trial Fable 5.1** (~$100 usage credits) on this research,
and fill me in on **new Agent-Sudo/secrets-proxy research** that may change direction (verify both-host
state first). Open decision on the table: **LifeOS voice-port pivot** vs the bespoke voice build, and
the LifeOS-vs-Gbrain context-layer overlap it forces.
## Update instructions
Update this file during every session debrief that touches this project — keep "Current state",
+10 -1
View File
@@ -6,7 +6,16 @@ re-verify versions, repos, and Linux support before acting on anything.
| Topic | File | Status | Key open decision |
|---|---|---|---|
| Local AI coding stack (inference engine, coding harness, LifeOS, voice, skill porting) | [local-ai-coding-stack-research.md](local-ai-coding-stack-research.md) | Surface-level, in progress | **Goal framing first:** cost vs. independence vs. fully-local — everything downstream (engine, harness, hybrid) follows from this. |
| **Infrastructure synthesis** — the whole landscape + reconciliation of the new research against decisions already made | [infrastructure-synthesis.md](infrastructure-synthesis.md) | Synthesized 2026-09-08 | **Confirm the July→Sept reality** (Max upgrade / Voice-Chat / Tailscale-Twingate / Agent-Sudo) before acting on anything. |
| Local AI coding stack (inference engine, coding harness, LifeOS, voice, skill porting) | [local-ai-coding-stack-research.md](local-ai-coding-stack-research.md) | Surface-level, in progress | **Goal framing:** RESOLVED by the existing vision = cost-reduction + tooling-independence, Claude stays the brain (NOT fully-local). See synthesis Part 3. |
### Where the prior infrastructure research lives (memory corpus)
Not duplicated here — cited in [infrastructure-synthesis.md](infrastructure-synthesis.md). Key files:
`project_two_horizon_operating_model` (governing filter), `project_ai_infrastructure_vision`
(Horizon-B master), `reference_open_source_ai_tools`, `reference_available_tools`,
`security_principles` + `feedback_autonomous_security_constrain_not_gate` + `playbook_zero_trust_security`,
`project_jenkins_hermes_boundary`, `project_tailscale_twingate_deploy`, `project_voice_chat_claude_code`,
`agent-builder/` (Agent-Sudo / Constrained Autonomy design + gameplan).
## Cross-cutting open decisions
Pulled up from the individual briefs so they don't get buried:
@@ -0,0 +1,169 @@
# Infrastructure Synthesis — where we are + what the new research changes
**Synthesized:** 2026-09-08. **Sources:** the memory corpus (see citations inline) + the new
[local-ai-coding-stack-research.md](local-ai-coding-stack-research.md).
> ✅ **STATUS CONFIRMED BY USER 2026-09-08** (resolves the ~6-week gap since the last recorded session, 2026-07-29):
> - **Claude Max: NOT upgraded — still Pro.** Justification stalled by a business shake-up (below). So everything "blocked on Max" (Gbrain/memory Track 2) stays blocked.
> - **Move: DONE** (ran long). User currently on **WiFi with poor quality** (new house has no grounding wire → EMF from appliances hurts speed/reception); plan to **wire critical infra to Ethernet tomorrow (2026-09-09)**, mobile stays wireless.
> - **Tailscale + Twingate: still PAUSED, now UNBLOCKED** by the completed move → can proceed. (Ethernet-wiring tomorrow is a natural predecessor.)
> - **Voice-Chat: PARKED, and the user wants to PIVOT** — new research indicates **LifeOS already ships this feature fully built (macOS)**; plan would be to **port it to Linux instead of building the bespoke faster-whisper+XTTS stack**. ⟶ reopens the locked Voice-Chat design; see Part 3 row 6 + the LifeOS row. (Note: the local-stack brief §5/§6 already covers this port and warns the LifeOS **Linux support is fragile** — VoiceServer's Linux fixes, PR #288, were lost in the v3.0 restructuring.)
> - **Agent-Sudo / secrets-proxy: NO significant progress.** User has **additional research that may change direction** — awaiting fill-in before any work. Primary-side check today: `agent-sudo-daemon` is `active` but **`NRestarts=17598`** (was 0 when healthy in July) = likely crash-loop; container check inconclusive (server-01 offline + docker perms). Full both-host verify = next session once server-01 is back.
> - **Businesses:** the **digital** ones (KDP — J.M. Hartley, Etsy — ConfettiPrintCo) are **NOT deployed**; work starts **Mon 2026-09-14**. Separately, the physical **Boring Burger Company** partnership restructured — see the business-reset note below. **No revenue yet** → Horizon B stays parked.
>
> **⚠️ Business partnership reset (needs proper capture in business memory — flagged, not yet saved):** Austin left after a falling-out; **Hailee** is the new **33% partner** (with the user + Tyler). Two employees came with her: **Grim** (Hailee's husband) and **Josh** (a close friend) — Josh runs the burger grill, Grim the fryer at pop-up events. **Switched fries → chips** for now until the fry recipe is perfected. Tyler's hours were also reduced. This is why the Max expense can't yet be justified.
---
## Part 1 — The infrastructure as it stands (recap)
### Governing filter — Horizon A vs Horizon B (`project_two_horizon_operating_model`)
- **Horizon A (now):** two jobs only — (1) make a business produce revenue, (2) make our interim
collaboration faster at #1. Two daily slots, ≤5 background agents.
- **Horizon B (parked):** the autonomous self-hosted AI ecosystem. Gated on
**revenue → ~$3,660 128GB MS-S1 Max → local 70b+ brain.** "A nice goal, not necessary."
- **The filter for every idea:** *"Does this move revenue, or make us faster at moving revenue,
BEFORE the hardware exists?"* No + assumes local/autonomous → tag `blocked_on_hardware`, park it.
- **Model/effort SETTLED (2026-07-15):** Opus 4.8 / medium standing default. Bake-off closed, Sonnet
cancelled, cost not a selection input (billing hard-capped, >$20/mo impossible).
### Cost / billing constraints (`feedback_no_claude_api`, `project_claude_max_upgrade`)
- **Claude Pro only.** No Anthropic API key; API billing intentionally OFF. `claude -p` runs on SDK
credits, **$20/mo hard cap.** Never recommend an API-key-dependent tool.
- **Claude Max ($100/mo)** was planned for **~2026-08-13** (Tyler covering) — *status unconfirmed.*
- What we have locally: Ollama on server-01 **and** primary — `nomic-embed-text` (embeddings) +
`llama3.1:8b` (simple/mundane chat only). No capable local thinking model until MS-S1 Max.
### Security stack (`security_principles`, `feedback_autonomous_security_constrain_not_gate`, `playbook_zero_trust_security`)
- **9 standing principles (P1–P9):** network vs at-rest encryption are separate; keys managed apart
from data (Vault/Bitwarden); audit findings are sensitive (security_audit DB + Bitwarden notes);
**Authelia over Authentik**; eliminate unnecessary exposure rather than authing it; Authelia bypass
rules for API/webhook paths and for services with strong native 2FA; **zero-trust internal**
(Docker net segmentation + TLS on Vault/Postgres + Tailscale later).
- **Autonomy security doctrine:** secure by **constraining capability, not gating access.** A
credential the autonomous machine must hold is theater. Split **EXECUTE (vetted, fully autonomous)**
from **EXPAND-allowlist (the only human touch, rare/async)**. Circuit-breaker splits by trip cause
(benign → auto re-arm Tier-3 only; attack-shaped → hard latch, human reset). Tier-4 = irreversible
**or** "guards-the-ceiling" actions, human-only by design. Prefer no new tool.
### Automation / deploy fabric (`reference_available_tools`, `project_jenkins_hermes_boundary`, `project_cicd_jenkins`)
- **Vault** (AppRole; dynamic IP via `docker inspect`), **secrets-proxy** (`/shell` with env_secrets —
the only sanctioned way to touch secrets), **sudo-bridge** (scoped root per server), **Bitwarden
Bridge**, **NTFY** (alerts + approval gates), **PostgreSQL**, **Gitea** (+ registry), **Traefik**
(HTTP origin; Cloudflare Tunnel terminates TLS), **N8N** (prod on server-01 :5678, sandbox :5679).
- **Jenkins = ALL deploys** (build, secret-inject, compose up, rollback). **Hermes = ALL monitoring**
(health, alerts, auto-restart, escalate) — **Hermes never redeploys itself; it triggers Jenkins.**
- **Coolify is RETIRED** (both servers, since 2026-06-25) — `coolify.managed` labels are vestigial;
`coolify-proxy` is just a standalone `traefik:v3.6`. ~7 N8N workflows still call the dead Coolify
API = silent debt awaiting Jenkins migration.
### Autonomy build — Agent-Sudo + Constrained Autonomy (`agent-builder/context.md`, git log)
- **Agent-Sudo deployed + ENFORCING on both hosts.** Replaces sudo-bridge. Design D1–D10 + D-CB1–D-CB9
locked. **P3 deployed on both hosts** (server-01 :8082, primary :8084). Old sudo-bridge still owns
primary :8082 — **cutover (#146) not done** as of late July.
- **Remaining scope = tiers 1/2/3.** Tier-1 undo was **broken** (`scoped_undo.py` vs `app.py` tier
mismatch — auto-executes with no undo capture); tier-2/3 were 503 stubs. Two ~10-min human decisions
(tier-1 undo, primary tier-2) were the next unblock. **Confirm current state before acting.**
- Recent commits (Sep 8 session start log) show continued Agent-Sudo/secrets-proxy wiring
(`c0293fa`, `1d37e45`, `506a481`) — so this progressed; verify against the artifacts.
### Networking — the immediate-future plan (`project_tailscale_twingate_deploy`, `feedback_cloudflared_file_size`)
- **PAUSED 2026-07-29** mid-preflight. Goal: retire the cloudflared inbound path.
- **Locked:** two independent systems — **me → Tailscale** (own fleet mesh: phone/tablet/laptop/both
servers as WireGuard peers; media + admin), **everyone else → Twingate** (scoped per-identity guest
access). Tailscale **cloud** control plane for v1 (not Headscale). **Host-terminated on both
servers**, no subnet routing (OPNsense segmentation preserved). Outbound-only, **zero inbound WAN
ports**; accepted trade — inline IDS/IPS goes blind inside the tunnel (compensated by overlay ACLs);
add explicit OPNsense **outbound** rules.
- **OPEN (resume here):** Tailscale container-networking mode — rec = `network_mode: host` +
`NET_ADMIN` + `/dev/net/tun` (documented exception to the compose template); Twingate connector =
standard bridged. Then keys→Vault, ACLs, Twingate resources, rollback → then write the
background-agent prompt. *Scheduled for Aug 1–2 weekend — status unconfirmed.*
### Memory / context / retrieval (`project_ai_infrastructure_vision` §Memory, `reference_open_source_ai_tools`)
- Memory must be **pull, not push** (every inject channel has a budget). `MEMORY.md` = bounded terse
index; cold store = memory dir; per-prompt `UserPromptSubmit` hook runs `semantic_recall.py`.
- **Direction = ADOPT Gbrain** (markdown-in-git + PGlite knowledge graph + MCP subprocess, ingests
Obsidian) as Track 2; **Obsidian = source of truth, Hermes = curator.** Companion tools: Graphify,
Claudian. The load-bearing lesson: the **lint/curation pass** decides whether the system compounds
or rots. **All of this is BLOCKED on the Max upgrade.**
- ⚙️ **Operational finding today:** `/recall`'s semantic layer errors with *"No route to host"* on its
Ollama embed endpoint, even though **localhost:11434 is UP**. The recall script points at an
unreachable host → misconfig. **In-scope fix for this claude-config directory.**
### Hardware roadmap (`project_ai_infrastructure_vision`)
- Now: LMDE 7 desktop + **RTX 2060 Super, 8GB (Turing SM 7.5)**.
- Phase 1: **MS-S1 Max 128GB** (unified LPDDR5 APU) → unlocks local LLM, Agent Builder, semantic
memory Phase 2. Phase 3: **NVIDIA GPU server (~$2.5–3.5k)** → NVENC + larger models.
---
## Part 2 — The new local-AI-coding-stack research, folded in
The brief (`research/local-ai-coding-stack-research.md`) untangles three independent layers:
**inference engine** (Ollama/vLLM/llama.cpp/LM Studio) · **coding harness** (Claude Code / jcode /
OpenCode / Aider / Cline / Goose) · **context-intent layer** (LifeOS). Governing constraint: the
**8GB VRAM ceiling** (~7–8B @ Q4). Headline findings:
1. **"Ollama went paid" was FALSE** — local Ollama is still free (MIT). Only *Ollama Cloud* is paid.
2. **vLLM not worth it on the 2060 Super** — Turing has no Marlin/FP8; single-user has no concurrency
win; Ollama/llama.cpp is more graceful on tight VRAM. Revisit only after a 12–16GB **Ampere+** card.
3. **Wiring local → Claude Code:** set `ANTHROPIC_BASE_URL` + dummy `ANTHROPIC_AUTH_TOKEN`.
**LM Studio 0.4.1+ exposes a native Anthropic `/v1/messages`** (no shim). Ollama's OpenAI `/v1`
needs a **LiteLLM** translation shim. **Reality check:** 7–8B models stumble on Claude Code's
agent loop — fine for autocomplete/boilerplate, not a drop-in agent driver.
4. **Harness field:** **jcode** = RAM-tiny + swarm/multi-agent, but *very immature/chaotic* (canonical
repo unclear, Zig-vs-Rust confusion). **OpenCode** = the most direct terminal-native Claude Code
replacement (75+ providers, client/server) — the safer daily driver. Aider (git-native), Cline
(VS Code), Goose, OpenHands (heaviest) round it out.
5. **LifeOS** (Daniel Miessler, formerly PAI) = a **context/intent layer** (TELOS format), *composes
with* a coding harness, does not replace it. Harness-agnostic by design but **Linux support is
documented better than it works** (VoiceServer had macOS-only deps; PR #288's Linux fixes were lost
in a restructuring). Still needs a model behind it — not a free-local path.
6. **Voice:** OS-level Whisper dictation (simple) vs LifeOS VoiceServer (fuller, needs macOS→LMDE
port: `afplay`→`mpg123`/`mpv`, `osascript`→`notify-send`, `launchctl`→systemd user service).
---
## Part 3 — Does the new research change any decision already made?
| # | Existing decision | New research says | Verdict |
|---|---|---|---|
| 1 | **Goal framing** was left open in the new brief (cost / independence / fully-local). | Brief §0: "decide this first." | **Already answered by the vision.** `project_ai_infrastructure_vision`: Claude is *deliberately kept as the reasoning brain*; everything else self-hosted; hybrid. So goal = **cost-reduction-over-time + tooling-independence, NOT fully-local/$0** (option (c) contradicts the vision). The brief re-opened a settled question — no change. |
| 2 | **Horizon A/B filter.** | Nearly the entire brief is self-hosted-brain / harness-migration / local-model work. | **No change, and this is the key point:** almost all of it is **Horizon B (parked)**, gated on revenue→MS-S1 Max. The only Horizon-A-eligible slice ("makes collaboration faster now") is **voice/dictation**, which is already the active, locked Voice-Chat project. **This research unlocks no new active work** — it's destination planning. |
| 3 | Harness migration candidates = **OpenHands (primary) + Cline** (`reference_open_source_ai_tools`). | Recommends **OpenCode** as the safer terminal-native daily driver; jcode only if swarm earns it; Aider/Goose also viable. | **Refinement, not reversal.** OpenCode was *not evaluated* in the old memory — add it (and jcode's immaturity warning) to the shortlist. Revisit when migration is actually on (Horizon B). |
| 4 | Inference engine = **Ollama** (`reference_available_tools`). | Stay on Ollama for this GPU; vLLM not worth it on Turing. | **Confirmed, sharper reasoning.** Nuance: vLLM/Marlin/FP8 need **NVIDIA Ampere+**, so that advice maps to the **Phase-3 NVIDIA GPU server**, *not* the MS-S1 Max (a unified-memory APU). Engine choice will differ by which hardware lands. |
| 5 | (none) Wiring local → Claude Code. | LM Studio native Anthropic endpoint (no shim); Ollama needs LiteLLM. | **New actionable reference detail.** Capture for the eventual hybrid wiring (Horizon B). |
| 6 | **Voice-Chat design LOCKED** — faster-whisper + Coqui XTTS-v2 on server-01 GPU, PTT, Jarvis (`project_voice_chat_claude_code`). | Offers plain Whisper dictation **or** LifeOS VoiceServer. | **No change — the locked plan supersedes** the brief's surface treatment (we already chose bespoke over VoiceServer). ⚠️ but its **Aug 14 deadline has passed** — confirm status. |
| 7 | Memory/context = **ADOPT Gbrain + Obsidian source-of-truth + Hermes curator** (`project_ai_infrastructure_vision`). | Introduces **LifeOS** as a context/intent layer that composes with a harness. | **New open question, does NOT override Gbrain.** LifeOS's TELOS *intent* layer is arguably additive; its *retrieval/skills* overlap Gbrain+Hermes. Serious Linux-maturity warnings. Decide additive-vs-redundant later — **do not adopt on faith.** Horizon B. |
| 8 | Local Ollama is free (`reference_available_tools`). | Confirms — "Ollama went paid" was a mistaken premise. | **Confirmed.** Kills a false premise that could have driven a needless migration. |
**Bottom line:** the new research **reverses nothing**. It (a) *confirms* Ollama + the 8GB ceiling +
Claude-stays-the-brain, (b) *refines* the harness shortlist (add OpenCode; note jcode immaturity) and
the engine/hardware mapping (vLLM ↔ Ampere+ only), (c) *adds one genuinely new open question*
(**LifeOS** vs the chosen Gbrain/Obsidian/Hermes context stack), and (d) mostly lands in **Horizon B**,
so it changes the *destination map*, not the *current worklist*.
---
## Part 4 — Open decisions & next actions
**Needs the user (decisions):**
1. **Confirm the July→September reality** — Max upgrade? Voice-Chat status vs the Aug 14 deadline?
Tailscale/Twingate (Aug 1–2)? Agent-Sudo tier-1/2/3 + primary cutover (#146)? This gates everything.
2. **LifeOS** — evaluate as additive intent layer vs redundant with Gbrain/Hermes (Horizon B; don't
adopt on faith given Linux-maturity warnings).
**Horizon-A eligible (could act now, cheap, "faster collaboration"):**
3. **OS-level Whisper dictation** as a lightweight interim, independent of the full Voice-Chat build.
4. **Fix `/recall`'s embed endpoint** (points at an unreachable Ollama host; localhost:11434 is up) —
in-scope for this directory, restores semantic memory search.
**Horizon-B / when hardware or migration is on (park, don't pull forward):**
5. Benchmark 2–3 candidate 7–8B coding models (Qwen-Coder, DeepSeek-Coder) at Q4 for fit + context +
tool-calling reliability through a harness.
6. Update `reference_open_source_ai_tools` with OpenCode + jcode + the engine/hardware mapping.
7. Stand up one hybrid path (Ollama/LM Studio + a harness, wired via `ANTHROPIC_BASE_URL` + LiteLLM).
> All harness/model/LifeOS findings move fast — re-verify versions, canonical repos, and Linux
> support before acting (the brief itself flags this).