docs(claude-config): harness-layer grill-me decision + research prompts A/B/C

- decisions: AI-harness LAYER entry (scope, 4-pain rubric, R1-R3 hard rules,
  Ubuntu excluded -> Docker Sandboxes revisit trigger marked dead)
- config/prompts: A harness layer, B jcode vs executors, C HyperAgent marketing
  fit, README spawn table; mirrored to Obsidian Resources/Prompts
- context.md: 23:05 current-value block supersedes 22:45

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X9iCzmxK2zbb8H1Ld3f8AN
This commit is contained in:
Backtalk6858
2026-09-08 22:57:13 -05:00
parent 5462928be7
commit 61c544118b
6 changed files with 222 additions and 6 deletions
+5 -5
View File
@@ -1,10 +1,10 @@
# Claude & Claude Code — Research & Config Home — Context
> **NEXT SESSION PLAN — CURRENT VALUE (rewritten 2026-09-08 22:45, Fable 5.1, supersedes the 21:40/21:43 blocks below).** Still on FABLE, credits ~$12 used. Research + decisions + preflight ONLY, no build work.
> 1. **AI-harness grill-me FIRST** (personal_projects 221): "what do I want out of an AI harness" — one question at a time with recommendations. Inputs: `research/local-ai-coding-stack-research.md` §3–§6, `reference_open_source_ai_tools`, `project_voice_chat_claude_code`, synthesis Part 3. Known pain points to test every candidate against: (a) 200K context → end-of-session checklist at 80% (context-monitor.py), (b) MEMORY.md auto-memory injector cap ~24.4 KB (memory_tier_manager.py WARN 17 KB / ACT 22 KB; tiers are the workaround), (c) no API key / Pro OAuth only / $20 cap, (d) hooks + skills + memory stack must survive (session-start.sh, security-enforcement.py, automation-detector.py, 3 skills + 15 queued), (e) LMDE 7 / Debian 13 Linux support.
> 2. **Write THREE background-agent prompts** (use playbook_background_agent_prompts + mandatory wrap-up block; user wants them saved so Opus 4.8 can run them if credits run out): (A) AI-harness research — LifeOS (current) vs free/OSS harnesses, must do everything the grill-me says; (B) jcode vs Claude Code — is the "bigger context window" real, does it fix (a)+(b), canonical repo (`local-ai-coding-stack-research.md` §3 says repo/language unclear), and find any better alternative to BOTH; (C) HyperAgent — what it is, cost, no-API-key fit, use for API-business client acquisition + KDP/Etsy product marketing (business_projects 41; run it in the api-digital-business conversation). Save all three in Obsidian at `Resources/Prompts/` (NEW folder — no prompts folder existed as of 2026-09-08; agent prompts previously lived only in agent-builder/*.md) and in `claude-config/config/prompts/`.
> 3. Spawn the agents one at a time (≤5 background rule) → results into `research/` + INDEX rows → decisions into DECISIONS.md → preflight.
> 4. **Skills track exists** (2026-09-08 22:30): `skill` is an automation type; rule = `feedback_skill_vs_playbook.md` (Desktop store), injected in every dir via session-start.sh; build queue personal_projects 223–237 (first three: 223 scaffold, 224 end-session, 210 secret-pull, 225 pg-query). Stop hook `automation-detector.py` FIXED (INSERT column mismatch had silently killed every tag capture; now types + reports failures). Row 238 = security-enforcement.py crash on nested quoting.
> **NEXT SESSION PLAN — CURRENT VALUE (rewritten 2026-09-08 23:05, Fable 5.1, supersedes the 22:45 block, which is now history).** Still on FABLE, credits ~$14 used. Research + decisions + preflight ONLY, no build work.
> 1. **Harness grill-me DONE 22:50** → `decisions/DECISIONS.md` "AI-harness LAYER" entry + Obsidian `Resources/Grill-Me/2026-09-08 — AI Harness Layer.md`. Key corrections: **LifeOS is NOT installed** (premise was false; we run Claude Code + our own layer); scope = the layer around Claude Code; 4 pains = rubric; hard rules R1 Pro-OAuth / R2 merge-not-replace / R3 free+OSS; distro soft (LMDE now, Arch/Omarchy or Fedora possible, **Ubuntu never** → Docker Sandboxes revisit trigger marked DEAD); one Incus-VM trial; voice rides on the verdict; fixed candidates + native baseline.
> 2. **Prompts A/B/C WRITTEN 23:00** → `config/prompts/` (README = spawn table) + Obsidian `Resources/Prompts/` (new folder). DB: rows 221 (in_progress), 239, business 41 annotated.
> 3. **Spawn order A → B → C, one at a time** (general-purpose agent, paste body verbatim). A = `research/ai-harness-layer-research.md`; B = `research/jcode-vs-claude-code.md` (reuses A's R1 OAuth finding); C spawns from `/opt/appdata/docker/Business/API Idea` → `research/hyperagent-marketing-fit.md` + copy to Obsidian `Business/Research/`. After each: check INDEX.md row, follow-up grill-me, DECISIONS.md entry. Then preflight only; Opus 4.8 executes.
> 4. **Skills track exists** (22:30): `skill` automation type; rule `feedback_skill_vs_playbook.md` (Desktop store), injected everywhere via session-start.sh; build queue personal_projects 223–237 (first: 223 scaffold, 224 end-session, 210 secret-pull, 225 pg-query). `automation-detector.py` FIXED (INSERT column mismatch silently killed every tag capture) — **uncommitted in /opt/appdata/docker** with session-start.sh; row 238 = security-enforcement.py crash on nested quoting.
## What this project does
The dedicated home for everything about **how we use Claude and Claude Code to do work** —
@@ -0,0 +1,70 @@
# Prompt A — AI harness LAYER research (LifeOS vs OSS vs our own stack)
Tracks: personal_projects row 221. Spawn from `~/Desktop/claude/claude-config`. Research-only.
---
You are a research analyst producing one scored comparison report. You will read the web and local files and write exactly one Markdown file; you will not install anything, change any config, write to any database, or need any credential.
## Task
Decide which **layer around Claude Code** (memory, skills, hooks, context routing, voice) best fixes four named pains under three hard rules, and name the single candidate worth a sandboxed trial. The executor (Claude Code itself) is NOT under evaluation here — that is a separate report.
## Context / pre-fetched data
- Current stack (the baseline): Claude Code on Claude Pro (OAuth, no API key, $20/mo hard cap). Custom layer at `/opt/appdata/docker/.claude/hooks/` (session-start.sh, context-monitor.py, security-enforcement.py, automation-detector.py, track-file-edits.sh, conflict-detector.py), skills at `/opt/appdata/docker/.claude/skills/` (diagnose, grill-me, recall; 15 more queued), six memory stores under `~/.claude/projects/*/memory/` with per-store MEMORY.md indexes, Postgres-backed semantic recall (`/opt/appdata/docker/.claude/scripts/semantic_recall.py`, nomic-embed-text via Ollama), an Obsidian vault as knowledge source of truth, and an end-of-session checklist that fires at 80% context.
- Machine: LMDE 7 (Debian 13 base), RTX 2060 Super 8 GB. User may move to Arch (Omarchy) or Fedora later; will never use Ubuntu.
- Prior reading you must use: `research/local-ai-coding-stack-research.md` §4–§6 (LifeOS lineage, fragile Linux, VoiceServer PR #288 lost in v3.0, issue #685), `research/infrastructure-synthesis.md` Part 3 rows 6–7 (LifeOS vs Gbrain/Obsidian/Hermes overlap; voice decision), `~/.claude/projects/-opt-appdata-docker/memory/reference_open_source_ai_tools.md`.
- LifeOS is NOT installed anywhere today (verified 2026-09-08). Any claim that "we run LifeOS" is false.
## The four pains (score each candidate 0–3 per pain, with one sentence of evidence)
1. **Context**: 200K window, checklist forced at 80%. A layer cannot enlarge the window; score how much it REDUCES per-session injection (context.md, MEMORY.md, hook output) or defers loading to on-demand retrieval.
2. **Memory cap**: the auto-memory injector caps MEMORY.md at ~24.4 KB; tiers are the workaround. Score retrieval-on-demand vs inject-everything.
3. **Rule adherence**: playbooks unread, tags not emitted, checklist steps skipped. Our own fix is converting 14 playbooks to skills. Score only what a candidate adds BEYOND skills (enforcement hooks, structured routing, self-curating skills).
4. **Voice + life ops**: PTT voice, calendar, mail, TELOS-style intent. Score Linux-working voice specifically; macOS-only = 0.
## The three hard rules (any fail = DISQUALIFIED, still list it with the reason)
- **R1 Pro OAuth only.** Must work with a Claude Pro subscription through Claude Code's own login. Verify Anthropic's CURRENT (2026) terms on third-party use of subscription OAuth tokens; cite the page and date. A candidate that needs `ANTHROPIC_API_KEY` for Claude is out.
- **R2 Merge, not replace.** Must coexist with our `~/.claude/settings.json` hooks and skills. Read each candidate's installer/docs: does it own settings.json, overwrite hooks, or merge? Cite the file.
- **R3 Free and open source.** State the license and any paid tier. Paid tiers are out; an OSS core is still scored.
Soft factors (score 0–3): Linux health (open issues mentioning Linux/Debian/Arch/Fedora in the last 90 days, and whether Linux is CI-tested), training-loop fit (can it log decisions as JSONL for a future local model), maintenance burden (commits last 90 days, bus factor).
## Fixed candidate list (do not add others; note if one is dead)
1. **Baseline**: our current stack + Claude Code native features on Pro (auto-compact, `--continue`/`--resume`, memory tool, 1M-context model availability on Pro, project/user CLAUDE.md, native skills). Score honestly; this row may win.
2. **LifeOS / PAI** (danielmiessler/LifeOS) — pin the current version; check VoiceServer Linux state; Fabric relationship.
3. **SuperClaude** framework.
4. **Claude-Flow** (ruvnet).
5. **oh-my-claudecode** (or the current name of that project).
6. **claude-mem**.
7. **Basic Memory** (basicmachines).
8. **Letta** (formerly MemGPT) used as a memory server behind Claude Code.
## Step-by-step
1. Read the three local files listed above first (one turn).
2. For each candidate, in ONE turn per candidate, run the web searches together: repo + license + last commit, install method vs settings.json, Linux issues (90 days), auth model (OAuth vs API key).
3. Verify R1 once, centrally: Anthropic's current policy on subscription OAuth in third-party tools. This single fact may disqualify most of the list; state it plainly at the top of the report.
4. Fill the scoring table. Total = sum of pain scores (max 12) + soft factors (max 9). Disqualified rows show "DQ" plus the failed rule.
5. Write the report to `research/ai-harness-layer-research.md` with sections: Verdict (one paragraph: the single trial candidate, or "baseline wins"); R1 finding with citation; scoring table; per-candidate notes (5–10 lines each); what a sandboxed trial in the Incus VM on server-01 should test (a checklist of ≤8 checks); what to steal from losing candidates even if we do not adopt them; sources with dates.
6. Append one row to `research/INDEX.md` in the existing table format.
## Output rules
- Every score has a one-line evidence sentence with a URL. No unsourced claims.
- Mark anything you could not verify as UNVERIFIED rather than guessing.
- Do not recommend anything that requires an Anthropic API key, a subscription, or Ubuntu.
## MANDATORY WRAP-UP (required regardless of success or failure)
Before stopping for ANY reason — task complete, error, or approaching turn limit — output this JSON as your final message. Do not stop without it.
{
"status": "succeeded|partially_succeeded|failed",
"actions_taken": ["action 1 — outcome", "action 2 — outcome"],
"actions_failed": ["action — reason it failed"],
"notes": "anything relevant for the next session",
"trial_candidate": "<name or 'baseline'>",
"r1_oauth_finding": "<one sentence + URL>"
}
If you hit --max-turns before finishing, set status="partially_succeeded" and list what remains in actions_failed.
--max-turns 15
If you issue the same tool call or command twice with identical arguments, STOP immediately and output the mandatory wrap-up with status=partially_succeeded.
@@ -0,0 +1,70 @@
# Prompt B — jcode vs Claude Code, and any better EXECUTOR than both
Tracks: personal_projects row 239. Spawn from `~/Desktop/claude/claude-config`. Research-only. Run AFTER prompt A so its R1 (OAuth) finding can be reused.
---
You are a research analyst testing two specific claims about jcode and then ranking coding-agent executors against fixed rules. You will read the web and local files and write exactly one Markdown file; you will not install anything, change any config, write to any database, or need any credential.
## Task
Answer three questions with evidence: (1) Does jcode give a real advantage over Claude Code for this user? (2) Do its claimed advantages fix the two named pains? (3) Is there a better alternative to BOTH under the hard rules?
## Context / pre-fetched data
- The user's two pains this report must test against:
- **(a) Context**: Claude Code sessions run on a 200K window and our hook forces an end-of-session checklist at 80%. The user's hope: jcode has "a bigger context window so we won't have to run the checklist as often."
- **(b) Memory injection cap**: Claude Code's auto-memory injector caps the per-store MEMORY.md at ~24.4 KB (our memory_tier_manager.py warns at 17 KB, acts at 22 KB). The user's hope: jcode "fixes the CLAUDE.md size limitation."
- Ground truth to hold the claims against: a harness cannot enlarge a model's context window; only the model (and the provider's plan) sets it. A harness CAN change what it injects, how it compacts, and whether memory is retrieved on demand.
- Prior reading: `research/local-ai-coding-stack-research.md` §3 (jcode: ~27.8 MB RAM per session vs ~386 MB; swarm/multi-agent; canonical repo unclear between github.com/jrish11/jcode and github.com/1jehuang/jcode; language described as Zig vs Rust; ~v0.81.3; ~16 Homebrew installs in a year). Section §7 recommends OpenCode as the stable daily driver.
- Existing shortlist memory: `~/.claude/projects/-opt-appdata-docker/memory/reference_open_source_ai_tools.md` (OpenHands, Cline, Goose, Aider, Pi).
- If `research/ai-harness-layer-research.md` exists, read its "R1 finding" section and reuse it instead of re-researching Anthropic's OAuth policy.
- Constraints: Claude Pro (OAuth login, no API key, $20/mo hard cap); LMDE 7 now, maybe Arch/Fedora later, never Ubuntu; our hooks/skills/memory layer must survive (see prompt A's R2); RTX 2060 Super 8 GB means local models top out at ~7–8B Q4.
## Hard rules (fail = DISQUALIFIED for daily use with Claude; still list it)
- **R1** Can use a Claude Pro subscription without an API key, under Anthropic's current terms.
- **R2** Supports hooks (pre/post tool, session start/stop) and skills/commands so our layer can port. Name the equivalent mechanisms or "none".
- **R3** Free and open source; state license.
Note: a candidate that fails R1 for Claude but is FREE with another model on its own subscription-free tier (for example Gemini CLI with a personal Google account) is NOT disqualified as a secondary tool — score it in a separate "non-Claude executor" row and say what it would be used for.
## Fixed candidate list
1. **Claude Code** (baseline, current). Include what Pro offers today: 1M-context model availability, auto-compact behavior, `/compact` and `--resume`, the memory tool, native skills and hooks.
2. **jcode** — first establish the canonical repo, language, license, maintainer count, and release cadence; then verify each claim from §3 and the two user hopes.
3. **OpenCode**.
4. **Aider**.
5. **Goose**.
6. **Cline** (VS Code / JetBrains).
7. **OpenHands** (CLI + SDK).
8. **Gemini CLI** (free personal tier; different model; 1M context) — scored as a possible SECONDARY executor for long-context reading tasks only.
## Step-by-step
1. Read the local files listed above (one turn).
2. jcode: one turn of searches to settle repo/language/license/maintainers; one turn to verify the RAM and swarm claims; one turn to check its context-management and memory-file behavior (does it compact? does it read CLAUDE.md-style files? any size limits?).
3. For each remaining candidate, ONE turn: auth model with Claude (OAuth vs API key), hooks/skills mechanism, context/compaction behavior, memory-file handling and any size limits, license, Linux health (issues last 90 days).
4. Write the claim-test table: rows = pains (a) and (b); columns = Claude Code today, jcode, best alternative; cells = "fixes / partially / does not" + one evidence sentence + URL.
5. Write the executor scoring table: hard rules R1–R3 (pass/fail), then 0–3 each for: context handling, memory handling, hook/skill parity with our layer, Linux health, maintenance (commits/maintainers last 90 days), local-model support for the 8 GB card.
6. Write `research/jcode-vs-claude-code.md` with sections: Verdict (three sentences answering the three questions); claim-test table; executor scoring table; jcode fact sheet (repo, language, license, version, maintainers, install size); per-candidate notes; what would have to be true for switching to be worth it (a short list of triggers); sources with dates.
7. Append one row to `research/INDEX.md` in the existing table format.
## Output rules
- Every claim has a URL and a date. UNVERIFIED beats a guess.
- Say plainly if "bigger context window" turns out to be a model property, not a jcode property.
- Never recommend a tool that needs an Anthropic API key, a paid plan, or Ubuntu.
## MANDATORY WRAP-UP (required regardless of success or failure)
Before stopping for ANY reason — task complete, error, or approaching turn limit — output this JSON as your final message. Do not stop without it.
{
"status": "succeeded|partially_succeeded|failed",
"actions_taken": ["action 1 — outcome", "action 2 — outcome"],
"actions_failed": ["action — reason it failed"],
"notes": "anything relevant for the next session",
"jcode_verdict": "<advantage: yes|no|conditional + one sentence>",
"pain_a_fixed_by": "<tool or 'none'>",
"pain_b_fixed_by": "<tool or 'none'>",
"best_executor": "<name>"
}
If you hit --max-turns before finishing, set status="partially_succeeded" and list what remains in actions_failed.
--max-turns 15
If you issue the same tool call or command twice with identical arguments, STOP immediately and output the mandatory wrap-up with status=partially_succeeded.
@@ -0,0 +1,50 @@
# Prompt C — HyperAgent: can it join the marketing system?
Tracks: business_projects row 41. Spawn from the api-digital-business conversation (`/opt/appdata/docker/Business/API Idea`). Research-only. Copy the report to Obsidian `Business/Research/` at wrap-up.
---
You are a research analyst assessing whether a tool called "HyperAgent" can be added to an existing small-business marketing system. You will read the web and local files and write exactly one Markdown file; you will not install anything, change any workflow, write to any database, or need any credential.
## Task
Determine what HyperAgent is, what it costs, whether it can run under this user's constraints, and exactly where (if anywhere) it would plug into three marketing jobs: API-business client acquisition, KDP book marketing, and Etsy product marketing.
## Context / pre-fetched data
- **Disambiguation is step one.** At least three things carry the name: Hyperbrowser's open-source **HyperAgent** (TypeScript browser-automation agent on Playwright, LLM-driven), the FSoft-AI4Code **HyperAgent** research system (generalist software-engineering agents, a paper), and assorted SaaS products. The user heard of it as "an AI agent to supercharge marketing and client acquisition." Identify which one that is, and evaluate the others only in one line each.
- **The marketing system today** (API business, `/opt/appdata/docker/Business/API Idea/.claude/context.md`): `lead_pull.py` searches GitHub repos → keyword exclusion + Claude relevance gate → `api_business.pending_leads` (cap 28) → N8N "Daily Outreach Emails" workflow (id 4fuzFJAclba2Syy7, currently PAUSED pending N8N migration to server-01) sends via Brevo → "Brevo Email Open Notification" workflow tracks opens/clicks → `rejected_leads` + `lead_review_log` tables log the Claude gate's decisions as training data. Product: Dividend Tracker API on RapidAPI.
- **Two more digital businesses start 2026-09-14, not deployed:** KDP (author name J.M. Hartley) and Etsy (shop ConfettiPrintCo). No marketing system exists for them yet.
- **Constraints:** no Anthropic API key (Pro plan OAuth only; `claude -p` runs on SDK credits under a $20/mo hard cap); local models available = Ollama `llama3.1:8b` and `nomic-embed-text` on an 8 GB GPU; everything self-hosted on Docker (server-01 192.168.1.90 and primary 192.168.1.88); N8N is the orchestration layer; Jenkins = deployments, Hermes = monitoring (rule: evaluate Jenkins/Hermes fit for any automation); every automated system must log decisions as training data.
- **Platform-risk rule:** any tool that drives a logged-in browser session on Etsy, Amazon KDP, LinkedIn, GitHub, or Gmail must be checked against that platform's automation terms. Account bans on Etsy/KDP would kill a business before it starts. State the risk per use case.
## Step-by-step
1. Read the API-business context.md (one turn).
2. Disambiguate HyperAgent (one turn of searches): official site/repo, license, pricing, last commit, maintainers, required LLM providers.
3. Constraint check (one turn): does it run with an OpenAI-compatible local endpoint (Ollama) or only with cloud keys? Can it run headless in Docker on Linux? Any per-run cloud cost (hosted browsers, proxies)?
4. Fit per job (one turn each, three jobs): for each of (A) API client acquisition, (B) KDP marketing, (C) Etsy marketing — what concrete task would HyperAgent do that N8N + Python + Claude gate cannot already do; where it would sit in the pipeline (before lead_pull, replacing lead_pull, after outreach, or a new pipeline); platform-ToS risk; local-model quality risk at 8B.
5. Alternatives (one turn): name up to three free/OSS tools that do the same job with less risk (for example Playwright scripts under N8N, browser-use, Skyvern, Stagehand) and say in one line each whether they beat HyperAgent here.
6. Write `~/Desktop/claude/claude-config/research/hyperagent-marketing-fit.md` with sections: Verdict (adopt / trial / reject, one paragraph); What it is (fact sheet); Constraint verdict (no-API-key, cost, Docker/Linux, local-model); Fit table (rows A/B/C; columns: task, pipeline position, adds beyond current stack, ToS risk, model-quality risk, verdict); Jenkins/Hermes fit; training-loop hook (where its decisions would be logged, which table); Alternatives table; a proposed grill-me question list (≤6 questions) for the user before any build; sources with dates.
7. Append one row to `~/Desktop/claude/claude-config/research/INDEX.md` in the existing table format.
## Output rules
- Every claim has a URL and a date; UNVERIFIED beats a guess.
- Do not recommend anything requiring an Anthropic API key, a paid plan above $20/mo total, or a hosted browser service with per-run fees unless the fee is stated and under the cap.
- No build steps, no workflow JSON, no DB writes.
## MANDATORY WRAP-UP (required regardless of success or failure)
Before stopping for ANY reason — task complete, error, or approaching turn limit — output this JSON as your final message. Do not stop without it.
{
"status": "succeeded|partially_succeeded|failed",
"actions_taken": ["action 1 — outcome", "action 2 — outcome"],
"actions_failed": ["action — reason it failed"],
"notes": "anything relevant for the next session",
"which_hyperagent": "<one line>",
"constraint_verdict": "<runs under constraints: yes|no|conditional + one sentence>",
"fit_verdict": {"api_clients": "adopt|trial|reject", "kdp": "adopt|trial|reject", "etsy": "adopt|trial|reject"}
}
If you hit --max-turns before finishing, set status="partially_succeeded" and list what remains in actions_failed.
--max-turns 15
If you issue the same tool call or command twice with identical arguments, STOP immediately and output the mandatory wrap-up with status=partially_succeeded.
+18
View File
@@ -0,0 +1,18 @@
# Background-agent prompts (research track, Sept 2026)
Canonical copies live here; mirrored to Obsidian `Resources/Prompts/`. Each prompt follows
`playbook_background_agent_prompts` (bounded task, fixed candidate list, mandatory wrap-up JSON,
`--max-turns 15`, loop detection). They are **research-only**: no installs, no DB writes, no secrets.
| # | File | Tracks | Spawn from | Output |
|---|---|---|---|---|
| A | [A_ai-harness-layer-research.md](A_ai-harness-layer-research.md) | personal_projects 221 | claude-config | `research/ai-harness-layer-research.md` |
| B | [B_jcode-vs-claude-code.md](B_jcode-vs-claude-code.md) | personal_projects 239 | claude-config | `research/jcode-vs-claude-code.md` |
| C | [C_hyperagent-marketing-fit.md](C_hyperagent-marketing-fit.md) | business_projects 41 | `/opt/appdata/docker/Business/API Idea` conversation | `research/hyperagent-marketing-fit.md` (copy to Obsidian `Business/Research/`) |
**Spawn rule:** one at a time (A → B → C), `subagent_type: general-purpose`, paste the prompt body
verbatim. After each: read the report, add the INDEX.md row if the agent missed it, then run the
follow-up grill-me before the next spawn. Works on Fable 5.1 or Opus 4.8 unchanged.
**Grill-me that shaped A and B:** `decisions/DECISIONS.md` 2026-09-08 "AI-harness layer" entry
(Obsidian `Resources/Grill-Me/2026-09-08 — AI Harness Layer.md`).
+9 -1
View File
@@ -15,6 +15,14 @@ entry links back to the research that drove it, so the reasoning survives.
---
## 2026-09-08 — DECIDED (grill-me 22:50): AI-harness LAYER scope, rubric, hard rules, trial depth
- **Decision:** ADOPT the framing — research the **layer around Claude Code** (memory/skills/hooks/routing/voice), not the executor; premise "we run LifeOS" is FALSE (nothing installed). Rubric = 4 pains (200K/80% checklist; ~24.4 KB MEMORY.md cap; rule adherence beyond the skills track; voice + life ops). Hard rules = R1 Pro OAuth/no API key, R2 merge-not-replace our hooks, R3 free + OSS. Distro = soft factor (LMDE 7 now; Arch/Omarchy or Fedora possible; **Ubuntu never**). Depth = research then ONE sandboxed trial in the server-01 Incus VM. Voice rides on the harness verdict. Fixed candidates + native-baseline row (baseline may win).
- **Why:** every candidate is the same category as the layer we already built; the only honest comparison is against that baseline plus Claude Code's own newer features.
- **Research:** prompts in [../config/prompts/](../config/prompts/) (A harness layer, B jcode/executor, C HyperAgent); outputs land in `research/`. Grill-me note: Obsidian `Resources/Grill-Me/2026-09-08 — AI Harness Layer.md`.
- **Implementation:** pending — spawn A → B → C one at a time; follow-up grill-me after each.
- **Revisit when:** Anthropic changes subscription-OAuth terms; distro changes; skill run logs show rule adherence solved; Max upgrade lands.
- **Side effect:** the Docker Sandboxes "revisit when a host moves to Ubuntu 24.04+" trigger (CA-D11 entry below) is **DEAD** — Ubuntu is excluded; fallback remains `@anthropic-ai/sandbox-runtime`.
## 2026-09-08 — DECIDED: the four small Agent-Sudo / recovery calls (grill-me 21:15–21:30)
- **Decision:** **A1 = (a)** tier 1 captures scoped-undo (implement D4; fix the test that pins tier-3-only). **A3 = (c)** primary tier-2 stays 503 — dissolved by CA-D11. **A7(A) = no**, an explicit tier-4 SUDO.md rule is not an attack signal (audit-only). **Recovery layer = host systemd level-0** (`control-plane-up` oneshot+timer, fixed command list, `ip_nonlocal_bind=1`) with Hermes as level-1; row 193 recommendation = keep fail-closed once the daemon has backoff. A7(A) and the recovery shape were defaults I took — object to reopen.
- **Why:** the 09-07 outage was boot order (LAN-IP port binds before WiFi), not a crash; an irreversible tier-1 is a gate removed without a constraint added.
@@ -32,7 +40,7 @@ entry links back to the research that drove it, so the reasoning survives.
actions but never said where the agent runs. The hypervisor already exists (Incus qemu driver, KVM).
- **Research:** [../research/autonomy-isolation-evaluation.md](../research/autonomy-isolation-evaluation.md) §3–§4
- **Implementation:** [GAMEPLAN_security-infra-deploy.md](GAMEPLAN_security-infra-deploy.md) Step 3; CA-D11 amendment lands with it.
- **Revisit when:** a host moves to Ubuntu 24.04+ (Docker Sandboxes becomes supported), or Incus VM
- **Revisit when:** ~~a host moves to Ubuntu 24.04+ (Docker Sandboxes becomes supported)~~ (DEAD 2026-09-08 22:55: Ubuntu excluded by the user), or Incus VM
overhead proves too high (fallback: `@anthropic-ai/sandbox-runtime`).
## 2026-09-08 — DECIDED (grill-me 21:12): re-scope Constrained Autonomy around auto mode