Files
claude-projects/claude-config/config/prompts/media-pipeline/RA-0_release_identity_capture.md
T
2026-10-02 00:19:54 -05:00

5.4 KiB

RA-0 — release-identity archive + capture + backfill (so saved .torrent files can be deleted) (spawned 2026-10-01, main session 771b2744)

Bounded background agent, media_pipeline (personal_projects #61; future_features #11 = later Prowlarr build). Budget: --max-turns 55. Same tool call twice with identical arguments → STOP with a partial wrap-up.

OWNER DECISION 2026-10-01 21:29: Prowlarr option A chosen, but ONLY do now the steps needed so the owner can safely delete the saved .torrent files; the Prowlarr build itself is deferred (future_features #11). Read first: /opt/appdata/docker/research/source_reacquisition_2026-10-01.md ("store_now" list + RA-0 phase). Known: download_jobs has infohash on only 1708 of 3524 rows (Movies 18 of 455).

GOAL: before any .torrent file is deleted, every torrent the owner has ever kept (saved .torrent files, qBittorrent BT_backup on BOTH qBittorrent instances — public qbittorrent and qbittorrent-private — and autobrr history) has its identity preserved in Postgres, and future downloads capture it automatically.

DECISIONS (pre-made):

  • D1 Credential-free source of truth: parse the .torrent FILES directly (BT_backup dirs + the owner's saved-.torrent folder(s) — locate them read-only via docker inspect mounts of the qbittorrent/qui/autobrr containers and the compose dirs; ask nothing). Write a stdlib-only Python bencode parser (no pip installs). Per torrent extract: infohash v1 (sha1 of the bencoded info dict) and v2 if present, name, total size, file list (path + size), comment, created by, creation date, tracker HOSTS only (strip passkeys/paths/query — NEVER store a full announce URL), private flag, source file path. Never print or store passkeys/announce URLs/cookies.
  • D2 autobrr history: read autobrr's sqlite DB READ-ONLY (copy it to your scratchpad first, query the copy) for release name, indexer, torrent id/guid/info url (strip any key/passkey query params), filter, timestamp; join on infohash or release name. If the DB is not readable without root → ⏸ OWNER-COMMAND-REQUEST.
  • D3 Storage: a NEW table media_pipeline.release_identity (do not alter download_jobs in this run): id, infohash_v1 (unique), infohash_v2, release_name, total_size, files jsonb, comment, created_by, creation_date, tracker_hosts text[], private boolean, indexer, indexer_torrent_id, info_url_sanitized, qbit_instance, source ('bt_backup'| 'saved_torrent'|'autobrr'|'download_jobs'), first_seen, download_job_id (nullable FK-less int, matched by infohash or name). Write the DDL as a migration file; do NOT run it on prod. Run it on the SANDBOX DB on server-01 (media-pipeline-sandbox-postgres-1, db media_pipeline_sandbox, user sandbox) and load the full backfill there to prove it.
  • D4 Backfill artifact: generate a single SQL (or COPY CSV + load script) with every row; keep a JSONL archive too (the same data) under /opt/appdata/docker/docker-compose/media-downloader-local/release-identity/ (create the dir). The main session reviews and applies the DDL + backfill to prod afterwards.
  • D5 Capture going forward: in media-downloader-local (find the code path that copies a finished torrent), add a call that records the torrent's identity into release_identity at copy time, reading the .torrent from qBittorrent's BT_backup by infohash (or the qBittorrent API only if the code already uses it with existing creds). Edit ONLY the sandbox twin (/opt/appdata/docker/docker-compose/server-01/media-pipeline-sandbox/media-downloader-local.py), scp it to the same path on server-01, add unit tests (fake .torrent built in tmp), run the full sandbox test suite. If a sandbox e2e is cheap (seedbox fixture → downloader-sandbox), run it and show the release_identity row. No live edit.
  • D6 Coverage report (the owner's go/no-go for deleting .torrent files): for every saved .torrent file → is its infohash in the backfill? For every download_jobs row → identity found or not (by infohash, else name). List any .torrent that could NOT be parsed. Verdict: SAFE_TO_DELETE (100% of saved .torrent files archived) or not, plus the exact list of gaps.

HARD RULES: read-only on every live container and file (cp to scratchpad for parsing is fine); NO deletion of any .torrent; no prod DB writes (SELECT only); no container restarts except media-pipeline-sandbox-* (log ELEVATED); no git; no sudo (⏸ OWNER-COMMAND-REQUEST block: host, command, why, paste_back, resume_at); never read .env files or print secrets; strip passkeys everywhere (grep your own outputs for 'passkey', '/announce?', 'torrent_pass' and 32+ hex query params before finishing — must be 0 hits). PERSIST: append "## RA-0 — 2026-10-01 (background agent)" to /opt/appdata/docker/docker-compose/media-downloader-local/.claude/context.md; run python3 /opt/appdata/docker/.claude/scripts/embed_memory_dir.py --only-recent 5. FINAL message = wrap-up JSON only: status, project "media_pipeline", torrent_sources[] (path, host, count), parsed_count, unparsed[], autobrr_rows_joined, migration_path, backfill_path, archive_path, sandbox_load_proof (row count), capture_change_summary, tests_passed, coverage {saved_torrents_total, archived, gaps[]}, download_jobs_coverage {total, with_identity}, verdict (SAFE_TO_DELETE|NOT_YET), passkey_scan_hits (must be 0), apply_commands (prod DDL + backfill, for the main session), actions_taken[], actions_failed[], unverified_claims[], next_step, notes.