Files
claude-projects/claude-config/config/prompts/media-pipeline/MP-27_t2_retry_fix.md
T
2026-10-02 00:19:54 -05:00

3.8 KiB

MP-27 — T2: environmental failures must not burn retries into skipped_permanent — sandbox → GO/NO-GO (spawned 2026-10-01, main session 771b2744)

Bounded background agent, media_pipeline (personal_projects #61). Budget: --max-turns 45. Owner-approved fix order step 3 (2026-10-01). Same tool call twice with identical arguments → STOP with a partial wrap-up.

READ FIRST: /opt/appdata/docker/docker-compose/media-transcoder/design/REVIEW_2026-10-01_full_code_review.md (T2, and T-S1/T-S2 for context) and the COMMON block in /home/administrator/.claude/projects/-opt-appdata-docker/memory/playbook_media_pipeline_phases.md ("### COMMON block") — follow its HARD RULES + ACCESS MAP + script-edit workflow exactly (sandbox only; record BASELINE = live sha256 of /opt/appdata/docker/docker-compose/media-transcoder/media-transcoder.py, currently eca0129a…; scp the working copy to server-01 AND confirm the container hash — MP-24 lesson: the server-01 copy must match). Prod: 65 sources stranded as skipped_permanent (transcode_jobs ids 3350-3414; Blue Bloods S02/S05, Peaky Blinders S02, JJK S01 v2, JJK S03) after an environmental failure (read-only tmp / DNS) used all MAX_RETRY_FAILURES in ~3 minutes (media-transcoder.py ~3842-3843, 5001-5003, 5305-5315).

DECISIONS (pre-made — apply exactly, one feature):

  • D1 Classify failures: ENVIRONMENTAL = the failure is not about the input file (OSError EROFS/ENOSPC/EACCES on the tmp/output dirs, DNS/connection errors to media-api/DB, ffmpeg/ffprobe exit caused by a missing/unwritable output dir, subprocess timeout while the host is unhealthy — find the actual error strings the code sees; list them). INPUT = everything else (decode errors, verify failures, corrupt source).
  • D2 ENVIRONMENTAL failures do NOT increment the per-file failure count. Instead: exponential backoff for that file (start 5 min, cap 6 h), one ntfy alert per distinct environmental error class per 6 h (existing notify helper), and a preflight health check before each job (tmp + output dirs writable, media-api reachable) that pauses intake while unhealthy (log once per state change).
  • D3 skipped_permanent is NOT permanent across restarts for rows whose last error was environmental: on startup, those rows are eligible again (log "[REQUEUE-ENV] …"). INPUT failures keep today's MAX_RETRY_FAILURES semantics.
  • D4 No change to verification, encode, A/V, picker, or naming logic. TESTS: unit tests for the classifier (each env error string → ENVIRONMENTAL; a decode error → INPUT), backoff schedule, startup eligibility, and that INPUT failures still reach skipped_permanent after MAX_RETRY_FAILURES. Sandbox e2e: make sandbox-tmp read-only (chmod inside the sandbox transcoder container on the sandbox tmp mount, or the equivalent) → stage one fixture → expect backoff + no failure-count burn + 1 ntfy log line (ntfy disabled in sandbox: assert the log line) → restore writability → the job completes done. Then an INPUT-failure fixture (truncate a copy of a fixture to 1 MB) → reaches skipped_permanent after MAX_RETRY_FAILURES as today. Full sandbox test suite must pass. Clean sandbox at rest (COMMON rules). NOT IN SCOPE: requeueing the 65 prod rows (the main session does that after promotion), any other review finding. OUTPUT: wrap-up JSON only: status, project "media_pipeline", verdict (GO|NO-GO), baseline_sha256, working_copy_sha256 (primary twin AND server-01 container must match), env_error_strings[], tests_passed, e2e[] (case, result), promote_commands (BASELINE-guarded cp + verify + owner restart command), prod_requeue_plan (how to make the 65 rows eligible after promotion — SQL or restart — exact commands, NOT run), actions_taken[], actions_failed[], containers_restarted[], unverified_claims[], next_step, notes. PERSIST per COMMON (context.md block "## MP-27 — 2026-10-01 (background agent)" + embed).