prompts(media-pipeline): MP-20, MP-20b, MP-21, MP-21b sandbox agent prompts (verbatim)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
Backtalk6858
2026-09-25 21:27:33 -05:00
parent 3cab46b9bb
commit 9f8adf653e
5 changed files with 309 additions and 1 deletions
+1 -1
View File
@@ -33,7 +33,7 @@ follow-up grill-me before the next spawn. Works on Fable 5.1 or Opus 4.8 unchang
## Project subfolders (added 2026-09-25)
Execution/build/design prompts that belong to an ongoing project set live in a subfolder, saved **verbatim as spawned**:
- `media-pipeline/` — MP-10 build, AV-CHECK, AV-FIX, VAAPI QP calibration, VAAPI build (2026-09-24); MP-18 build, MP-19 design, MP-19b design, MP-19b R2 build (2026-09-25). Index/decisions: memory `playbook_media_pipeline_phases.md`.
- `media-pipeline/` — MP-10 build, AV-CHECK, AV-FIX, VAAPI QP calibration, VAAPI build (2026-09-24); MP-18 build, MP-19 design, MP-19b design, MP-19b R2 build, MP-15a density scan, MP-15 v2 hybrid picker, MP-20 HEVC copy assembly (NO-GO), MP-20b direct-map, MP-21 A/V Stage-1 start_time (NO-GO), MP-21b A/V mux timeline (2026-09-25). Index/decisions: memory `playbook_media_pipeline_phases.md`.
- `infra/` — NTFY-272 hardening (2026-09-24); #273 image-autoupdate research + P1a pin proof, cloudflared replacement research (2026-09-25). #273 P2 prompt lives verbatim in memory `playbook_image_autoupdate_phases.md`.
The 13 files above were recovered from the session transcript on 2026-09-25 because they had been spawned inline without being saved first.
- `wireguard/` — #274 home-access tunnel phase prompts: `W1_build_prep.md` (2026-09-25; write files only, no deploy). Design: `/opt/appdata/docker/research/diy_wireguard_design.md`.
@@ -0,0 +1,75 @@
# MP-20 — HEVC video-copy assembly failure + truncated output marked done — sandbox fix, verdict only
Tracks: personal_projects 61 (media_pipeline). Evidence: MP-15 v2 agent run 2026-09-25 (fixture M1 failed at final
assembly: "Can't write packet with unknown timestamp"; 9.3 s partial output then marked done by
`[SKIP] Output already exists`). Pre-decided in the main session 2026-09-25 evening. Written + spawned 2026-09-25.
---
You are a bounded background agent for the media_pipeline project. PHASE MP-20 — fix the HEVC video-copy path of the R2 chunked pipeline and harden the output-exists skip, in the SANDBOX, verify, deliver a verdict. Budget: --max-turns 40 (reason: reproduce on old + new code, root-cause, two fixes, unit tests, 4 sandbox transcodes, regression batch).
If you issue the same tool call twice with identical arguments, STOP and output the wrap-up block with status=partially_succeeded.
HARD RULES
- NEVER restart, stop, recreate, or exec a mutating command in ANY live/prod container, on any host. Read-only `docker exec ... sha256sum|ffprobe|cat` and `docker ps/inspect` on live containers are allowed.
- Container restart allowlist starts EMPTY. Before you start/stop/restart a SANDBOX container, add it to your allowlist and log "ELEVATED <name>" in actions_taken; verify healthy after; list it in containers_restarted. Only `media-pipeline-sandbox-*` containers on server-01 may be elevated. Throwaway `docker run --rm` containers you create are fine.
- NO git add/commit/push. NO writes to the primary host's Postgres. NO edits to the LIVE media-transcoder.py (/opt/appdata/docker/docker-compose/media-transcoder/). NO promotion — verdict only; the main session promotes.
- Issue every mutating command as its OWN tool call; verify with a SEPARATE call.
- WAITING: never hand-write nested-quote wait loops. Write wait conditions as small script FILES on the target host (SQL in a .sql file run with `psql -f`), run with `timeout 900`. Before wrap-up confirm no leftover waiters on both hosts (`ps -eo pid,args | grep "[t]imeout"`; kill only your own). Do NOT leave long-running background jobs and go idle — run each batch in the foreground with a timeout.
- Mask secrets. No sudo. If a security hook BLOCKS an action, do not work around it: partial wrap-up.
- Do not make design decisions beyond the pre-made ones below. If reality contradicts them (e.g. the root cause is not what D1 assumes and D1 does not fix it), STOP after step 2 and report with evidence.
ACCESS MAP
- server-01: ssh administrator@192.168.1.90 (batch read commands per ssh call).
- Sandbox dir (server-01, and a git-tracked twin on the primary host at the same path): /opt/appdata/docker/docker-compose/server-01/media-pipeline-sandbox/
- Sandbox containers: media-pipeline-sandbox-media-transcoder-sandbox-1, media-pipeline-sandbox-postgres-1 (DB media_pipeline_sandbox, user sandbox, no password via docker exec).
- Baseline reset: `TRUNCATE pipeline_learning, transcode_jobs, transcode_overrides, pipeline_events, download_jobs, rename_jobs, show_configs RESTART IDENTITY;` then `docker exec -i media-pipeline-sandbox-postgres-1 psql -U sandbox -d media_pipeline_sandbox < init/02-seed.sql`; ALWAYS restart the sandbox transcoder after a reset (done-set is cached at startup). Also clear the transcoder's persistent done-state under its tmp root (find it: grep TranscodeState / state file path in the script) if a file would be skipped.
- sandbox-output/, sandbox-renamed/, sandbox-tmp/ hold ROOT-owned files: clear them from inside the container (`docker exec media-pipeline-sandbox-media-transcoder-sandbox-1 find <in-container path> -mindepth 1 -delete`; paths via docker inspect).
- STAGE = cp a fixture into sandbox-input/<CATEGORY DIR>/<Show>/ — category dirs: "H264 ENGLISH DUBBED ANIME TO BE TRANSCODED", "H264 ENGLISH SUBBED ANIME TO BE TRANSCODED", "H264 LIVE ACTION SERIES TO BE TRANSCODED", "H264 MOVIES TO BE TRANSCODED". sandbox-input/ must hold only the 4 empty category dirs at the end.
- Sandbox = software libx265 (NVIDIA host, no VAAPI) — expected; the copy path uses no encoder.
- Cut/remux fixtures with a throwaway container: `docker run --rm --entrypoint ffmpeg -v <dir>:/w <sandbox transcoder image>` — mount at /w, NEVER /lib.
- Script workflow: the primary working copy /opt/appdata/docker/docker-compose/server-01/media-pipeline-sandbox/media-transcoder.py is the MP-15 v2 build (sha256 starts d1d19b02 — record the full sha as BASELINE; it = live 7617fc0a + the approved-but-unpromoted MP-15 v2 picker). Build your fix ON TOP of it (both ship together). Edit; `python3 -m py_compile`; scp to server-01; restart the sandbox transcoder; confirm the in-container sha256.
- Pre-R2 reference code (for the step-1 comparison ONLY, never deployed to the sandbox container): `git -C /opt/appdata/docker show 249bbe8:docker-compose/media-transcoder/media-transcoder.py` (sha e3b6f751; MP-10 shadow + AV-FIX, before R2/MP-18/VAAPI). Live = 7617fc0a (commit 8e62433).
- Logs: `ssh administrator@192.168.1.90 "docker logs --since <ts> media-pipeline-sandbox-media-transcoder-sandbox-1 2>&1 | grep -E '<pattern>'"`.
- Unit tests next to the working copy on primary: test_vaapi_cmd.py, test_verify_output_streams.py, test_av_stage2_suppression.py, test_mp19b_resume.py, test_mp15_picker.py (111 total today). All must still pass, plus your new test file.
BACKGROUND (main session — do not re-diagnose beyond step 2)
- R2 (MP-19b) pipeline: per-part video-only encodes cut at keyframes (`.tmp` → rename) into `job_r_<sha1>` dirs; one full audio encode (audio.mka); final mux `build_mux_cmd()` (~L2227) = concat demuxer over the parts with option-B `duration` directives (`concat_durations()` ~L2803, `write_concat_list()` ~L2815) + audio.mka + subs/metadata/chapters from the source, all `-c copy`. Assembly call site ~L3909–3912.
- HEVC sources take the copy branch (~L2035–2048): `plan["encoder"]="copy"`, `plan["video_args"]=["-c:v","copy"]`. On R2 these copied parts then fail the final concat mux with "Can't write packet with unknown timestamp" (stream-copied HEVC parts carry no/invalid pts for the concat demuxer, and/or the `duration` directives don't fit copied GOP boundaries — confirm in step 2).
- The failed assembly left a partial (9.3 s) file AT the final output path; the next scan's secondary guard `output_already_exists()` (~L4724, call ~L5153–5157) only checks size > 0, so it logged `[SKIP] Output already exists (marking done)` → a truncated file is treated as finished. Output is written straight to its final `.mkv` name, which filebot-monitor could also pick up mid-write (its watcher filters by suffix: VIDEO_EXTENSIONS = .mkv .mp4 .avi .m4v .mov .wmv .flv .webm .mpg .mpeg — any other suffix is ignored).
- Prod exposure is low (0 copy-path jobs in 90 days, 0 jobs since R2 went live), but any HEVC download will hit it.
PRE-MADE DECISIONS (binding; main session 2026-09-25)
- D1 Copy path = NO chunking. When `plan["encoder"] == "copy"`, produce exactly ONE video part covering the whole source (`-map 0:v:0 -c:v copy`, no -ss/-t), and write the concat list with that single part and NO duration directive (or bypass concat and map the part directly — pick whichever is the smaller change and say which). Rationale: a copy takes seconds, so there is nothing to resume; one part avoids cross-part timestamp problems. Audio/subs/metadata/chapters flow unchanged. If the single part still fails the mux with the same error, add `-fflags +genpts` to the part's input for the copy path only, and record that it was needed.
- D2 Atomic output: assemble into `<output_path>.partial` (a suffix filebot ignores), then `os.replace()` to the final name ONLY after the mux exits 0 and the completeness check (D3) passes. On any failure delete the `.partial` and leave the job in its failed state (existing failure handling/logging); never leave a file at the final name. Applies to ALL encoders, not just copy.
- D3 Completeness check helper `output_is_complete(output_path, source_duration)`: ffprobe must read the file, it must have ≥1 video stream, and its format duration must be within max(2.0 s, 1 %) of the source duration (use the source duration the job already probed — do not add a second source probe if one exists). Used by D2 before the rename AND by `output_already_exists()`.
- D4 Skip guard: `output_already_exists()` returns True only if `output_is_complete()` passes (probe the source duration there; if the SOURCE cannot be probed, keep today's size>0 behaviour and log WARNING). An incomplete file is renamed to `<name>.incomplete-<YYYYmmdd-HHMMSS>` (suffix filebot ignores; never deleted — owner inspects), logged at WARNING `[SKIP-GUARD] Incomplete output quarantined: …`, and the input is queued normally.
- D5 Constants OUTPUT_DURATION_TOL_S = 2.0, OUTPUT_DURATION_TOL_PCT = 0.01 at module level with a comment citing MP-20. Update the function docstrings + module fix history (MP-20 entry). Existing log wording unchanged elsewhere.
FIXTURES (server-01 sandbox-fixtures/)
- mp10/"MP10 M1 (1993).mkv" — the movie that failed (ffprobe first: confirm the video codec is hevc; if it is NOT hevc, find what sent it down the copy branch and report).
- CREATE `sandbox-fixtures/mp20/`:
- "MP20 H1 - S01E01.mkv" = an HEVC dubbed-anime-layout episode clip ≥ 400 s (so R2 would make ≥3 parts at SEGMENT_SECONDS=180). Source: take an existing HEVC fixture if one is ≥400 s; otherwise build it by encoding an existing ≥400 s sandbox fixture with libx265 (`-c:v libx265 -preset ultrafast -x265-params keyint=48`, audio/subs copied) in the throwaway container. Record how it was made.
- "MP20 T1 - S01E02.mkv" = a TRUNCATED copy of H1 (first 10 s, `-t 10 -c copy`) used to test D4.
- Regression fixtures (H.264, encode path, already exist): mp18/ progressive + interlaced clips, mp19/ join clip (read names with ls).
STEPS
0 Baseline: BASELINE sha (full) of the working copy; confirm sandbox container runs the same hash; live container sha = 7617fc0a (read-only); sandbox-input empty; fixtures present; create mp20/ fixtures and ffprobe them; reset DB; restart transcoder (ELEVATE).
1 Reproduce on CURRENT code: stage M1 (movies) and H1 (dubbed) — one batch at a time. Record for each: job state, the assembly stderr, output duration vs source, and whether the next scan logs `[SKIP] Output already exists`. Then, WITHOUT deploying old code, run the pre-R2 (249bbe8) copy-path assembly logic by hand in a throwaway container only if needed to answer "did pre-R2 handle HEVC copy correctly?" (a quick read of its copy branch + one manual ffmpeg of its command form is enough). Record the answer (regression yes/no).
2 Root cause: show the exact failing ffmpeg command (from the log) and the minimal manual reproduction in a throwaway container; confirm that the D1 single-part form succeeds manually. If it does not, STOP and report (see HARD RULES).
3 Implement D1–D5. Add `test_mp20_copy_assembly.py` (no ffmpeg needed — patch subprocess/ffprobe helpers): copy plan → exactly one part, no duration directive; encode plan → unchanged multi-part list; assembly writes `.partial` then renames only on success; failure → no file at the final name; `output_is_complete` tolerance edges (source 1400 s: 1398.1 ok, 1385.9 fails; source 60 s: 58.1 ok, 57.9 fails); `output_already_exists` with complete / truncated (→ quarantined, returns False) / unprobeable-source (size>0 fallback) files. Run ALL unit test files — all pass. py_compile; scp; restart (ELEVATE); confirm in-container hash.
4 Sandbox transcodes (end-to-end), one category batch at a time:
a) H1 (dubbed, HEVC): job done; output duration within tolerance of source; video stream copied (codec hevc, same frame count ±1 as the source via `ffprobe -count_packets`); audio + subs + forced disposition present as the MP-15 v2 picker decides; `[VERIFY]` lines no fail; no `.partial` left.
b) M1 (movies, HEVC): same checks; forced sub = "English Forced".
c) T1 skip-guard: after H1's output exists, REPLACE it (from inside the container) with the 10 s truncated file under the same name, clear the done-state so the secondary guard is exercised, restart; EXPECT `[SKIP-GUARD] Incomplete output quarantined`, a `.incomplete-*` file, and H1 re-transcoded to a complete output.
d) Regression (encode path): one mp18 interlaced clip + the mp19 join clip — job done, output duration within tolerance, parts > 1 where expected, `.partial` absent, [VERIFY] no fail. Run `tools/av_align.py` (sandbox dir) on the mp19 output vs its source and record head_frame_offset + max audio drift (expect ≤ ±1 frame / ≤ 40 ms, as before).
Un-stage/clear between batches; reset DB + restart if the done-set would skip a file.
5 Verdict: GO iff step 1 reproduced the bug, steps 3–4 meet every expectation, and all unit tests pass. NO promotion. On NO-GO give the failing expectation with evidence.
6 Clean up: sandbox-input back to 4 empty dirs; clear output/renamed/tmp (including any .incomplete-*); reset DB; sandbox transcoder RUNNING idle on the new working copy; keep sandbox-fixtures/mp20/* (regression); no leftover waiters on either host; delete your temp scripts.
SCOPE ALLOWLIST: primary working copy server-01/media-pipeline-sandbox/media-transcoder.py + new test_mp20_copy_assembly.py next to it; server-01 sandbox dir contents (script copy, sandbox-input/output/renamed/tmp, sandbox-fixtures/mp20/) + sandbox DB; your temp scripts; the context.md append. Nothing else.
PERSIST BEFORE YOU FINISH
- Append (one `cat >>` heredoc; append only) "## MP-20 HEVC copy assembly — 2026-09-25 (background agent)" to /opt/appdata/docker/docker-compose/media-downloader-local/.claude/context.md: What was done / Decisions / Current state / Next step.
- Run: python3 /opt/appdata/docker/.claude/scripts/embed_memory_dir.py --only-recent 5
- FINAL message = wrap-up JSON only, always:
{"status":"succeeded|partially_succeeded|failed","project":"media_pipeline","phase":"MP-20","actions_taken":[],"actions_failed":[],"files_touched":[],"containers_restarted":[],"unverified_claims":[],"baseline_sha256":"","new_sha256":"","repro":{"M1":{},"H1":{},"pre_r2_regression":null},"root_cause":"","genpts_needed":null,"transcode_results":{"H1":{},"M1":{},"T1_guard":{},"regress_mp18":{},"regress_mp19":{}},"av_align_mp19":{"head_frame_offset":null,"max_drift_ms":null},"unit_tests":{"passed":null,"failed":null},"verdict":"GO|NO-GO","diff_summary":"","next_step":"","notes":""}
@@ -0,0 +1,74 @@
# MP-20b — HEVC video-copy: map video straight from the source + truncated-output guard — sandbox fix, verdict only
Tracks: personal_projects 61 (media_pipeline). Evidence: MP-15 v2 agent run 2026-09-25 (fixture M1 failed at final
assembly: "Can't write packet with unknown timestamp"; 9.3 s partial output then marked done by
`[SKIP] Output already exists`). Supersedes MP-20 (NO-GO 2026-09-25: its D1 was wrong; its step-2 diagnosis is inlined
below). Revised decisions by the main session 2026-09-25 evening. Written + spawned 2026-09-25.
---
You are a bounded background agent for the media_pipeline project. PHASE MP-20b — fix the HEVC video-copy path of the R2 chunked pipeline and harden the output-exists skip, in the SANDBOX, verify, deliver a verdict. Budget: --max-turns 40 (reason: two fixes, unit tests, 4 sandbox transcodes, regression batch + 2 av_align runs; diagnosis is already done).
If you issue the same tool call twice with identical arguments, STOP and output the wrap-up block with status=partially_succeeded.
HARD RULES
- NEVER restart, stop, recreate, or exec a mutating command in ANY live/prod container, on any host. Read-only `docker exec ... sha256sum|ffprobe|cat` and `docker ps/inspect` on live containers are allowed.
- Container restart allowlist starts EMPTY. Before you start/stop/restart a SANDBOX container, add it to your allowlist and log "ELEVATED <name>" in actions_taken; verify healthy after; list it in containers_restarted. Only `media-pipeline-sandbox-*` containers on server-01 may be elevated. Throwaway `docker run --rm` containers you create are fine.
- NO git add/commit/push. NO writes to the primary host's Postgres. NO edits to the LIVE media-transcoder.py (/opt/appdata/docker/docker-compose/media-transcoder/). NO promotion — verdict only; the main session promotes.
- Issue every mutating command as its OWN tool call; verify with a SEPARATE call.
- WAITING: never hand-write nested-quote wait loops. Write wait conditions as small script FILES on the target host (SQL in a .sql file run with `psql -f`), run with `timeout 900`. Before wrap-up confirm no leftover waiters on both hosts (`ps -eo pid,args | grep "[t]imeout"`; kill only your own). Do NOT leave long-running background jobs and go idle — run each batch in the foreground with a timeout.
- Mask secrets. No sudo. If a security hook BLOCKS an action, do not work around it: partial wrap-up.
- Do not make design decisions beyond the pre-made ones below. If reality contradicts them (e.g. the D1 build still fails M1 end-to-end), STOP and report with evidence.
ACCESS MAP
- server-01: ssh administrator@192.168.1.90 (batch read commands per ssh call).
- Sandbox dir (server-01, and a git-tracked twin on the primary host at the same path): /opt/appdata/docker/docker-compose/server-01/media-pipeline-sandbox/
- Sandbox containers: media-pipeline-sandbox-media-transcoder-sandbox-1, media-pipeline-sandbox-postgres-1 (DB media_pipeline_sandbox, user sandbox, no password via docker exec).
- Baseline reset: `TRUNCATE pipeline_learning, transcode_jobs, transcode_overrides, pipeline_events, download_jobs, rename_jobs, show_configs RESTART IDENTITY;` then `docker exec -i media-pipeline-sandbox-postgres-1 psql -U sandbox -d media_pipeline_sandbox < init/02-seed.sql`; ALWAYS restart the sandbox transcoder after a reset (done-set is cached at startup). Also clear the transcoder's persistent done-state under its tmp root (find it: grep TranscodeState / state file path in the script) if a file would be skipped.
- sandbox-output/, sandbox-renamed/, sandbox-tmp/ hold ROOT-owned files: clear them from inside the container (`docker exec media-pipeline-sandbox-media-transcoder-sandbox-1 find <in-container path> -mindepth 1 -delete`; paths via docker inspect).
- STAGE = cp a fixture into sandbox-input/<CATEGORY DIR>/<Show>/ — category dirs: "H264 ENGLISH DUBBED ANIME TO BE TRANSCODED", "H264 ENGLISH SUBBED ANIME TO BE TRANSCODED", "H264 LIVE ACTION SERIES TO BE TRANSCODED", "H264 MOVIES TO BE TRANSCODED". sandbox-input/ must hold only the 4 empty category dirs at the end.
- Sandbox = software libx265 (NVIDIA host, no VAAPI) — expected; the copy path uses no encoder.
- Cut/remux fixtures with a throwaway container: `docker run --rm --entrypoint ffmpeg -v <dir>:/w <sandbox transcoder image>` — mount at /w, NEVER /lib.
- Script workflow: the primary working copy /opt/appdata/docker/docker-compose/server-01/media-pipeline-sandbox/media-transcoder.py is the MP-15 v2 build (sha256 starts d1d19b02 — record the full sha as BASELINE; it = live 7617fc0a + the approved-but-unpromoted MP-15 v2 picker). Build your fix ON TOP of it (both ship together). Edit; `python3 -m py_compile`; scp to server-01; restart the sandbox transcoder; confirm the in-container sha256.
- Pre-R2 reference code (for the step-1 comparison ONLY, never deployed to the sandbox container): `git -C /opt/appdata/docker show 249bbe8:docker-compose/media-transcoder/media-transcoder.py` (sha e3b6f751; MP-10 shadow + AV-FIX, before R2/MP-18/VAAPI). Live = 7617fc0a (commit 8e62433).
- Logs: `ssh administrator@192.168.1.90 "docker logs --since <ts> media-pipeline-sandbox-media-transcoder-sandbox-1 2>&1 | grep -E '<pattern>'"`.
- Unit tests next to the working copy on primary: test_vaapi_cmd.py, test_verify_output_streams.py, test_av_stage2_suppression.py, test_mp19b_resume.py, test_mp15_picker.py (111 total today). All must still pass, plus your new test file.
BACKGROUND (main session — do not re-diagnose beyond step 2)
- R2 (MP-19b) pipeline: per-part video-only encodes cut at keyframes (`.tmp` → rename) into `job_r_<sha1>` dirs; one full audio encode (audio.mka); final mux `build_mux_cmd()` (~L2227) = concat demuxer over the parts with option-B `duration` directives (`concat_durations()` ~L2803, `write_concat_list()` ~L2815) + audio.mka + subs/metadata/chapters from the source, all `-c copy`. Assembly call site ~L3909–3912.
- HEVC sources take the copy branch (~L2035–2048): `plan["encoder"]="copy"`, `plan["video_args"]=["-c:v","copy"]`. R2 already makes exactly ONE part for them (`seg_len = 0 if source_is_hevc`, ~L3757–3760) and the concat list has no duration directive.
- ROOT CAUSE (MP-20 agent, proven 2026-09-25 by manual runs on M1 in a throwaway container): M1 (a stream-copy cut from t=300) starts with open-GOP leading pictures — first keyframe pts 0.082, next packets pts 0.000. The intermediate video-only part (`ffmpeg -i src -map 0:v:0 -an -sn -dn -c:v copy -f matroska v_0000.tmp.mkv`) shifts the keyframe to 0, pushing those leading pictures negative; they are written with pts=N/A (3 packets), and the later `-c copy` mux stops at the first of them → 1 packet, 9.286 s. Results: current concat form FAIL; part rebuilt with `-fflags +genpts` FAIL; part mapped directly without concat FAIL; **video mapped straight from the source in the final mux OK (92.59 s, 2218/2218 packets)**; part built with -copyts OK; pre-R2 249bbe8 segment-muxer form OK. It is an R2 regression. MP-20 fixture H1 (starts on a keyframe) passes on current code: 420.127 s vs 421.127 s source.
- The failed assembly left a partial (9.3 s) file AT the final output path; the next scan's secondary guard `output_already_exists()` (~L4724, call ~L5153–5157) only checks size > 0, so it logged `[SKIP] Output already exists (marking done)` → a truncated file is treated as finished. Output is written straight to its final `.mkv` name, which filebot-monitor could also pick up mid-write (its watcher filters by suffix: VIDEO_EXTENSIONS = .mkv .mp4 .avi .m4v .mov .wmv .flv .webm .mpg .mpeg — any other suffix is ignored).
- Prod exposure is low (0 copy-path jobs in 90 days, 0 jobs since R2 went live), but any HEVC download will hit it.
PRE-MADE DECISIONS (binding; main session 2026-09-25)
- D1 (REVISED) Copy path = NO intermediate video part. When `plan["encoder"] == "copy"`: skip the video part encode entirely (no v_*.mkv, no concat list) and in `build_mux_cmd()` map the video straight from the SOURCE input (`-map 2:v:0`, the source is already input 2) with the audio.mka as input 1 — drop the concat input for this path and renumber inputs cleanly (e.g. `[audio.mka, source]` → `-map 1:v:0 -map 0:a:0`, subs/metadata/chapters from the source). Everything else (-c copy, dispositions, stream order v,a,s) unchanged. Resume/state: the copy job has no parts; make sure plan-hash, has_resume_state, part-count logging and the "all parts done" check handle zero parts without crashing or re-looping (read those code paths first; keep the change minimal and say what you touched). Rationale: the video never goes through an intermediate file, so leading pictures keep their source timestamps (proven: variant D). Do NOT use -copyts.
- D2 Atomic output: assemble into `<output_path>.partial` (a suffix filebot ignores), then `os.replace()` to the final name ONLY after the mux exits 0 and the completeness check (D3) passes. On any failure delete the `.partial` and leave the job in its failed state (existing failure handling/logging); never leave a file at the final name. Applies to ALL encoders, not just copy.
- D3 Completeness check helper `output_is_complete(output_path, source_duration)`: ffprobe must read the file, it must have ≥1 video stream, and its format duration must be within max(2.0 s, 1 %) of the source duration (use the source duration the job already probed — do not add a second source probe if one exists). Used by D2 before the rename AND by `output_already_exists()`.
- D4 Skip guard: the SECONDARY guard (`output_already_exists()` call for NOT-done inputs, ~L5153–5157) returns True only if `output_is_complete()` passes (probe the source duration there; if the SOURCE cannot be probed, keep today's size>0 behaviour and log WARNING). An incomplete file is renamed to `<name>.incomplete-<YYYYmmdd-HHMMSS>` (suffix filebot ignores; never deleted — owner inspects), logged at WARNING `[SKIP-GUARD] Incomplete output quarantined: …`, and the input is queued normally. The requeue PEEK for already-done inputs (~L5118–5122, `if not output_already_exists(peek_outp)`) runs every scan for every done input: it must stay a CHEAP existence check (today's size>0 logic, no ffprobe, never quarantines). Implement this by keeping a cheap helper (e.g. `output_present()`) for the peek and using the completeness check only in the secondary guard.
- D5 Constants OUTPUT_DURATION_TOL_S = 2.0, OUTPUT_DURATION_TOL_PCT = 0.01 at module level with a comment citing MP-20. Update the function docstrings + module fix history (MP-20 entry). Existing log wording unchanged elsewhere.
FIXTURES (server-01 sandbox-fixtures/)
- mp10/"MP10 M1 (1993).mkv" — HEVC movie, 92.59 s, 2218 video packets, open-GOP leading pictures: the FAILING case.
- mp20/ (already created by MP-20, keep): "MP20 H1 - S01E01.mkv" (421.127 s HEVC 10-bit Bookworm S01E27 copy, 2x opus jpn/eng, 2x ASS incl. forced "Signs & Songs" — the PASS case), "MP20 T1 - S01E02.mkv" (first 10 s of H1, video+audio only, 10.225 s, root-owned — for the D4 guard test).
- Regression fixtures (H.264, encode path, already exist): mp18/ progressive + interlaced clips, mp19/ join clip (read names with ls).
STEPS
0 Baseline: BASELINE sha (full, expect d1d19b02…) of the working copy; confirm the sandbox container runs the same hash; live container sha = 7617fc0a (read-only); sandbox-input empty; fixtures present (ffprobe M1, H1, T1); reset DB; restart transcoder (ELEVATE). Do NOT re-run the MP-20 reproduction.
1 Read the code paths D1 changes (part planning ~L3750–3910, has_resume_state, compute_plan_hash, build_mux_cmd, the assembly call site) and list in notes exactly which functions you will edit.
2 Implement D1–D5. Add `test_mp20_copy_assembly.py` (no ffmpeg needed — patch subprocess/ffprobe helpers): copy plan → no video part, mux cmd maps video from the source input and has no concat input; encode plan → unchanged multi-part concat cmd (compare against the current build_mux_cmd output); zero-part copy state → resume/hash code does not crash; assembly writes `.partial` then renames only on success; failure → no file at the final name; `output_is_complete` tolerance edges (source 1400 s: 1398.1 ok, 1385.9 fails; source 60 s: 58.1 ok, 57.9 fails); secondary guard with complete / truncated (→ quarantined, returns False) / unprobeable-source (size>0 fallback) files; the done-input peek never calls ffprobe and never quarantines. Run ALL unit test files — all pass. py_compile; scp; restart (ELEVATE); confirm in-container hash.
3 Sandbox transcodes (end-to-end), one category batch at a time:
a) M1 (movies, HEVC, the failing case): job done; output duration within tolerance of 92.59 s; video codec hevc with 2218 ±1 packets (`ffprobe -count_packets`); forced sub = "English Forced"; `[VERIFY]` no fail; no `.partial` left.
b) H1 (dubbed, HEVC): same checks vs 421.127 s; eng dub audio; forced+default "Signs & Songs". Also run `tools/av_align.py` (sandbox dir) on H1 output vs source and record head_frame_offset + max drift — MP-20 saw the A/V detector apply a −994 ms correction to this clean BD source; report whether the output is actually out of sync (measure only, do NOT change A/V logic).
c) T1 skip-guard: after H1's output exists, REPLACE it (from inside the container) with T1 under H1's output name, clear the done-state so the secondary guard is exercised, restart; EXPECT `[SKIP-GUARD] Incomplete output quarantined`, a `.incomplete-*` file, and H1 re-transcoded to a complete output.
d) Regression (encode path): one mp18 interlaced clip + the mp19 join clip (read names with ls) — job done, output duration within tolerance, parts > 1 where expected, `.partial` absent, [VERIFY] no fail. av_align on the mp19 output vs its source: head_frame_offset + max drift (expect ≤ ±1 frame / ≤ 40 ms, as before).
Un-stage/clear between batches; reset DB + restart if the done-set would skip a file.
4 Verdict: GO iff steps 2–3 meet every expectation (3b's av_align result is informational, not a gate) and all unit tests pass. NO promotion. On NO-GO give the failing expectation with evidence.
5 Clean up: sandbox-input back to 4 empty dirs; clear output/renamed/tmp (including any .incomplete-*); reset DB; sandbox transcoder RUNNING idle on the new working copy; keep sandbox-fixtures/mp20/*; no leftover waiters on either host; delete your temp scripts.
SCOPE ALLOWLIST: primary working copy server-01/media-pipeline-sandbox/media-transcoder.py + new test_mp20_copy_assembly.py next to it; server-01 sandbox dir contents (script copy, sandbox-input/output/renamed/tmp; sandbox-fixtures read-only) + sandbox DB; your temp scripts; the context.md append. Nothing else.
PERSIST BEFORE YOU FINISH
- Append (one `cat >>` heredoc; append only) "## MP-20b HEVC copy direct-map — 2026-09-25 (background agent)" to /opt/appdata/docker/docker-compose/media-downloader-local/.claude/context.md: What was done / Decisions / Current state / Next step.
- Run: python3 /opt/appdata/docker/.claude/scripts/embed_memory_dir.py --only-recent 5
- FINAL message = wrap-up JSON only, always:
{"status":"succeeded|partially_succeeded|failed","project":"media_pipeline","phase":"MP-20b","actions_taken":[],"actions_failed":[],"files_touched":[],"containers_restarted":[],"unverified_claims":[],"baseline_sha256":"","new_sha256":"","functions_edited":[],"transcode_results":{"M1":{},"H1":{},"T1_guard":{},"regress_mp18":{},"regress_mp19":{}},"av_align":{"H1":{"head_frame_offset":null,"max_drift_ms":null,"detector_correction_ms":null},"mp19":{"head_frame_offset":null,"max_drift_ms":null}},"unit_tests":{"passed":null,"failed":null},"verdict":"GO|NO-GO","diff_summary":"","next_step":"","notes":""}
@@ -0,0 +1,75 @@
# MP-21 — A/V Stage-1 "container offset" shifts audio when a track has a nonzero start_time — diagnose (STOP gate) → fix in sandbox → verdict only
Tracks: personal_projects 276 (+ 61 media_pipeline). Evidence: MP-20b sandbox run 2026-09-25 (fixture mp20/H1: eng dub
opus start_time 1.001 s → Stage 1 applied −994 ms → output eng audio constant ~994 ms EARLY vs the source, no drift).
Main session read the code + queried prod before writing this (lesson from MP-20: read the code path first; unproven
root cause = diagnosis phase with a STOP gate, fixes as options). Written + spawned 2026-09-25 evening.
---
You are a bounded background agent for the media_pipeline project. PHASE MP-21 — prove why the A/V Stage-1 container-offset correction puts audio out of sync, measure the damage on real production outputs (read-only), fix it in the SANDBOX, verify, deliver a verdict. Budget: --max-turns 50 (reason: diagnosis with manual ffmpeg runs, 2 prod read-only measurements, 2 candidate fixes, unit tests, 5 sandbox transcodes with av_align each, prod-impact inventory).
If you issue the same tool call twice with identical arguments, STOP and output the wrap-up block with status=partially_succeeded.
HARD RULES
- NEVER restart, stop, recreate, or exec a mutating command in ANY live/prod container, on any host. Read-only `docker exec ... sha256sum|ffprobe|cat` and `docker ps/inspect` on live containers are allowed.
- PRODUCTION MEDIA IS READ-ONLY: you may ffprobe / read / decode (av_align) files under /media/mediashare and the seedbox WebDAV mount, and COPY a file out to a scratch dir to cut a fixture. Never move, rename, delete, retag or remux IN PLACE any production file. No re-transcoding of prod content — remediation is the owner's decision after your inventory.
- Container restart allowlist starts EMPTY. Before you start/stop/restart a SANDBOX container, add it to your allowlist and log "ELEVATED <name>" in actions_taken; verify healthy after; list it in containers_restarted. Only `media-pipeline-sandbox-*` containers on server-01 may be elevated. Throwaway `docker run --rm` containers you create are fine.
- NO git add/commit/push. Prod Postgres (primary host, DB media_pipeline): SELECT only, never write. NO edits to the LIVE media-transcoder.py (/opt/appdata/docker/docker-compose/media-transcoder/). NO promotion — verdict only.
- Issue every mutating command as its OWN tool call; verify with a SEPARATE call.
- WAITING: never hand-write nested-quote wait loops. Write wait conditions as small script FILES on the target host (SQL in a .sql file run with `psql -f`), run with `timeout 900`. Before wrap-up confirm no leftover waiters on both hosts (`ps -eo pid,args | grep "[t]imeout"`; kill only your own). Do NOT leave long-running background jobs and go idle — run each batch in the foreground with a timeout.
- Mask secrets. No sudo. If a security hook BLOCKS an action, do not work around it: partial wrap-up.
- STOP GATE: if step 2 does NOT confirm that the output's A/V differs from the source as played BECAUSE of the Stage-1 correction (i.e. the mechanism below is wrong), STOP after step 2 and report with evidence. Do not invent a different fix.
- Do not change Stage 2 (content analysis; already suppressed by AV_STAGE2_APPLY=0), the learned-skip heuristic thresholds, or anything outside the A/V Stage-1 path + its DB/learning bookkeeping.
ACCESS MAP
- server-01: ssh administrator@192.168.1.90 (batch read commands per ssh call).
- Sandbox dir (server-01, and a git-tracked twin on the primary host at the same path): /opt/appdata/docker/docker-compose/server-01/media-pipeline-sandbox/
- Sandbox containers: media-pipeline-sandbox-media-transcoder-sandbox-1, media-pipeline-sandbox-postgres-1 (DB media_pipeline_sandbox, user sandbox, no password via docker exec).
- Baseline reset: `TRUNCATE pipeline_learning, transcode_jobs, transcode_overrides, pipeline_events, download_jobs, rename_jobs, show_configs RESTART IDENTITY;` then `docker exec -i media-pipeline-sandbox-postgres-1 psql -U sandbox -d media_pipeline_sandbox < init/02-seed.sql`; ALWAYS restart the sandbox transcoder after a reset (done-set cached at startup). Note the learned-skip corpus lives in pipeline_learning: after a reset every (codec,padding) profile runs full_detection — Stage 1 runs on BOTH paths, so that is fine.
- sandbox-output/, sandbox-renamed/, sandbox-tmp/ hold ROOT-owned files: clear them from inside the container (`docker exec media-pipeline-sandbox-media-transcoder-sandbox-1 find <in-container path> -mindepth 1 -delete`; paths via docker inspect).
- STAGE = cp a fixture into sandbox-input/<CATEGORY DIR>/<Show>/ — category dirs: "H264 ENGLISH DUBBED ANIME TO BE TRANSCODED", "H264 ENGLISH SUBBED ANIME TO BE TRANSCODED", "H264 LIVE ACTION SERIES TO BE TRANSCODED", "H264 MOVIES TO BE TRANSCODED". sandbox-input/ must hold only the 4 empty category dirs at the end.
- Sandbox = software libx265 (NVIDIA host, no VAAPI) — expected; this fix is encoder-independent.
- Throwaway ffmpeg/ffprobe: `docker run --rm --entrypoint ffmpeg -v <dir>:/w <sandbox transcoder image>` (mount at /w, NEVER /lib). No host ffmpeg on primary. The same image can run on the primary host for prod read-only measurements (mount prod dirs :ro).
- Script workflow: primary working copy /opt/appdata/docker/docker-compose/server-01/media-pipeline-sandbox/media-transcoder.py = 776041cf… (= live 7617fc0a + MP-15 v2 picker + MP-20b HEVC-copy/.partial/skip-guard, all GO, not yet promoted — record the full sha as BASELINE). Build ON TOP of it. Edit; `python3 -m py_compile`; scp to server-01; restart the sandbox transcoder; confirm the in-container sha256.
- A/V measurement tool: `server-01/media-pipeline-sandbox/tools/av_align.py` (read it first: how it picks the audio stream and whether it measures on the PRESENTED timeline, i.e. honours stream start_time). It needs matching audio content: when the source's a:0 is a different language than the output's audio, measure against a throwaway remux of the source's matching track (`-map 0:v:0 -map 0:a:<n> -c copy`) — MP-20b did this; its source-vs-output H1 result: av_rel_ms −994/−995 at t=30,120,200,300,380, head_frame_offset 0.
- Logs: `ssh administrator@192.168.1.90 "docker logs --since <ts> media-pipeline-sandbox-media-transcoder-sandbox-1 2>&1 | grep -E '<pattern>'"`.
- Prod DB (read-only): `PG=$(docker ps --format '{{.Names}}' | grep '^postgres-'); docker exec "$PG" psql -U postgres -d media_pipeline -c "<one SELECT>"`. transcode_jobs columns: id, playback_verified, av_offset_ms (APPLIED), av_sync_confidence, audio_initial_padding_samples, audio_start_time_offset_ms (Stage-1 measured), deletion_eligible_after, source_deleted, queued_at, finished_at, input_path, output_path, status, encoder, detection_method, audio_codec, error_msg. Do not guess other column names.
- Unit tests next to the working copy on primary: test_vaapi_cmd.py, test_verify_output_streams.py, test_av_stage2_suppression.py, test_mp19b_resume.py, test_mp15_picker.py, test_mp20_copy_assembly.py (128 total). All must still pass, plus your new test file.
BACKGROUND (read by the main session 2026-09-25 — line numbers approximate, verify)
- `_detect_container_av_offset(streams, audio_stream_idx)` (~L655): offset_ms = (video.start_time − audio.start_time) × 1000; comment says "Positive: audio ahead of video (adelay needed); Negative: audio behind video (trim needed)".
- Applied on BOTH detection paths when |offset| ≥ AV_CONTAINER_THRESHOLD_MS: learned-skip path in `_detect_audio_delay()` (~L3493–3612, `audio_delay_ms = _av_container_off`), and `detect_av_sync_offset()` Stage 1 (~L717–765, returns container_off before Stage 2).
- The correction is applied in `audio_correction_args()` (~L2226): >0 → `adelay=<ms>:all=1`; <0 → `atrim=start=<|ms|/1000>,asetpts=PTS-STARTPTS`, inside `build_audio_cmd()` (~L2237: `ffmpeg -i src -map <audio_map> -vn -sn -dn <audio_args> <correction> -f matroska audio.mka`). audio.mka is then muxed -c copy with the video (concat of parts for the encode path; straight from the source for the HEVC copy path since MP-20b).
- SUSPECTED MECHANISM (unproven): ffmpeg keeps the audio's source timestamps (first pts ≈ 1.001 s) in audio.mka, so the source's start_time offset is ALREADY preserved through the pipeline; the negative-branch `atrim=start=1.0` trims ~nothing (the audio frames' pts are already ≥ 1.0) and `asetpts=PTS-STARTPTS` then re-bases the audio to 0 → audio plays ~1 s early. I.e. Stage 1 "corrects" an offset the pipeline never loses. The positive branch (adelay) may have the mirror-image problem (double delay) — untested. Whether the VIDEO side also keeps its source start_time (encode path: concat of parts; copy path: straight from source) is unknown and matters for the fix.
- PROD IMPACT (main session SELECT 2026-09-25, 752 jobs with detection data): applied |av_offset_ms| ≥ 400 on ~53 jobs — Jujutsu Kaisen (aac, avg −1034 ms, 23 eps, 2026-08-09), Assassination Classroom (opus, avg −672, 14, 2026-04-24), Ascendance of a Bookworm (opus, −994, 10, 2026-04-24; e.g. output_path ".../Media to be Renamed/Anime/English Dubbed Anime/Ascendance of a Bookworm/Ascendance.of.a.Bookworm.S01E31.1080p.Bluray.Opus2.0.10bit.x265-Headpatter.mkv" — renamed since; find the current file read-only under "/media/mediashare/MY MEDIA/MY MEDIA/" e.g. the ANIME tree), Yu Yu Hakusho The Movie (aac, −1900, 1), plus 4 positive singles (Slime +2120, The Pitt +720, Farming Life +820, one .ffprobe_tmp +1840). Sources were deleted for most (source_deleted=t); the seedbox mount "/media/mediashare/MY MEDIA/MY MEDIA/seedbox local/webdav/Finished/Qbittorrent/" still had Bookworm S01E27 (MP-20 cut H1 from it).
CANDIDATE FIXES (to TEST in step 3 — the criterion picks, not preference)
- C1 "Stage 1 = record only": keep measuring the container offset (DB column audio_start_time_offset_ms, log line), apply 0 on both paths, because the pipeline preserves stream start times. Learning outcome follows the APPLIED value (as Stage-2 suppression does). Only valid if the output keeps BOTH the audio and the video source start times on both paths (encode + HEVC copy).
- C2 "timeline-correct correction": keep applying Stage 1, but make audio.mka's timeline match the output video's timeline exactly (e.g. re-base audio to the video's first pts, no asetpts to 0 when the pts are already correct). Use only if C1 fails the criterion on some path.
- CRITERION (both paths): for every fixture, output av_rel_ms within ±40 ms of the SOURCE AS PLAYED at ≥4 points and head_frame_offset 0 (av_align), and output duration within MP-20 tolerance. If both pass, choose C1 (simpler, no correction to get wrong). If neither passes, STOP and report.
FIXTURES
- server-01 sandbox-fixtures/mp20/"MP20 H1 - S01E01.mkv": HEVC copy path; eng opus start_time 1.001 s (the case that failed). a:0 is jpn — measure against an eng-track remux.
- CREATE sandbox-fixtures/mp21/ (throwaway container, from existing sandbox fixtures; record exact commands):
- "MP21 N1 - S01E01.mkv": H.264 (encode path), audio start_time +1.0 s later than video — from the mp19 join clip: `-i clip -itsoffset 1.0 -i clip -map 0:v:0 -map 1:a:0 -c copy` (then ffprobe: audio start_time ≈ 1.0, video ≈ 0).
- "MP21 P1 - S01E02.mkv": H.264, VIDEO start_time +0.7 s later than audio (positive Stage-1 case) — same method with the offset on the video input.
- "MP21 Z1 - S01E03.mkv": the unmodified mp19 join clip copy (control: offset 0 → no correction; must be unchanged by the fix).
- Prod read-only measurement pair: the current (renamed) prod output of Ascendance of a Bookworm S01E27 if one exists (else another Bookworm episode listed above with its seedbox source) + its seedbox source. Copy nothing into prod dirs.
STEPS
0 Baseline: BASELINE sha (full, expect 776041cf…) = working copy = sandbox in-container; live = 7617fc0a (read-only); sandbox-input empty; create mp21/ fixtures + ffprobe start_times; reset DB; restart transcoder (ELEVATE).
1 Validate the tool: av_align of each fixture against ITSELF (and H1 vs its eng-track remux) → expect av_rel ≈ 0, head_frame_offset 0; state in notes whether av_align honours start_time. If av_align cannot measure presented-timeline offsets correctly, STOP (tool problem, report).
2 DIAGNOSE (STOP gate): on CURRENT code, transcode N1 then P1 (one batch each; record the Stage-1 log line + applied ms, the audio command from the log, ffprobe start_time/first pts of audio.mka if still in the job dir or reproduce the audio cmd by hand in a throwaway container, and output stream start_times). av_align output vs source. Also run by hand (throwaway container) the H1 audio command with and without the correction and show first pts. EXPECT: N1 output audio ~1 s early; P1 output off by ~0.7 s (report direction); and the audio.mka pts evidence explains it. Then the PROD read-only measurement: av_align the prod Bookworm output vs its seedbox source (matching track) and record av_rel_ms. If prod shows ~−1 s → the bug is live in the library; say so.
3 Implement C1 (and C2 only if C1 fails the criterion); unit tests `test_mp21_av_stage1.py` (no ffmpeg): C1 → applied 0 on BOTH paths for ±offsets, container offset still logged + passed to the DB writer as audio_start_time_offset_ms, learning outcome "clean" when applied 0; audio_correction_args unchanged for explicit Stage-2/forced values (FORCE_AUDIO_DELAY_MS sandbox switch still works); Z1-style 0 offset unchanged. Update docstrings (+ the misleading sign comment) and the module fix history (MP-21). Run ALL unit test files — all pass. py_compile; scp; restart (ELEVATE); confirm hash.
4 Sandbox transcodes on the fixed code, one batch each, av_align every output vs its source (matching track): N1, P1, Z1 (encode path), H1 (HEVC copy path). CRITERION above for all four; Z1 and H1 durations/streams otherwise identical to the MP-20b results (H1: 10072 video pkts, forced "Signs & Songs [EMBER]"). Record per fixture: applied ms, av_rel_ms list, head_frame_offset, verify result.
5 PROD IMPACT INVENTORY (read-only): one SELECT listing every prod job with |av_offset_ms| ≥ AV_CONTAINER_THRESHOLD_MS (id, input_path, output_path, av_offset_ms, audio_start_time_offset_ms, detection_method, finished_at, source_deleted); for each, check read-only whether a source still exists on the seedbox mount (same filename search) and whether the output still sits in "Media to be Renamed" or was renamed (best-effort filename search). Write it to /opt/appdata/docker/docker-compose/media-transcoder/design/MP21_av_stage1_prod_impact.csv (gitignored-style data file, fine to create) + a short markdown summary /opt/appdata/docker/docker-compose/media-transcoder/design/MP21_av_stage1_findings.md (mechanism, evidence, chosen fix, prod counts: affected / source-still-on-seedbox / needs re-download). No remediation.
6 Verdict: GO iff step 2 confirmed the mechanism, step 4 meets the criterion for all 4 fixtures, and all unit tests pass. NO promotion. On NO-GO give the failing expectation with evidence.
7 Clean up: sandbox-input back to 4 empty dirs; clear output/renamed/tmp; reset DB; sandbox transcoder RUNNING idle on the new working copy; keep sandbox-fixtures/mp21/* and mp20/*; remove any scratch copies of prod files and throwaway remuxes; no leftover waiters on either host; delete your temp scripts.
SCOPE ALLOWLIST: primary working copy server-01/media-pipeline-sandbox/media-transcoder.py + new test_mp21_av_stage1.py next to it; server-01 sandbox dir contents (script copy, sandbox-input/output/renamed/tmp, sandbox-fixtures/mp21/) + sandbox DB; the two new files in media-transcoder/design/ named above; scratch dirs you create and delete; your temp scripts; the context.md append. Nothing else. Production media and the prod DB are read-only.
PERSIST BEFORE YOU FINISH
- Append (one `cat >>` heredoc; append only) "## MP-21 A/V Stage-1 start_time — 2026-09-25 (background agent)" to /opt/appdata/docker/docker-compose/media-downloader-local/.claude/context.md: What was done / Decisions / Current state / Next step.
- Run: python3 /opt/appdata/docker/.claude/scripts/embed_memory_dir.py --only-recent 5
- FINAL message = wrap-up JSON only, always:
{"status":"succeeded|partially_succeeded|failed","project":"media_pipeline","phase":"MP-21","actions_taken":[],"actions_failed":[],"files_touched":[],"containers_restarted":[],"unverified_claims":[],"baseline_sha256":"","new_sha256":"","av_align_honours_start_time":null,"mechanism_confirmed":null,"mechanism":"","diagnosis":{"N1":{},"P1":{},"H1_manual":{},"prod_bookworm":{"file":"","av_rel_ms":[],"head_frame_offset":null}},"chosen_fix":"C1|C2|none","functions_edited":[],"results":{"N1":{},"P1":{},"Z1":{},"H1":{}},"prod_impact":{"jobs_affected":null,"source_on_seedbox":null,"source_gone":null,"csv":"","findings_md":""},"unit_tests":{"passed":null,"failed":null},"verdict":"GO|NO-GO","diff_summary":"","next_step":"","notes":""}
@@ -0,0 +1,84 @@
# MP-21b — keep each stream's source start offset through the final mux + make Stage 1 record-only — sandbox fix → verdict only
Tracks: personal_projects 276 (+ 61 media_pipeline). Evidence: MP-20b sandbox run 2026-09-25 (fixture mp20/H1: eng dub
opus start_time 1.001 s → Stage 1 applied −994 ms → output eng audio constant ~994 ms EARLY vs the source, no drift).
Supersedes MP-21 (NO-GO at its STOP gate 2026-09-25; its measured diagnosis is inlined below as FACT). Written + spawned
2026-09-25 late evening.
---
You are a bounded background agent for the media_pipeline project. PHASE MP-21b — make the output keep every stream's source start offset (audio included) and stop the Stage-1 correction from double-counting, in the SANDBOX; verify on 7 fixtures across both mux paths; inventory the affected prod jobs (read-only); deliver a verdict. Budget: --max-turns 55 (reason: 2 mux candidates tested by hand, unit tests, 7 sandbox transcodes with av_align + subtitle-timing checks, prod inventory).
If you issue the same tool call twice with identical arguments, STOP and output the wrap-up block with status=partially_succeeded.
HARD RULES
- NEVER restart, stop, recreate, or exec a mutating command in ANY live/prod container, on any host. Read-only `docker exec ... sha256sum|ffprobe|cat` and `docker ps/inspect` on live containers are allowed.
- PRODUCTION MEDIA IS READ-ONLY: you may ffprobe / read / decode (av_align) files under /media/mediashare and the seedbox WebDAV mount, and COPY a file out to a scratch dir to cut a fixture. Never move, rename, delete, retag or remux IN PLACE any production file. No re-transcoding of prod content — remediation is the owner's decision after your inventory.
- Container restart allowlist starts EMPTY. Before you start/stop/restart a SANDBOX container, add it to your allowlist and log "ELEVATED <name>" in actions_taken; verify healthy after; list it in containers_restarted. Only `media-pipeline-sandbox-*` containers on server-01 may be elevated. Throwaway `docker run --rm` containers you create are fine.
- SANDBOX RESTART BLOCKED? (MP-21's restart was denied by the auto-mode classifier.) Do NOT retry or work around. Use the OWNER-COMMAND relay: end your run with a final message headed exactly `⏸ OWNER-COMMAND-REQUEST` with fields `host:` (server-01 — ssh administrator@192.168.1.90), `command:` (exact, e.g. `docker restart media-pipeline-sandbox-media-transcoder-sandbox-1`), `why:`, `paste_back:` (the output + `docker ps --filter name=media-pipeline-sandbox-media-transcoder-sandbox-1` status), `resume_at:` (step number). Then stop; the main session relays and resumes you with the output. Max 3 relays per run; batch work so one restart covers as much as possible.
- NO git add/commit/push. Prod Postgres (primary host, DB media_pipeline): SELECT only, never write. NO edits to the LIVE media-transcoder.py (/opt/appdata/docker/docker-compose/media-transcoder/). NO promotion — verdict only.
- Issue every mutating command as its OWN tool call; verify with a SEPARATE call.
- WAITING: never hand-write nested-quote wait loops. Write wait conditions as small script FILES on the target host (SQL in a .sql file run with `psql -f`), run with `timeout 900`. Before wrap-up confirm no leftover waiters on both hosts (`ps -eo pid,args | grep "[t]imeout"`; kill only your own). Do NOT leave long-running background jobs and go idle — run each batch in the foreground with a timeout.
- Mask secrets. No sudo. If a security hook BLOCKS an action, do not work around it: partial wrap-up.
- STOP GATE: if in step 2 NEITHER mux candidate passes the CRITERION on the hand-built runs, STOP after step 2 and report with evidence. Do not invent a third design.
- Do not change Stage 2 (content analysis; already suppressed by AV_STAGE2_APPLY=0), the learned-skip heuristic thresholds, or anything outside the A/V Stage-1 path + its DB/learning bookkeeping.
ACCESS MAP
- server-01: ssh administrator@192.168.1.90 (batch read commands per ssh call).
- Sandbox dir (server-01, and a git-tracked twin on the primary host at the same path): /opt/appdata/docker/docker-compose/server-01/media-pipeline-sandbox/
- Sandbox containers: media-pipeline-sandbox-media-transcoder-sandbox-1, media-pipeline-sandbox-postgres-1 (DB media_pipeline_sandbox, user sandbox, no password via docker exec).
- Baseline reset: `TRUNCATE pipeline_learning, transcode_jobs, transcode_overrides, pipeline_events, download_jobs, rename_jobs, show_configs RESTART IDENTITY;` then `docker exec -i media-pipeline-sandbox-postgres-1 psql -U sandbox -d media_pipeline_sandbox < init/02-seed.sql`; ALWAYS restart the sandbox transcoder after a reset (done-set cached at startup). Note the learned-skip corpus lives in pipeline_learning: after a reset every (codec,padding) profile runs full_detection — Stage 1 runs on BOTH paths, so that is fine.
- sandbox-output/, sandbox-renamed/, sandbox-tmp/ hold ROOT-owned files: clear them from inside the container (`docker exec media-pipeline-sandbox-media-transcoder-sandbox-1 find <in-container path> -mindepth 1 -delete`; paths via docker inspect).
- STAGE = cp a fixture into sandbox-input/<CATEGORY DIR>/<Show>/ — category dirs: "H264 ENGLISH DUBBED ANIME TO BE TRANSCODED", "H264 ENGLISH SUBBED ANIME TO BE TRANSCODED", "H264 LIVE ACTION SERIES TO BE TRANSCODED", "H264 MOVIES TO BE TRANSCODED". sandbox-input/ must hold only the 4 empty category dirs at the end.
- Sandbox = software libx265 (NVIDIA host, no VAAPI) — expected; this fix is encoder-independent.
- Throwaway ffmpeg/ffprobe: `docker run --rm --entrypoint ffmpeg -v <dir>:/w <sandbox transcoder image>` (mount at /w, NEVER /lib). No host ffmpeg on primary. The same image can run on the primary host for prod read-only measurements (mount prod dirs :ro).
- Script workflow: primary working copy /opt/appdata/docker/docker-compose/server-01/media-pipeline-sandbox/media-transcoder.py = 776041cf… (= live 7617fc0a + MP-15 v2 picker + MP-20b HEVC-copy/.partial/skip-guard, all GO, not yet promoted — record the full sha as BASELINE). Build ON TOP of it. Edit; `python3 -m py_compile`; scp to server-01; restart the sandbox transcoder; confirm the in-container sha256.
- A/V measurement tool: `server-01/media-pipeline-sandbox/tools/av_align.py` (read it first: how it picks the audio stream and whether it measures on the PRESENTED timeline, i.e. honours stream start_time). It needs matching audio content: when the source's a:0 is a different language than the output's audio, measure against a throwaway remux of the source's matching track (`-map 0:v:0 -map 0:a:<n> -c copy`) — MP-20b did this; its source-vs-output H1 result: av_rel_ms −994/−995 at t=30,120,200,300,380, head_frame_offset 0.
- Logs: `ssh administrator@192.168.1.90 "docker logs --since <ts> media-pipeline-sandbox-media-transcoder-sandbox-1 2>&1 | grep -E '<pattern>'"`.
- Prod DB (read-only): `PG=$(docker ps --format '{{.Names}}' | grep '^postgres-'); docker exec "$PG" psql -U postgres -d media_pipeline -c "<one SELECT>"`. transcode_jobs columns: id, playback_verified, av_offset_ms (APPLIED), av_sync_confidence, audio_initial_padding_samples, audio_start_time_offset_ms (Stage-1 measured), deletion_eligible_after, source_deleted, queued_at, finished_at, input_path, output_path, status, encoder, detection_method, audio_codec, error_msg. Do not guess other column names.
- Unit tests next to the working copy on primary: test_vaapi_cmd.py, test_verify_output_streams.py, test_av_stage2_suppression.py, test_mp19b_resume.py, test_mp15_picker.py, test_mp20_copy_assembly.py (128 total). All must still pass, plus your new test file.
FACTS (measured by MP-21, 2026-09-25 — do not re-diagnose)
- The A/V tool is sound: `tools/av_align.py` measures on the PRESENTED timeline (self-comparisons ±4 ms; Z1-vs-N1 audio_lag +1000, Z1-vs-P1 video_lag +710). head_frame_offset is frame-index based (fps filter pads the start), so judge start shifts by av_rel/audio_lag/video_lag, not head offset. The transcoder image's ffmpeg SIGSEGVs decoding some libx265 sandbox outputs at certain seek points (N1 output t=119) — pick other points; not in scope.
- ROOT CAUSE, NEGATIVE case (audio starts later than video, e.g. N1 audio 1.0 s, H1 eng 1.001 s, prod Bookworm): audio.mka DOES keep the source audio start (H1 no correction: audio.mka start_time 1.002). The final `-c copy` mux (`build_mux_cmd()`, no -copyts) re-bases EACH INPUT by its own start_time, so audio.mka's 1.0 s is dropped → output audio start 0 → audio ~1 s early. This happens with or without the Stage-1 correction (its `atrim=start=…,asetpts=PTS-STARTPTS` is effectively a no-op). Hand-built H1 copy mux WITH -copyts: audio start 1.002, av_rel −1…+0.5 ms (fixed, copy path only; encode path untested).
- ROOT CAUSE, POSITIVE case (video starts later, P1 video 0.7 s): the encode path already reproduces the video's 0.7 s start (the -r CFR padding; output video matches the source's presented timeline, video_lag ≈ 0), so Stage 1's `adelay=700ms` double-counts → audio +707 ms late. Stage 1 must become record-only.
- PROD: Bookworm S01E27 → library "ANIME/ENGLISH DUBBED ANIME/Ascendance of a Bookworm/Season 03/Ascendance of a Bookworm - S03E01 - The Beginning of Winter.mkv" (prod job 1805, av_offset_ms −994; matched by duration/tracks, not hash): audio ~1 s early (plus pre-R2 video drift 43–376 ms). Main-session SELECT: ~53 jobs with |applied| ≥ 400 ms (Jujutsu Kaisen 23, Assassination Classroom 14, Bookworm 10, YYH movie 1, 4 positive singles) — but the NEGATIVE bug hits EVERY job whose selected audio starts later than the video, applied or not, so the inventory must select on audio_start_time_offset_ms too.
- CODE (working copy, verify line numbers): `_detect_container_av_offset()` ~L655 (offset = (video_start − audio_start) × 1000; its sign comment is misleading); applied in `_detect_audio_delay()` learned-skip path (~L3550) and `detect_av_sync_offset()` Stage 1 (~L750); `audio_correction_args()` ~L2226; `build_audio_cmd()` ~L2237; `build_mux_cmd()` (MP-20 form: encode = [concat, audio.mka, source]; copy = [audio.mka, source], video from the source); `assemble_output()` ~L3683; JobState/`compute_plan_hash()` ~L2297 for resume.
PRE-MADE DECISIONS (binding)
- D1 Stage 1 = RECORD ONLY on both detection paths: still measure + log the container offset and write it to audio_start_time_offset_ms; applied audio_delay_ms from Stage 1 = 0; learning outcome follows the APPLIED value (as Stage-2 suppression does). Fix the misleading sign comment. Keep `audio_correction_args()` + FORCE_AUDIO_DELAY_MS working for explicit values (sandbox test switch).
- D2 The final mux must place audio.mka on the output timeline at the SAME offset the source audio had relative to the source's timeline origin. Test BOTH candidates by hand first (step 2), then implement the one the CRITERION picks:
M-A (preferred if it passes): probe audio.mka's start_time S after the audio encode (store S in the job state so a resumed job reuses it, and include it in nothing that would invalidate old parts needlessly — say what you did) and put `-itsoffset S` immediately before `-i audio.mka` in `build_mux_cmd()` for both paths; skip when |S| < 0.001. Video/subs/metadata inputs untouched.
M-B: `-copyts` on the final mux (proven on the H1 copy path). RISK to test: with a source whose GLOBAL start is non-zero (fixture G1), concat parts start at 0 while subs/source streams keep absolute timestamps → subs/video could separate.
Choose M-A if both pass; if only one passes, that one; if neither, STOP (gate).
- D3 Do not change the video path, part planning, Stage 2, or learned-skip thresholds. Update docstrings + module fix history (MP-21b entry: mechanism, both fixes, fixture evidence).
CRITERION (every fixture, on the path it exercises)
- av_align output vs the source AS PLAYED (matching audio track; eng-track remux where a:0 differs): av_rel within ±40 ms at ≥4 points (or audio_lag/video_lag within ±40 ms where av_rel is null).
- Subtitle timing (fixtures with subs): for the first 5 subtitle packets, (sub pts − first video pts) in the output equals the source's within ±40 ms (`ffprobe -show_packets -select_streams s:0 / v:0 -read_intervals %+30`).
- Output duration within the MP-20 tolerance; video packet count unchanged vs the MP-20b results where known (M1 2218, H1 10072).
FIXTURES (server-01 sandbox-fixtures/)
- mp21/ (exist): N1 (encode, audio +1.0 s), P1 (encode, video +0.7 s), Z1 (encode, control 0/0).
- mp20/ (exist): H1 (copy path, eng a:1 start 1.001 s, ASS subs incl. forced "Signs & Songs [EMBER]").
- mp10/ (exist): "MP10 M1 (1993).mkv" (copy path, open-GOP leading pictures — regression for MP-20b; must stay 2218 packets).
- CREATE in mp21/ (throwaway container; append commands to build_mp21.sh):
- "MP21 HP1 - S01E04.mkv": copy-path POSITIVE case = first 120 s of H1 with its video delayed 0.7 s (`-itsoffset 0.7 -i H1 -i H1 -map 0:v:0 -map 1:a -map 1:s -c copy -t 120`, adjust as needed; ffprobe to confirm video start ≈0.7, audio ≈0 / eng ≈1.0).
- "MP21 G1 - S01E05.mkv": encode-path GLOBAL offset with subs = an existing sandbox fixture that HAS a text subtitle track on the H.264 encode path (e.g. mp10 "MP10 D3 - S01E03.mkv"; read layouts) with ALL streams shifted +1.4 s (`-output_ts_offset 1.4 -map 0 -c copy`); ffprobe: format start ≈1.4.
STEPS
0 Baseline: BASELINE sha (full, expect 776041cf…) = working copy = sandbox in-container (the MP-21 run could not restart after its DB reset: the container is on 776041cf but its cached done-set may hold N1/P1 → reset DB + restart before step 3 anyway); live = 7617fc0a read-only; sandbox-input empty; create HP1 + G1 and ffprobe start_times of all 7 fixtures.
1 Read build_mux_cmd, assemble_output, the audio-encode call site and JobState, and list in notes the functions you will edit.
2 HAND TEST both mux candidates (throwaway containers, from the current code's own audio cmd with NO correction + the current mux cmd plus M-A or M-B): run on N1 (encode), G1 (encode, subs), H1 (copy, subs) — use a real part/concat for the encode path (encode a short part yourself or reuse the transcoder's cmd form). Apply the CRITERION. Record a table candidate × fixture. Choose per D2 or STOP.
3 Implement D1 + the chosen D2 + D3. Unit tests `test_mp21_av_timeline.py` (no ffmpeg): Stage 1 applied 0 for ±offsets on both detection paths, offset still passed to the DB writer, learning outcome "clean"; build_mux_cmd adds the chosen timeline handling on BOTH paths (and nothing when |S| < 1 ms for M-A); resume reuses S; encode/copy cmd otherwise byte-identical to baseline; FORCE_AUDIO_DELAY_MS still yields the correction args. Run ALL unit test files — all pass. py_compile; scp; reset DB; restart (ELEVATE or relay); confirm in-container hash.
4 Sandbox transcodes on the fixed code (batch several fixtures per restart where categories allow; stage episode-style names), CRITERION on every output: N1, P1, Z1, G1 (encode path); H1, HP1, M1 (copy path). Record per fixture: Stage-1 measured ms, applied ms, av_rel/lag list, sub timing deltas, duration, packet count, [VERIFY] result.
5 PROD IMPACT INVENTORY (read-only): one SELECT of every prod job whose selected audio starts later than its video beyond AV_CONTAINER_THRESHOLD_MS (audio_start_time_offset_ms ≤ −threshold) OR had a Stage-1 correction applied (|av_offset_ms| ≥ threshold): id, input_path, output_path, av_offset_ms, audio_start_time_offset_ms, detection_method, finished_at, source_deleted. For each: best-effort read-only filename search for the source on the seedbox mount and for the current library file. Write /opt/appdata/docker/docker-compose/media-transcoder/design/MP21_av_prod_impact.csv + /opt/appdata/docker/docker-compose/media-transcoder/design/MP21_av_findings.md (mechanism, fix, fixture table, prod counts: affected / source on seedbox / source gone, and the recommended remediation order). No remediation.
6 Verdict: GO iff step 2 chose a candidate, step 4 meets the CRITERION on all 7 fixtures, and all unit tests pass. NO promotion. On NO-GO give the failing expectation with evidence.
7 Clean up: sandbox-input back to 4 empty dirs; clear output/renamed/tmp; reset DB; sandbox transcoder RUNNING idle on the new working copy (restart via ELEVATE or relay); keep sandbox-fixtures/mp21/*, mp20/*; remove scratch copies/remuxes; no leftover waiters on either host; delete your temp scripts.
SCOPE ALLOWLIST: primary working copy server-01/media-pipeline-sandbox/media-transcoder.py + new test_mp21_av_timeline.py next to it; server-01 sandbox dir contents (script copy, sandbox-input/output/renamed/tmp, sandbox-fixtures/mp21/, .transcoder-config.json auto-registrations) + sandbox DB; the two new files in media-transcoder/design/ named above; scratch dirs you create and delete; your temp scripts; the context.md append. Nothing else. Production media and the prod DB are read-only.
PERSIST BEFORE YOU FINISH
- Append (one `cat >>` heredoc; append only) "## MP-21b A/V mux timeline — 2026-09-25 (background agent)" to /opt/appdata/docker/docker-compose/media-downloader-local/.claude/context.md: What was done / Decisions / Current state / Next step.
- Run: python3 /opt/appdata/docker/.claude/scripts/embed_memory_dir.py --only-recent 5
- FINAL message = wrap-up JSON only, always:
{"status":"succeeded|partially_succeeded|failed","project":"media_pipeline","phase":"MP-21b","actions_taken":[],"actions_failed":[],"files_touched":[],"containers_restarted":[],"owner_relays":0,"unverified_claims":[],"baseline_sha256":"","new_sha256":"","hand_test":{"M-A":{"N1":{},"G1":{},"H1":{}},"M-B":{"N1":{},"G1":{},"H1":{}}},"chosen_mux":"M-A|M-B|none","functions_edited":[],"results":{"N1":{},"P1":{},"Z1":{},"G1":{},"H1":{},"HP1":{},"M1":{}},"prod_impact":{"jobs_affected":null,"source_on_seedbox":null,"source_gone":null,"csv":"","findings_md":""},"unit_tests":{"passed":null,"failed":null},"verdict":"GO|NO-GO","diff_summary":"","next_step":"","notes":""}