metadata: unify the music pipeline into one per-album process #241

Open
opened 2026-09-17 02:52:01 -04:00 by mysticalsoap · 2 comments
Owner

Problem

Music metadata is handled by four scripts that grew separately: music_audit (detect), music_retag (apply pinned retags), music_tag_sync (dates + comments), album_art_sync (art + the beets import that feeds everything). Four state files, two overrides files, load-bearing execution order in run_all.sh, and cross-script nudges (retag deletes an album's album_art_state entry to force a same-pass re-fetch). It works, but the couplings are invisible and each new concern adds another fragment.

Shape

One per-album pipeline (movie jobs stay separate):

  1. Evaluate the release (today's audit checks: duplicate-inode orphans, sparse identity tags).
  2. Severely mangled + human-pinned mb_release -> dehardlink, full retag, then the normal passes.
  3. Sparse + unpinned -> ntfy once, wait for a pin in the one overrides file.
  4. Every album: normalize DATE to the MB original release date, strip comment noise (today's tag_sync).
  5. Art missing -> non-destructive sidecar fetch (today's album_art_sync).
    One state file, one overrides file, one report. Existing scripts become stages; musiclib.py carries the shared pieces.

Carried constraints

  • Tag writes only after dehardlink; seeds keep their bytes, disk is the cost (standing policy).
  • beets db stores music-root-relative paths; only the beet CLI with config loaded resolves them — the bare library API absolutizes against cwd. Subprocess the CLI.
  • Quiet retag can mis-order tracks on bare cue-split files (the What's Going On overhaul saw beets swap tracks 6/8); identity-tag presence checks don't catch it. The approved-in-principle filename-seeded-tags preprocessing belongs in the retag stage.
  • A retag must refresh the main beets library entry and invalidate that album's art state.
## Problem Music metadata is handled by four scripts that grew separately: music_audit (detect), music_retag (apply pinned retags), music_tag_sync (dates + comments), album_art_sync (art + the beets import that feeds everything). Four state files, two overrides files, load-bearing execution order in run_all.sh, and cross-script nudges (retag deletes an album's album_art_state entry to force a same-pass re-fetch). It works, but the couplings are invisible and each new concern adds another fragment. ## Shape One per-album pipeline (movie jobs stay separate): 1. Evaluate the release (today's audit checks: duplicate-inode orphans, sparse identity tags). 2. Severely mangled + human-pinned `mb_release` -> dehardlink, full retag, then the normal passes. 3. Sparse + unpinned -> ntfy once, wait for a pin in the one overrides file. 4. Every album: normalize DATE to the MB original release date, strip comment noise (today's tag_sync). 5. Art missing -> non-destructive sidecar fetch (today's album_art_sync). One state file, one overrides file, one report. Existing scripts become stages; musiclib.py carries the shared pieces. ## Carried constraints - Tag writes only after dehardlink; seeds keep their bytes, disk is the cost (standing policy). - beets db stores music-root-relative paths; only the beet CLI with config loaded resolves them — the bare library API absolutizes against cwd. Subprocess the CLI. - Quiet retag can mis-order tracks on bare cue-split files (the What's Going On overhaul saw beets swap tracks 6/8); identity-tag presence checks don't catch it. The approved-in-principle filename-seeded-tags preprocessing belongs in the retag stage. - A retag must refresh the main beets library entry and invalidate that album's art state.
Author
Owner

Constraint discovered on the first bulk music_tag_sync run: two instances raced (a hook-triggered dag run vs a manual docker exec) and fought over dehardlink tmps. Per-file collisions are hardened (short pid-unique tmp names), but the unified pipeline should assume a single-writer rule: every run goes through the dagu queue (enqueue, never direct exec), so the queue serializes library writers. Also: editing run_all.sh while a run is active corrupted the in-flight bash read (02:20 run errored) -- the unified single-script entry point removes that hazard too.

Constraint discovered on the first bulk music_tag_sync run: two instances raced (a hook-triggered dag run vs a manual docker exec) and fought over dehardlink tmps. Per-file collisions are hardened (short pid-unique tmp names), but the unified pipeline should assume a single-writer rule: every run goes through the dagu queue (enqueue, never direct exec), so the queue serializes library writers. Also: editing run_all.sh while a run is active corrupted the in-flight bash read (02:20 run errored) -- the unified single-script entry point removes that hazard too.
Author
Owner

Stage: edition trim (from the Marquee Moon / Ants From Up There pass, PR #381)

An album that lands as a bonus-track edition (remaster, deluxe) should be cut down to the pinned release by the pipeline, not by hand. Today's manual sequence, which worked, maps onto one opt-in pin:

"Artist/Album": {"mb_release": "<release mbid>", "trim": true}

trim is explicit — a pin alone must never delete files.

  1. Lidarr release by id. PUT /api/v1/album/{id} with anyReleaseOk: false and releases[].monitored true only for the entry whose foreignReleaseId is the pin. The UI dropdown can't take an id and renders several editions identically (five "8 tracks, United States, 12" Vinyl" for Marquee Moon). Album id comes from GET /api/v1/album?artistId= matched on foreignAlbumId (the release-group) — beets already has that.
  2. Unmonitor the album for the duration. Between the switch and the final rescan Lidarr sees 0/N tracks, and an RSS sync in that window can re-grab the edition we're removing. Re-monitor at the end.
  3. Delete the extras. Resolve the pinned release's track list (ws/2/release/<mbid>?inc=recordings), keep the local files that match by position + title, delete the rest. Seeds are unaffected either way: the library copy is a hardlink or already copy-broken. Pre-retag files have no MB ids, so this is title/number matching — the cue-split caveat from the retag stage applies here too.
  4. Retag (existing stage) — writes the pinned release's ids into the tags.
  5. Lidarr RefreshArtist, then re-monitor. Order is fixed: Lidarr's identifier follows the MB ids in tags, and without them it scores by date/label, picks a different edition and rejects every file with "Album release not requested". Switching the release never remaps on its own, and a rescan before the retag fails the same way.
  6. Verify via album statistics.trackFileCount == trackCount; ntfy on mismatch.

Plumbing: the metadata container has no Lidarr credentials today (the hook goes the other way, Lidarr → Dagu). Step 1/5 need the Lidarr API key as a secret in media/metadata/secrets/, _FILE convention, read inside the container only.

## Stage: edition trim (from the Marquee Moon / Ants From Up There pass, PR #381) An album that lands as a bonus-track edition (remaster, deluxe) should be cut down to the pinned release by the pipeline, not by hand. Today's manual sequence, which worked, maps onto one opt-in pin: ```json "Artist/Album": {"mb_release": "<release mbid>", "trim": true} ``` `trim` is explicit — a pin alone must never delete files. 1. **Lidarr release by id.** `PUT /api/v1/album/{id}` with `anyReleaseOk: false` and `releases[].monitored` true only for the entry whose `foreignReleaseId` is the pin. The UI dropdown can't take an id and renders several editions identically (five "8 tracks, United States, 12" Vinyl" for Marquee Moon). Album id comes from `GET /api/v1/album?artistId=` matched on `foreignAlbumId` (the release-group) — beets already has that. 2. **Unmonitor the album for the duration.** Between the switch and the final rescan Lidarr sees 0/N tracks, and an RSS sync in that window can re-grab the edition we're removing. Re-monitor at the end. 3. **Delete the extras.** Resolve the pinned release's track list (`ws/2/release/<mbid>?inc=recordings`), keep the local files that match by position + title, delete the rest. Seeds are unaffected either way: the library copy is a hardlink or already copy-broken. Pre-retag files have no MB ids, so this is title/number matching — the cue-split caveat from the retag stage applies here too. 4. **Retag** (existing stage) — writes the pinned release's ids into the tags. 5. **Lidarr `RefreshArtist`, then re-monitor.** Order is fixed: Lidarr's identifier follows the MB ids in tags, and without them it scores by date/label, picks a different edition and rejects every file with "Album release not requested". Switching the release never remaps on its own, and a rescan before the retag fails the same way. 6. **Verify** via album `statistics.trackFileCount == trackCount`; ntfy on mismatch. Plumbing: the metadata container has no Lidarr credentials today (the hook goes the other way, Lidarr → Dagu). Step 1/5 need the Lidarr API key as a secret in `media/metadata/secrets/`, `_FILE` convention, read inside the container only.
Sign in to join this conversation.
No milestone
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
mysticalsoap/docker#241
No description provided.