Skip to content

fix(ddl): backfill album saves and reposts to playlist - #1008

Closed
rickyrombo wants to merge 1 commit into
mainfrom
fix/saves-reposts-album-to-playlist
Closed

fix(ddl): backfill album saves and reposts to playlist#1008
rickyrombo wants to merge 1 commit into
mainfrom
fix/saves-reposts-album-to-playlist

Conversation

@rickyrombo

Copy link
Copy Markdown
Contributor

What

Collapses save_type / repost_type 'album' back into 'playlist'670 saves and 528 reposts.

Pairs with OpenAudio/go-openaudio#428, which stops the indexer writing 'album'. These ship together: this backfill alone would be re-broken by the current indexer.

Why

An album is a playlist with is_album = true. The indexer briefly derived a separate 'album' type by reading playlists.is_album at index time — a side effect of an entity-id collision fix, not a deliberate convention. It first appears in prod saves and reposts on the same day, 2026-05-28, after seven years with zero.

is_album is mutable, but save_type is written once and is part of the primary key (user_id, save_item_id, save_type, txhash) — so the same chain history indexed at different times produced different rows. Nothing reads the distinction: every consumer is track / != track, or ORs the two together (get_account_playlists, reconcile_aggregates). The notification triggers already derive album from is_album at read time.

Two hazards, and how they're handled

Duplicate notifications. on_save / on_repost fire on AFTER INSERT OR UPDATE, and the notification group_id embeds the type ('save:<id>:type:<save_type>'). A plain UPDATE would mint a second favourite notification per row under a new group_id. Those two triggers are disabled for the backfill; trg_saves / trg_reposts (the pg_notify → search indexer) stay enabled so ES still sees the change.

Primary key. The type is part of the PK. Verified against prod data that no (user_id, item_id, txhash) has both a 'playlist' and an 'album' row, so this is a straight UPDATE with no conflict handling.

Aggregate counts are unaffected either way — handle_save's delta is transition-aware and evaluates to 0 when is_delete does not change.

Verification

Applied against a fixture mirroring the real wiring (enums, 4-column PK, all four triggers):

  • all albumplaylist; track untouched
  • 0 on_save / on_repost firings
  • pg_notify fired exactly once per updated row
  • a row that had both a playlist and an album version kept both (different txhash)
  • all triggers re-enabled (tgenabled = 'O') afterwards

⚠️ Not run against a real database. The repo's migration set can't be built from scratch locally (it assumes base tables created outside this repo), so the fixture covers the mechanism rather than the full schema.

Notes

  • Each table gets its own transaction with lock_timeout to keep the ACCESS EXCLUSIVE lock from ALTER TABLE ... DISABLE TRIGGER as short as possible. ~1.2k rows, so it should be milliseconds.
  • Re-running is a no-op once no 'album' rows remain.
  • 'album' is deliberately left in the savetype / reposttype enums — Postgres can't drop an enum value without rebuilding the type, and the indexer no longer writes it.

🤖 Generated with Claude Code

An album is a playlist with is_album = true. The indexer briefly derived a
separate 'album' save_type/repost_type by reading playlists.is_album at
index time; it no longer does (OpenAudio/go-openaudio#428). This backfills
the rows written while that was live: 670 saves and 528 reposts, all
first appearing on 2026-05-28.

Two hazards this has to avoid:

  * on_save/on_repost fire on AFTER INSERT OR UPDATE, and the notification
    group_id embeds the type ('save:<id>:type:<save_type>'). A plain UPDATE
    would mint a second favourite/repost notification per row under a new
    group_id, so those two triggers are disabled for the backfill.
    trg_saves/trg_reposts stay enabled so the search indexer still sees the
    rows change.
  * save_type/repost_type are part of the primary key. Verified against
    prod data that no (user_id, item_id, txhash) has both a 'playlist' and
    an 'album' row, so this is a straight UPDATE with no conflict handling.

Aggregate counts are unaffected either way: handle_save's delta is
transition-aware and evaluates to 0 when is_delete does not change.

Each table gets its own transaction to keep the ACCESS EXCLUSIVE lock taken
by ALTER TABLE ... DISABLE TRIGGER as short as possible. Re-running is a
no-op once no 'album' rows remain. The 'album' label is left in the
savetype/reposttype enums since Postgres cannot drop an enum value without
rebuilding the type.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@rickyrombo

Copy link
Copy Markdown
Contributor Author

Superseded by #1011, which combines this with the other two api-side changes. They're only correct together in one deploy — merging the module bump without the users backfill would fail ETL migration 0035's index creation and stop the indexer starting.

@rickyrombo rickyrombo closed this Aug 4, 2026
rickyrombo added a commit that referenced this pull request Aug 5, 2026
Supersedes #1008, #1009 and #1010, which were the same work split three
ways.

## Why one PR

These three changes are only correct **together, in one deploy**. Split,
each one alone breaks something:

| merged alone | result |
|---|---|
| the bump | ETL `0035` creates `users_current_uniq_idx`, **fails on the
existing duplicates**, `RunMigrations` errors and the indexer won't
start |
| `0236` (album) | the old indexer keeps deriving `album` from
`is_album` and undoes the backfill |
| `0237` (users) | harmless, but pointless without the index that stops
it recurring |

As one PR the deploy is atomic, and the ordering inside it is guaranteed
by existing machinery: `bridge migrate` runs as a pre-roll Job that
every serving Deployment `DependsOn` (serving pods get
`runMigrations=false`), so both ddl migrations complete before the
indexer starts and runs the ETL's.

## Contents

**`0237_users_one_current_row_backfill`** — deletes 5 duplicate
`is_current` rows from `users`. Small count, large blast radius: joins
from an entity to its owner's wallet fan out, measured at **+18 tracks
and +787 follows** on a production clone. Must precede the ETL index.

**`0236_saves_reposts_album_to_playlist`** — 670 saves and 528 reposts
written as `album` collapse to `playlist`. `on_save`/`on_repost` are
disabled for the update (their notification `group_id` embeds the type,
so a plain UPDATE would mint duplicate favourite notifications);
`trg_saves`/`trg_reposts` stay enabled so the search indexer sees the
change.

**`deps: pin pkg/etl v1.6.4`** — brings OpenAudio/go-openaudio#428
(album type), #425 (the `users` invariant + genesis-writer join
simplification) and #433 (`0035` no longer deletes anything).

## The delete moved out of the ETL migration

`v1.6.3`'s `0035` deleted the duplicates itself. Since ETL migrations
run automatically at indexer start, that made a `go get` able to remove
rows from this database. #433 split it: the index stays in the ETL
migration, the repair moved to `0237` here. **`v1.6.4` ships zero
`DELETE` statements** — verified against the resolved module, not just
the tag.

The comment above the ETL config now records that line, and its
corollary: an ETL migration can depend on a ddl one having run, and
`0035` fails loudly if `0237` hasn't.

## Verified

- Resolved module `pkg/etl@v1.6.4` contains `0035` with `CREATE UNIQUE
INDEX` and **0** `DELETE` statements.
- Both migration orders against fixtures: backfill→index applies cleanly
(`violations 0`, `indisvalid = t`); index→backfill fails with `could not
create unique index … Key (user_id)=(98311147) is duplicated`, which is
the intended signal that `0237` hasn't run.
- Both migrations idempotent; re-running is a no-op.
- `0236` fires **zero** `on_save`/`on_repost` triggers against a fixture
with the real wiring, and exactly one `pg_notify` per updated row.
- `go build ./...` and `go vet ./indexer/` clean.
- No FK references `users`; its triggers are INSERT / INSERT OR UPDATE,
so the delete fires neither.
- Cutting `pkg/etl/v1.6.4` did not move `openaudio/go-openaudio:stable`
— still the 2026-07-30 `v1.8.2` digest, so no node-operator rollout.

## Not established

The cause of the duplicate `users` rows. Both indexer create paths
reject an existing user, so a single writer can't produce them; a second
writer can, since check-then-act isn't atomic across transactions. Three
of five pairs put a bare-hex `txhash` next to a `0x`-prefixed one, which
fits but doesn't prove it. The index will surface it if it recurs.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant