Skip to content

Index coordinator wedges permanently after a worker is killed mid-run: orphaned generation markers in /tmp are never reaped on daemon restart #2271

Description

@sblattj

Summary

When an index worker is terminated mid-run (client/session exit, SIGKILL, crash), it can leave orphaned per-generation marker files in the daemon's /tmp coordination directory. From then on, every new index worker, including the daemon's own file-watcher auto-refresh, refuses to start with:

CBM index worker could not start: a pre-coordination or unverified CBM generation is active

A daemon restart does not clear these markers, so the wedge survives every restart. Only a full machine reboot (which wipes /tmp) clears it. In practice this recurs repeatedly on a busy machine that runs many concurrent MCP clients.

Environment

  • codebase-memory-mcp 0.10.8
  • macOS (per-user coordination dir at /tmp/cbm-daemon-<uid>/)

Root cause (as observed)

  • The coordination directory (/tmp/cbm-daemon-<uid>/) holds advisory locks plus per-generation marker files: cbm-<generation-id>.rw and cbm-<generation-id>.turn.
  • A cleanly completed generation is fine, but a generation whose worker is killed after it enters pre-coordination/admission but before it verifies leaves its .rw/.turn markers behind.
  • New workers treat that orphaned, unverified generation as "active" and refuse to proceed (correct as a safety check against clobbering an in-progress index), but there is no path that reaps it.
  • Because the markers live in /tmp and are not reaped on daemon startup, restarting the daemon re-reads the same orphaned state and stays wedged. Reboot is the only self-heal.

The index databases themselves stay healthy and queryable throughout; only new indexing / auto-refresh is blocked.

Repro (sketch)

  1. Start indexing a repo.
  2. Kill the index worker (or the client that spawned it) mid-run, before the generation verifies.
  3. Attempt any subsequent index (or let the file-watcher auto-refresh fire on a git change).
  4. It fails with the "pre-coordination or unverified CBM generation is active" error.
  5. Restart the daemon: still fails. Only a reboot clears it.

Impact

  • Silent, permanent loss of auto-refresh and manual re-indexing until reboot.
  • High-concurrency setups (many MCP clients / short-lived sessions) hit it often, since mid-run kills are common there.
  • The failure is easy to misread as a version/build conflict; the actual cause is an unreaped orphaned generation.

Requested fix

On daemon startup (and/or when no live worker holds the corresponding lock), reap orphaned/unverified generation markers so the coordinator self-heals without a reboot. Options:

  • On startup, remove .rw/.turn markers for generations whose owning process is no longer alive (advisory-lock probe / PID liveness).
  • Age out unverified generations after a bounded TTL (a real generation verifies in seconds).
  • Emit a clear diagnostic when a wedge is detected and reaped, distinct from a genuine version/build conflict.

Current workaround

Stop the daemon, clear the /tmp coordination directory (equivalent to what a reboot does), and let the daemon respawn. The index DBs are untouched by this.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    editor/integrationEditor compatibility and CLI integrationstability/performanceServer crashes, OOM, hangs, high CPU/memoryux/behaviorDisplay bugs, docs, adoption UX

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions