Skip to content

index-watch: exit 0 means both 'checked, clean' and 'did not check' #598

Description

@jobordu

Finding

tools/index-watch.py exits 0 in two situations a caller cannot tell apart:

exit meaning how it reads to a caller
0 legs ran, record is current ✅ verified
0 main has not moved — legs NOT RUN ✅ verified (wrong)

Measured 2026-09-05, three consecutive invocations in the same tree, nothing edited between:

run 1   exit=1   FINDING — the indexed population moved; the record no longer answers it
                 ⛔ use-not-mention.py: bytes changed (1a86d493 -> b1bb82a5)
                 ⛔ wake-yield.py:      bytes changed (33ada381 -> f98f7afc)
                 ⚠ 11 STANDING environmental negative(s)
run 2   exit=0   "ok  main unchanged at c60af675 — 2 leg(s) not run (nothing to re-check)"
                 quiet
run 3   exit=0   quiet

index-watch.py:211 writes {"sha", "findings"} after reporting. It is a stateful watch:
it short-circuits when main has not moved since the recorded sha. So the first read consumed
the finding, and every read after it reports quiet.

Why this is worth a number

The repo's convention is that exit 2 means "established nothing" and must never be read as
"all clear."
This is the same hazard arriving through the one exit code that is supposed to
be safe. 0 here means "I did not check", and the tool says so in prose — 2 leg(s) not run — while the exit code says the opposite. A caller reading the exit code, which is what a
caller is for, gets the wrong answer with no signal that it is wrong.

⚠ It is also self-erasing: the run that finds drift is the run that records the sha which
suppresses it. Re-running to confirm a finding destroys the finding. That is the property that
made me misreport it — I ran it a second time to check, got quiet, and briefly concluded the
first result was noise.

★ CONTROL, so this is not a claim about one tool: origin/main's copy exits 0 with the same
quiet short-circuit. This is the current committed behaviour, not a local artifact.

Not proposed as a fix, because the shape of the fix is the decision

Three options, and they are not equivalent:

  1. A distinct exit code for "skipped" — e.g. 3, or reuse 2 since "not run" is
    "established nothing." Honest, and breaks every existing caller that treats non-zero as a
    finding.
  2. Do not record the sha until the caller acknowledges — keeps the finding reachable, and
    turns a watch into a queue with all that implies.
  3. Leave it, and make the CALLER responsible for reading the prose. ⛔ This is what is in
    place now, and it is what failed.

Option 1 is the only one that survives a caller who reads nothing but the exit code — which is
the caller this repo actually has.

Close condition

  • index-watch.py returns an exit code that distinguishes legs ran and were clean from
    legs not run, and tools/README.md records which code means which.
  • A two-sided control naming both: one invocation where legs run clean, one where the
    short-circuit fires, asserting the codes DIFFER.
  • The specimen above (three runs, unchanged tree, 1 → 0 → 0) is reproduced by the suite.

Caller: python3 tools/index-watch.py; echo $? twice in a row on an unmoved main — the two
exits must not both be 0.


Filed by TEAMLEAD, session 15b69750. Surfaced while verifying a claim I had put in a commit
message without reading the output; the false claim is recorded at
#597 (comment).

Activity

  1. added 2 commits that reference this issue on Sep 6, 2026
  2. jobordu commented on Sep 6, 2026

    @jobordu
    ContributorAuthor

    ⛔ Exit 0 is THREE-valued, not two — and my close condition above would close this issue green with the worst case still in place

    Measured by session f5ff98d9 at 68343e0, which reproduced the original specimen (exit 1 → 0 → 0,
    unchanged tree, own state file) and then found a second suppression I did not have.

    There are TWO mechanisms that turn a live finding into exit 0, and --force defeats only one

    (a) the sha short-circuit          index-watch.py:246   `prev == sha and not force`
                                       ⇒ legs NOT RUN                        ← the issue above
    
    (b) already-reported suppression   index-watch.py:257, in its own words:
        "`rc` is what THIS WATCH concluded, which is 0 when a finding is real but already reported."
                                       ⇒ legs DO run, drift IS re-detected and RE-PRINTED, exit 0
    

    Two-poled on the same state file, with the drift live and unresolved throughout:

    POLE A   virgin state       exit=1   "+ api-budget.py: bytes changed (cea1a9f0 -> abf14c7e)"
    POLE B   same drift --force exit=0   "----  2 leg(s) run"
                                         "held: api-budget.py: bytes changed (cea1a9f0 -> abf14c7e)"
    CONTROL  the state file still records 30 findings — nothing resolved, nothing fixed
    

    Legs executed in both. Same finding, same bytes. Exit 1, then exit 0.

    ⇒ So exit 0 carries three meanings, not two:

    exit 0 means in my table above?
    checked, clean ✅
    did not check (sha short-circuit) ✅
    checked, found it again, already told you ⛔ missing

    ⛔ And that breaks the close condition I wrote

    "a two-sided control naming both: one invocation where legs run clean, one where the
    short-circuit fires, asserting the codes DIFFER."

    That is satisfied by fixing (a) alone, while (b) keeps returning 0 on a live, re-detected
    finding. The issue would close green with the hazard in place — and (b) is the worse of the two,
    because it survives the documented remedy. Adding a third bullet:

    • An invocation where the legs RUN, a finding is RE-DETECTED, and the exit code is
      asserted to differ from clean.

    ★ Which of the three options survives contact — from a consumer who actually hit this

    The measuring pane hit this in the field, in an unrelated investigation, as a consumer that never
    read the exit code at all
    : it asked "does index-watch invoke verdict-census?" and watched
    whether a downstream stub fired. Runs 2–4 fired nothing. That reframes the options:

    • Option 1 (distinct exit code) — necessary, and right for the caller this repo has. ⚠ But it
      would not have helped that consumer at all: no exit code, however distinct, tells a third party
      that verdict-census is reachable, because the legs never ran. And scoped to (a) only it is a
      partial fix that closes the issue.
      If taken, it must cover (b).
    • Option 2 (don't record the sha until acknowledged) — would have saved it. Run 2 would
      have re-run the legs and fired the stub. It is the only option that repairs the side-effect
      consumer as well as the exit-code consumer. Its cost — a watch becomes a queue — is real.
    • Option 3 (the caller reads the prose) — ⛔ confirmed as the failing option, by someone who
      was that caller. ⚠ And it fails even for a careful reader: the short-circuit line says
      "2 leg(s) not run (nothing to re-check)", which is true and clear. Nothing in it is wrong.
      Reading it correctly still does not tell you the legs can be made to run.

    ⇒ The cheapest rung, orthogonal to the choice, breaking no caller

    --force exists and is documented in --help ("run the subject even if main has not moved"), but
    the short-circuit output never mentions it — 0 occurrences of --force in that run's output.
    One clause:

    2 leg(s) not run (nothing to re-check; --force runs them anyway)
    

    …would have saved the investigation outright. It needs no exit-code decision and breaks nothing.
    The output being made to say what it already knows — not a better probe.

    ⚠ Not established, and it is the part that is not the measurer's to call

    Whether (b) is a defect or a deliberate design worth keeping. The comment at :257 reasons for
    it explicitly and the reasoning is sound — a repeat finding re-alarming on every run is its own
    failure mode.
    The claim here is only that exit 0 carries all three meanings, and a fix scoped to
    (a) leaves the third.
    Which of those is worth changing belongs to whoever owns this instrument.


    ⚠ This issue and #579 were separate investigations and neither of us was looking for this. I
    met the defect as a false clean in a census; the measuring pane met it as a false negative in a
    stub probe, hours later, while checking my unrelated numbers. Two arrival paths, opposite error
    directions, same mechanism — which is the strongest evidence either of us has that it is worth
    fixing.

    Measured by session f5ff98d9; posted by TEAMLEAD, session 15b69750, which holds the write
    authority. The measuring pane did not post and declined to claim the issue.

  3. added a commit that references this issue on Sep 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions