diff --git a/README.md b/README.md index fa77dc16d..fc7274d1f 100644 --- a/README.md +++ b/README.md @@ -1160,14 +1160,17 @@ rather than blurring it: among the metrics shown to track measured cognitive load directly (Peitek, Apel, Parnin, Brechmann & Siegmund, ICSE 2021, [doi:10.1109/ICSE43902.2021.00056](https://doi.org/10.1109/ICSE43902.2021.00056)) — `--readability` emits volume and stops there; difficulty and effort are computed nowhere in this - tree. **Unvalidated against human judgement, stated plainly rather than assumed:** this is a - deterministic ordering signal, not a checked one. Our own proxy measurement — 484 matched - before/after function pairs mined from 80 refactor/simplify/cleanup commits in this repository's own - history — found the lens agrees with the commit's implied readability direction on only 30.2% of - pairs, worse than chance. That is a construct-validity finding about the ordering claim, not a bug in - the arithmetic (a separate self-consistency check confirms the formula computes exactly what it says - it computes); the lens itself is unchanged pending a proper human study, and `--help=--readability` - carries the same caveat where a CLI reader meets it. + tree. **A ranking lens, never a grade, and here is what it actually orders:** on ripwire's own + history at the pinned `v0.6.2` tag (412 function pairs mined from 80 refactor/simplify/cleanup + commits), the order between two versions of a function followed the sign of its token-count change + in 96.0% of pairs — read a move as *more or fewer tokens*, not as *more or less readable*. Of the 412, + 154 (37.4%) ran the commit's implied direction and 258 (62.6%) ran opposite it; Halstead volume drove + 91.1% of those 258; a separate self-consistency check confirms the formula + computes exactly what it says it computes, so this is a construct-validity finding about the + ordering claim, not an arithmetic bug. (First recorded as 484 pairs / 30.2% from an unpinned `git + log --all` walk that does not reproduce — full derivation, the instrument fix and the retracted + figure: [`docs/EVALS.md` §8](docs/EVALS.md).) The lens itself is unchanged pending a proper human + study, and `--help=--readability` carries the same figures where a CLI reader meets them. - **`--naming-consistency`** is the *lexical* family's one exception to "evidence, never advice": every other lens in this panel tells you WHAT is wrong, never a computed fix. Case-style consistency is the one property with a corpus-derivable answer — on this repository's `src/`, camelCase is the diff --git a/docs/COMMANDS.md b/docs/COMMANDS.md index 8e491b610..ffdf44cac 100644 --- a/docs/COMMANDS.md +++ b/docs/COMMANDS.md @@ -1839,7 +1839,7 @@ $ ./build/ripwire . --clones **Answers:** rank functions least-readable first, by volume, token entropy and length per-function readability lens, LEAST readable first: vol= Halstead volume V (N*log2(eta)), ent= Shannon token entropy E, lines= L, posnett= sigmoid(8.87 - 0.033V + 0.40L - 1.5E) (Posnett/Hindle/Devanbu, MSR 2011). -APPROXIMATION, disclosed: ONE token-class table serves every language (keywords + punctuation = operators, identifiers + literals = operands), with no per-grammar refinement, so V is cross-language and not a per-grammar Halstead count. The formula was fitted on snippets of 20 lines or fewer, so it is a RANKING lens, not a grade: read the ORDER of the rows, not the number on any one of them. Pages with limit=N (offset=M); default 40 rows. Declarations with no body are not measured. UNVALIDATED (t14-cleanup #8): this is a deterministic ORDERING signal that has not been checked against human judgement of readability. Our own proxy measurement — 484 matched before/after function pairs from 80 refactor/simplify/cleanup commits in this repo's own history — found the lens agrees with the commit's implied readability direction on only 30.2% of pairs, which is worse than chance and suggests the ranking may run backwards more often than not. Treated here as a signal to weigh, never a verdict; do not read a low posnett= as proof a function needs work. +APPROXIMATION, disclosed: ONE token-class table serves every language (keywords + punctuation = operators, identifiers + literals = operands), with no per-grammar refinement, so V is cross-language and not a per-grammar Halstead count. The formula was fitted on snippets of 20 lines or fewer, so it is a RANKING lens, not a grade: read the ORDER of the rows, not the number on any one of them. Pages with limit=N (offset=M); default 40 rows. Declarations with no body are not measured. MEASURED: on ripwire's own history at the pinned v0.6.2 tag (412 function pairs mined from 80 refactor/simplify/cleanup commits), the order between two versions of a function tracked the sign of its token-count change in 96.0% of pairs (388/404 whose count changed). Read a move as more or fewer tokens, not as more or less readable; full derivation and history in docs/EVALS.md §8. Do not read a low posnett= as proof a function needs work. **Try it** @@ -2831,7 +2831,7 @@ $ ./build/ripwire . --quality-baseline --allow-dirty ### `--quality-delta` -**Answers:** before a PR: report ONLY what your change made worse, across 10 kinds agent self-check before a PR (pair with --test-gate): report ONLY what a change made worse vs the baseline (10 kinds: complexity/verbosity/nesting/params/dup/dead/api-surface + error-masking/short-horizon-churn/reuse-decline); +**Answers:** before a PR: report ONLY what your change made worse, across 10 kinds agent self-check before a PR (pair with --test-gate): report ONLY what a change made worse vs the baseline (10 kinds: complexity/verbosity/nesting/params/dup/dead/api-surface + error-masking/short-horizon-churn/new-clone-of-reused-helper); every finding is classified by ORIGIN: a symbol that EXISTED at the baseline and got worse (preexisting-worse="N", no attribute on the row) vs one that exists only because the code is NEW (new-symbol="N", origin="new-symbol" on the row). A small numeric delta is additionally sev="minor". EXIT 2 ONLY on preexisting-worse AND major AND unacked — the gating="N" header count. New-symbol rows are still PRINTED (they are the debt you are adding — read them), they just never gate; exit 0 means "nothing that already existed got worse", not "clean". Clone kinds classify by member set (new-symbol only if EVERY member is new); short-horizon-churn is preexisting by construction. LIMIT: origin is canonId (path::scope::name) identity, so a RENAMED/MOVED symbol reads as new and a regression carried in with the move will not gate. Test-fixture dirs + doc sections are exempt from dead-code/churn; churn needs COMMITTED thrash evidence (rewritten across recent commits AND again by this diff), never the current edit alone WHICH FLOOR IT COMPARES AGAINST, and a side effect: the sidecar is honored only when the sha it was pinned at EQUALS the current git HEAD (strict equality — an ancestor commit describes a DIFFERENT tree, so everything committed since would read as your regression). A sidecar pinned anywhere else is STALE: this verb then DELETES it from your working tree (self-heal, so the next run does not rediscover the dead pin) and auto-compares the working tree vs git HEAD instead. Re-pin with --quality-baseline. The read-only MCP quality_delta verb applies the SAME staleness test but never deletes. A sidecar at the current HEAD that ANOTHER ripwire build pinned (its producer stamp names other sources — a dead set depends on how calls were resolved) is FOREIGN: both arms ignore it, never delete it, and auto-compare vs git HEAD. Which floor was actually used is on every report as baseline=: sidecar | git-HEAD | git-HEAD (stale sidecar removed) | git-HEAD (stale sidecar ignored) | git-HEAD (foreign sidecar ignored) — the stale two say a stale sidecar existed, and 'removed' means the file is gone. A non-git root has no HEAD to fall back to, so its sidecar is honored whenever this build pinned it; without one there, or with another build's, the verb exits 1. @@ -3243,7 +3243,7 @@ $ ./build/ripwire . --safe-delete=DoesNotExist **Answers:** trace one variable's definitions and uses inside one function NAME-BASED intra-procedural def-use slice of variable VAR inside the ONE uniquely-resolved definition SYM (statement-level def-use edges as a queryable primitive — the ARISE result, arXiv:2605.03117). -One row per line touching VAR, source order: k=def|use|both| scope = a Python global/nonlocal statement, neither read nor write; t=param|decl| assign|call-arg|read|global|nonlocal = the strongest role on the line; CDATA = the trimmed source line; defs=/uses= count occurrences. JS/TS destructuring binders (`const {a, b} = o`, `[x] = arr`, destructured parameters) are locals whose def is the pattern line. A write hidden behind a call — receiver mutation, a by-reference/out-parameter, a function-like macro — is a use, never a def (stated in the legend). Bare --slice=SYM lists the sliceable locals ( rows) so a caller can pick VAR. LIMITS in the legend, not implied: no alias analysis. REACHING DEFINITIONS are FLOW-SENSITIVE inside the definition for C-family and Python (root reach="cfg": a def is killed by the next unconditional def on every path, defs join at if/elif/else, switch, loop back-edge, try/finally, for/while-else, match, #ifdef merges) and source-order for JS/TS/Go/Java/Rust (reach="linear", nothing joins); every use row carries rd= (the lines of the defs that reach it, "-" = none). The unit is the STATEMENT (uses read the entering state, defs apply after); a nested lambda/def body, ?:, short-circuit fold into their statement, goto is untracked, global/nonlocal is tracked like a local — each disclosed in the legend. Block scopes ARE separated: a name declared twice in the definition is two variables, each row of a shadowed name carries b= (the declaration line it binds to), the root bindings=, the inventory one per binding. SYM matching several definition sites REFUSES (exit 1) listing the file:name spellings that pick one, like --edit-check. Served: C/C++/ObjC (+CUDA/Metal), Python, JS/TS, Go, Java, Rust — other indexed languages refuse loudly (never an empty success). Single-root only. PREPROCESSOR (C-family): a `#if 0` body and the `#else` of `#if 1` are DEAD — dropped, the line count disclosed as preproc_rows=; every other conditional region (`#ifdef X`, `#ifndef X`, `#if defined(X)`, `#if EXPR`) is build-dependent and cannot be decided without the build's macro set, so its rows are KEPT and flagged pp="1", and a pp def never hides the unconditional def before it in a flow (both are reaching). LINE-SEEDED: --at=FILE:LINE beside --slice (or --slice=@FILE:LINE) is the ARISE (file, line[, variable]) seed — the definition sliced is the innermost one enclosing the line (a seed narrows an otherwise-ambiguous SYM; a seed enclosed by none of SYM's definitions refuses naming both). A seed line naming exactly ONE sliceable local pre-picks it (disclosed: seed= var_from="seed"); zero or several serve the inventory with seed_vars= and the candidate rows marked seed="1", never a guess. A plain identifier spec beside --at reads as the seed's VARIABLE (--slice=VAR --at=src/f.cpp:12). SINCE: --since=REV|DATE beside --slice=SYM:VAR adds a child carrying the DEPENDENCE diff of that variable against the committed tree at REV — one row per added or removed STATEMENT of the variable, one row per added or removed def-use edge. The unit is the STATEMENT and the key is the ROLE, never the line and never the text, so a re-wrap, a comment edit, an insertion above the definition, and a rename of an unrelated local all come back EMPTY. Empty means no def-use edge of that variable moved, never that the commit changed nothing — git diff answers the second question. status= names each way the symbol can be absent at REV, and comparable="0" says outright that no comparison was made and the emptiness is not evidence. Refused on the bare inventory: a dependence diff needs a seed variable. +One row per line touching VAR, ranked as the root's order="defuse" states: def-use coverage (distinct local names on the line) descending, then line — measured (docs/EVALS.md): puts a gold line first more often than a random shuffle. k=def|use|both| scope = a Python global/nonlocal statement, neither read nor write; t=param|decl| assign|call-arg|read|global|nonlocal = the strongest role on the line; CDATA = the trimmed source line; defs=/uses= count occurrences. JS/TS destructuring binders (`const {a, b} = o`, `[x] = arr`, destructured parameters) are locals whose def is the pattern line. A write hidden behind a call — receiver mutation, a by-reference/out-parameter, a function-like macro — is a use, never a def (stated in the legend). Bare --slice=SYM lists the sliceable locals ( rows) so a caller can pick VAR. LIMITS in the legend, not implied: no alias analysis. REACHING DEFINITIONS are FLOW-SENSITIVE inside the definition for C-family and Python (root reach="cfg": a def is killed by the next unconditional def on every path, defs join at if/elif/else, switch, loop back-edge, try/finally, for/while-else, match, #ifdef merges) and source-order for JS/TS/Go/Java/Rust (reach="linear", nothing joins); every use row carries rd= (the lines of the defs that reach it, "-" = none). The unit is the STATEMENT (uses read the entering state, defs apply after); a nested lambda/def body, ?:, short-circuit fold into their statement, goto is untracked, global/nonlocal is tracked like a local — each disclosed in the legend. Block scopes ARE separated: a name declared twice in the definition is two variables, each row of a shadowed name carries b= (the declaration line it binds to), the root bindings=, the inventory one per binding. SYM matching several definition sites REFUSES (exit 1) listing the file:name spellings that pick one, like --edit-check. Served: C/C++/ObjC (+CUDA/Metal), Python, JS/TS, Go, Java, Rust — other indexed languages refuse loudly (never an empty success). Single-root only. PREPROCESSOR (C-family): a `#if 0` body and the `#else` of `#if 1` are DEAD — dropped, the line count disclosed as preproc_rows=; every other conditional region (`#ifdef X`, `#ifndef X`, `#if defined(X)`, `#if EXPR`) is build-dependent and cannot be decided without the build's macro set, so its rows are KEPT and flagged pp="1", and a pp def never hides the unconditional def before it in a flow (both are reaching). LINE-SEEDED: --at=FILE:LINE beside --slice (or --slice=@FILE:LINE) is the ARISE (file, line[, variable]) seed — the definition sliced is the innermost one enclosing the line (a seed narrows an otherwise-ambiguous SYM; a seed enclosed by none of SYM's definitions refuses naming both). A seed line naming exactly ONE sliceable local pre-picks it (disclosed: seed= var_from="seed"); zero or several serve the inventory with seed_vars= and the candidate rows marked seed="1", never a guess. A plain identifier spec beside --at reads as the seed's VARIABLE (--slice=VAR --at=src/f.cpp:12). SINCE: --since=REV|DATE beside --slice=SYM:VAR adds a child carrying the DEPENDENCE diff of that variable against the committed tree at REV — one row per added or removed STATEMENT of the variable, one row per added or removed def-use edge. The unit is the STATEMENT and the key is the ROLE, never the line and never the text, so a re-wrap, a comment edit, an insertion above the definition, and a rename of an unrelated local all come back EMPTY. Empty means no def-use edge of that variable moved, never that the commit changed nothing — git diff answers the second question. status= names each way the symbol can be absent at REV, and comparable="0" says outright that no comparison was made and the emptiness is not evidence. Refused on the bare inventory: a dependence diff needs a seed variable. **Try it** @@ -4386,7 +4386,7 @@ _The session legend dictionary the MCP server serves as ripwire://legend-dict/fu ``` $ ./build/ripwire . --legend-dict -ripwire legend dictionary ripwire.dict/v1 dictv=de80b7c7b3ea84b5 entries=703 +ripwire legend dictionary ripwire.dict/v1 dictv=2d9cec7852fd809b entries=704 : the answer's rows come first; its root keeps only task= changed= from= to=, and this LAST child carries every other root attribute unchanged (schema= included); legend="ref": a definition is sent once per session (this dictionary's core, or the first answer that ne … [line truncated: 83 more bytes on this line] schema=ripwire.KEY/v1: the line ripwire.KEY/v1 below reads the answer's rows window: shown= total= capped= has_more= next_offset= offset= limit= page a list (capped=1 cut; next_offset= pastes as offset=) diff --git a/docs/EVALS.md b/docs/EVALS.md index 97eff9a02..b16761d7f 100644 --- a/docs/EVALS.md +++ b/docs/EVALS.md @@ -6867,6 +6867,22 @@ Listed because the reason is more useful than the silence. prototyped on two corpora and rejected — see "Shotgun Surgery — two formulations measured" at the end of this document. What ships for the smell is the co-change check `--situ` / `--pr-context` already carried; its backtest numbers there are the only ones this project publishes about it. +- **"484 matched pairs, 146 right-direction, 30.2%"** for `--readability`'s construct-validity proxy + (`--help=--readability`, this document, and README.md all carried it). It was first computed by + walking `git log --all` over a clone whose branch set changes every time a lane is pushed — a rerun + of the identical, unmodified script gave 413 pairs (38.3%) and, separately, 409 pairs (38.4%), + neither matching the published number and neither touching `src/readability.h`. A number that moves + when an unrelated lane is pushed is not measuring the lens; it is measuring which branches exist in + the shared `.git` right now. Pinned to the immutable `v0.6.2` tag, the same instrument reproduces + deterministically: **412 function pairs, 154 right-direction (37.4%)**. The mechanism behind the + inversion is not "the lens is wrong 63% of the time" — it is that the sign of a function's + token-count change predicts the lens's direction in **96.0%** of pairs (388/404 whose token count + changed), with the Halstead-volume term driving 91.1% of the 258 wrong-direction pairs (62.6%). Full protocol, + the instrument-fix note, and the decomposition: `docs/research/readability-construct-validity.md` + §3a/§3c — a draft investigation, PR [#313](https://github.com/redhat-et/ripwire/pull/313) (open; + cited here for the derivation only, nothing shipped depends on it merging). The shipped `--help` + text now states the pinned figures; this entry keeps the retracted number visible rather than + silently replacing it. --- @@ -13868,3 +13884,81 @@ The detected root comes from the first transcript whose first line names it, whi scans until one matches instead of reading only the first file: 113 of the 628 do not name the absolute root on their first line — that is the complement of the 515 above, and it is not a claim that those 113 print no banner, only that the root is not in it. + +## `--slice=SYM:VAR` def-use row order — PRE-REGISTERED 2026-09-22 (before any src/ change and before any number below was computed) + +`--slice=SYM:VAR` seed rows emitted in SOURCE order (the file's own line order), unstated on the root. The +red-first gate arm (commit `48a4f071`) registered a rule and a decision procedure before any implementation +code or any number existed: + +> **Rule under test: R1 def-use coverage.** A `--slice=SYM:VAR` row's score is the number of distinct +> sliceable locals with an occurrence on that line; rows emit score-descending, then line ascending, then +> binding line ascending (file is constant: one definition). Zero fitted parameters. +> +> **Measurement:** `bench/slice/run_slice_linerecall.py` from `lane/research-arise-slice` (LocBench V1 test, +> Python single-function rows), plus one scratch arm that reads the EMITTED order: per (scored instance, +> inventory variable) whose v1 rows hold a gold line, Recall@{1,3,5,10,20} and MRR of the `l=` values in +> emission order against gold & rows, versus a uniform random permutation of the same rows (200 shuffles, +> seed 20260920, own RNG). Population: every inventory variable (unseeded); the gold-touched subset is +> reported beside it, never instead of it. +> +> **Decision:** ADOPT R1 ordering iff the new binary's emitted-order MRR beats the random control's MRR on +> the all-inventory population. Otherwise stop ranking: emit in source order, and state that order in the +> header as a presentation order, not a ranking. + +Neither the harness (`bench/slice/run_slice_linerecall.py`) nor the scratch arm that reads emission order is +committed to this tree — both live on the unmerged `lane/research-arise-slice`, and the scratch arm is a +43-line diff over that harness. The numbers below are reproduced from that harness's output and from an +independent re-run against both binaries (base `15a20855` and this lane), not from a script this tree ships; +`docs/research/slice-line-recall.md`, cited by an earlier draft of this feature, does not exist on `main` or +on this lane and is not the record of this claim — this section is. + +## `--slice=SYM:VAR` def-use row order — MEASURED 2026-09-22 against the band above: **ADOPTED — the emitted-order MRR beats the random control** + +**Population, stated precisely (correcting the scratch arm's own label):** 478 (scored instance, inventory +variable) pairs, from 173 LocBench V1 Python instances, **whose v1 rows hold a gold line** — not "every +inventory variable, unseeded" as the scratch arm's `all_inventory` label implied. A pair with no gold line +among its rows scores 0 under every candidate order, so the ADOPT/DON'T-ADOPT decision is unaffected either +way; only the population's name was overstated. + +| arm | @1 | @3 | @5 | @10 | @20 | MRR | +| --- | --- | --- | --- | --- | --- | --- | +| source order (base `15a20855`, the "before") | 0.198 | — | — | — | — | 0.525 | +| random control (200 shuffles, seed 20260920, sequential stream) | 0.268 | 0.676 | 0.827 | 0.939 | 0.983 | 0.602 | +| def-use coverage (this lane, emitted order, the "after") | 0.285 | 0.743 | 0.847 | 0.937 | 0.983 | 0.628 | + +@3/@5/@10/@20 were not separately reported for source order at registration time — only MRR and @1 were +measured for that arm; the table states that gap rather than filling it in. + +The registered decision is the MRR row: **0.628 > 0.602 > 0.525** — def-use coverage beats the random +control, which beats source order. ADOPTED per the pre-registered rule. + +**The control's spread, so the margin has a scale.** Across the control's own 200 shuffles, the mean MRR's +standard deviation is 0.0117 (the registered sequential RNG stream) / 0.0116 (an independent per-pair RNG, +cross-check only); the emitted-order MRR exceeds 199 of those 200 shuffle means (197/200 under the +independent RNG) — roughly +2.2σ. A paired bootstrap of (emitted MRR − that pair's own control mean) over +the 478 pairs gives Δ = +0.026, 95% CI [0.004, 0.049] — excludes zero, narrowly. + +**The gain is concentrated at @1/@3, honestly split per pair.** Per pair against its own control mean: +better on 182, WORSE on 227, tied on 69 — def-use coverage loses the per-pair comparison more often than it +wins. The population-level win lives in where gold lands when the rule is right: @1/@3 favor def-use +clearly, @10/@20 tie the shuffle (0.937 vs 0.939, 0.983 vs 0.983 — the shuffle is marginally ahead at @10). +On average 43% of a pair's rows share the top def-use-coverage score, so line order (the tiebreak) still +does real work inside that band. + +**In-sample, honestly.** Zero fitted parameters — the rule is a fixed count-and-sort, not tuned against this +population — but the population itself is the LocBench V1 Python set the rule was measured on; there is no +held-out split. **Python only.** The coverage count follows the `--slice` inventory exactly (as registered), +and that inventory is uneven across languages the tool serves — Python lambda/JS arrow params are not +inventory locals while Python nested-def/Rust closure params are, Python `self` counts as a param local +while Rust `self` does not, Java/Python attribute identifiers share a name with same-spelled locals while +C++/JS field/property identifiers do not. None of this was measured for C/C++, JS/TS, Go, Java, C#, Ruby, +Rust, or Swift; the `order="defuse"` ranking ships for every served language on the strength of the Python +measurement alone, stated here rather than left implicit. + +**Falsifiable claim, restated to match what the table shows:** *"Ranking `--slice=SYM:VAR` rows by def-use +coverage puts a gold line first (rank 1) more often than a random shuffle of the same rows, on LocBench V1 +Python."* That is an @1 claim (0.285 vs 0.268), not a claim that the rule outranks a shuffle on every pair +or at every depth — @10/@20 tie, and the per-pair split has more losses than wins. `--help` and +`src/slice.h`'s `sliceDefUseRowOrder` comment are worded to this claim, not to the broader "ranks above it" +a first draft of this feature shipped. diff --git a/docs/captures/COMMANDS_showcase_2026-09-14.md b/docs/captures/COMMANDS_showcase_2026-09-14.md index 599301eb4..b2cfbb7aa 100644 --- a/docs/captures/COMMANDS_showcase_2026-09-14.md +++ b/docs/captures/COMMANDS_showcase_2026-09-14.md @@ -5351,7 +5351,7 @@ ripwire: --run-timeout=SECONDS modifies --run-trace — pass it too (e.g. ripwir *The session legend dictionary the MCP server serves as ripwire://legend-dict/full — one definition per line, headed by its dictv= version; no corpus needed. =roster lists the completeness attributes it defines.* ````` -ripwire legend dictionary ripwire.dict/v1 dictv=de80b7c7b3ea84b5 entries=703 +ripwire legend dictionary ripwire.dict/v1 dictv=2d9cec7852fd809b entries=704 : the answer's rows come first; its root keeps only task= changed= from= to=, and this LAST child carries every other root attribute unchanged (schema= included); legend="ref": a definition is sent once per session (this dictionary's core, or the first answer that ne … [line truncated: 83 more bytes on this line] schema=ripwire.KEY/v1: the line ripwire.KEY/v1 below reads the answer's rows window: shown= total= capped= has_more= next_offset= offset= limit= page a list (capped=1 cut; next_offset= pastes as offset=) @@ -5381,7 +5381,7 @@ ripwire.impact/v1 : transitive blast radius of of=: reach s ripwire.path/v1 : one DIRECTED call path from= to to=, each a hop; reachable=0 hops=0 when none ripwire.connect/v1 : minimal joining subgraph: groups, terminals, joins, edges, ripwire.at/v1 : enclosing-definition chain at p=:l=: sym= innermost, chain= outermost-first, spans -… [674 more display lines; full output is 68453 bytes on 704 raw line(s)] +… [675 more display lines; full output is 68632 bytes on 705 raw line(s)] ````` ## `./build/ripwire . --lint --lint-select=cache-` diff --git a/src/cli.h b/src/cli.h index 327a61c55..000643347 100644 --- a/src/cli.h +++ b/src/cli.h @@ -1387,13 +1387,12 @@ inline constexpr char kHelpHead[] = " The formula was fitted on snippets of 20 lines or fewer, so it is a RANKING lens, not a\n" " grade: read the ORDER of the rows, not the number on any one of them. Pages with limit=N\n" " (offset=M); default 40 rows. Declarations with no body are not measured.\n" - " UNVALIDATED (t14-cleanup #8): this is a deterministic ORDERING signal that has not been\n" - " checked against human judgement of readability. Our own proxy measurement — 484 matched\n" - " before/after function pairs from 80 refactor/simplify/cleanup commits in this repo's own\n" - " history — found the lens agrees with the commit's implied readability direction on only\n" - " 30.2% of pairs, which is worse than chance and suggests the ranking may run backwards more\n" - " often than not. Treated here as a signal to weigh, never a verdict; do not read a low\n" - " posnett= as proof a function needs work.\n" + " MEASURED: on ripwire's own history at the pinned v0.6.2 tag (412 function pairs mined\n" + " from 80 refactor/simplify/cleanup commits), the order between two versions of a\n" + " function tracked the sign of its token-count change in 96.0% of pairs (388/404 whose\n" + " count changed). Read a move as more or fewer tokens, not as more or less readable;\n" + " full derivation and history in docs/EVALS.md §8. Do not read a low posnett= as proof\n" + " a function needs work.\n" " --nonlocal-state per function, the non-local mutable state it can reach, most writes first\n" " per function, the NON-LOCAL MUTABLE STATE it can reach, MOST WRITES FIRST: writes= reads= are the\n" " distinct cells this function OR its transitive callees write / read; direct_writes= direct_reads=\n" @@ -1636,7 +1635,7 @@ inline constexpr char kHelpHead[] = " later --quality-delta against it carries baseline_absorbed=\"N\" — so a green exit beside that attribute reads as\n" " \"clean SINCE THE PIN\", never \"clean\". Refused alone.\n" " --quality-delta before a PR: report ONLY what your change made worse, across 10 kinds\n" - " agent self-check before a PR (pair with --test-gate): report ONLY what a change made worse vs the baseline (10 kinds: complexity/verbosity/nesting/params/dup/dead/api-surface + error-masking/short-horizon-churn/reuse-decline);\n" + " agent self-check before a PR (pair with --test-gate): report ONLY what a change made worse vs the baseline (10 kinds: complexity/verbosity/nesting/params/dup/dead/api-surface + error-masking/short-horizon-churn/new-clone-of-reused-helper);\n" " every finding is classified by ORIGIN: a symbol that EXISTED at the baseline and got worse (preexisting-worse=\"N\", no attribute on the row) vs one that exists only\n" " because the code is NEW (new-symbol=\"N\", origin=\"new-symbol\" on the row). A small numeric delta is additionally sev=\"minor\". EXIT 2 ONLY on preexisting-worse AND\n" " major AND unacked — the gating=\"N\" header count. New-symbol rows are still PRINTED (they are the debt you are adding — read them), they just never gate; exit 0 means\n" @@ -1817,7 +1816,9 @@ inline constexpr char kHelpHead[] = " --slice=SYM[:VAR] trace one variable's definitions and uses inside one function\n" " NAME-BASED intra-procedural def-use slice of variable VAR inside the ONE uniquely-resolved\n" " definition SYM (statement-level def-use edges as a queryable primitive — the ARISE result,\n" - " arXiv:2605.03117). One row per line touching VAR, source order: k=def|use|both|\n" + " arXiv:2605.03117). One row per line touching VAR, ranked as the root's order=\"defuse\"\n" + " states: def-use coverage (distinct local names on the line) descending, then line — measured\n" + " (docs/EVALS.md): puts a gold line first more often than a random shuffle. k=def|use|both|\n" " scope = a Python global/nonlocal statement, neither read nor write; t=param|decl|\n" " assign|call-arg|read|global|nonlocal = the strongest role on the line; CDATA = the trimmed source\n" " line; defs=/uses= count occurrences. JS/TS destructuring binders (`const {a, b} = o`,\n" diff --git a/src/compactlegend.h b/src/compactlegend.h index d2576f0f7..235d0c3d3 100644 --- a/src/compactlegend.h +++ b/src/compactlegend.h @@ -1017,6 +1017,7 @@ inline constexpr CompactCompletenessTerm kCompactAttributeReadings[] = // slice: src/slice.h (the root emit, kSliceCountsAttrXml) { "sym", "sym=/lang=: the sliced definition's name and language; p= is its file:line", false, "slice", MapHeaderRead::No, {}, "slice" }, // also defines lang= { "vars", "vars=N: sliceable local bindings in the definition, one v row each; name one to slice it", false, "slice", MapHeaderRead::No, {}, "slice" }, + { "order", "order=defuse: seed s rows (no v=) ranked by def-use coverage (distinct local names on the line) desc, then line; not source order — flow s rows (v=) keep their (d=,l=,v=) order", false, "slice", MapHeaderRead::No, {}, "slice" }, // slice.h sliceDefUseRowOrder { "counts", "counts=as-classified: defs=/uses=/vars=/steps= count what the name classifier rowed; neither floors nor totals", false, "slice", MapHeaderRead::No, {}, "slice" }, // flags: src/darkflags.h (the root emit) { "gates", "gates=/dark_gates=: gate rows (never cut) / those whose default keeps the guarded code out of the build", false, "flags", MapHeaderRead::No, {}, "flags" }, // also defines dark_gates= diff --git a/src/quality.h b/src/quality.h index 934107cbc..986383664 100644 --- a/src/quality.h +++ b/src/quality.h @@ -5286,9 +5286,13 @@ struct Regression // and the duplication kind reported this function as a fifth copy of it the moment it was written that way. struct FacetAttr { std::string_view kind; const char* attr; }; inline constexpr FacetAttr kFacetAttrs[] = { - { "short-horizon-churn", "churn" }, // self / ambient - { "api-surface", "surface" }, // new-symbol / contract-change - { "duplication", "idiom" }, // the recognized clone-body shape (cloneidiom.h) + { "short-horizon-churn", "churn" }, // self / ambient + { "api-surface", "surface" }, // new-symbol / contract-change + { "duplication", "idiom" }, // the recognized clone-body shape (cloneidiom.h) + { "new-clone-of-reused-helper", "idiom" }, // T13/fix2 gave this kind duplication's idiom demotion (Regression::facet + // is populated the same way, from the same CloneIdiomVerdict); this row was + // missing, so a demoted row's idiom name never reached the reader even + // though sev="minor" fired correctly — a disclosure gap, not a gating one. }; inline const char* facetAttrName( std::string_view kind ) noexcept diff --git a/src/slice.h b/src/slice.h index cf85f7a7c..866d0c943 100644 --- a/src/slice.h +++ b/src/slice.h @@ -6,7 +6,7 @@ // MOTIVATION. ARISE (arXiv:2605.03117) measured statement-level definition-use edges exposed as a // queryable agent primitive at +17pp Function Recall@1 on SWE-bench Lite. ripwire's graph stops at // symbol granularity; this is the bounded v1 of that primitive: one definition, one variable, its -// def/use statement rows in source order. +// def/use statement rows, emitted in the order="defuse" ranking (sliceDefUseRowOrder). // // HONESTY CONTRACT (all four limits are stated in the emitted legend, never implied): // • NAME-BASED — occurrences are identifier-name matches inside the definition's span. No alias @@ -3055,11 +3055,13 @@ inline std::string sliceLegendText( const SliceEmitOpts& opts ) "counts=as-classified — defs/uses/vars/steps count what the classifier rowed, neither floors nor totals. " ": k=def|use|both|scope, t=param|decl|assign|call-arg|read|global|nonlocal, b=declaration line a " "shadowed name binds to (0=unbound), pp=1 build-dependent preprocessor region, rd=lines of the defs reaching a use row " - "(-=none) per reach=cfg (flow-sensitive; C-family, Python) | linear (source order; JS/TS, Go, Java, Rust). Inventory " + "(-=none) per reach=cfg (flow-sensitive; C-family, Python) | linear (source order; JS/TS, Go, Java, Rust). " + "order=defuse: seed rows ranked by def-use coverage (distinct local names on the line) desc, then line — not " + "source order. Inventory " ", vars=count. " "bindings=shadow count; preproc_rows=lines dropped under #if 0; seed/var_from/seed_vars/seed=1 = line-seed disclosure. " "Flow rows add v=variable d=depth f=from-line; steps=flow rows, depth=bound, flow_truncated=1 bounded not complete, " - "flow_redundant=1 (both=, unseeded only) reaches no line the flat inventory does not — seed via --at=FILE:LINE for real reach. " + "flow_redundant=1 (both=, unseeded only) reaches no line the flat inventory does not — seed via at=FILE:LINE for real reach. " "Limits: a write hidden behind a call (receiver mutation, by-ref/out-param, macro) rows as a use; no alias analysis; " "the statement is the unit (nested bodies/?:/short-circuit fold, goto untracked); no control dependence; block " "scopes separated; C-family #if 0 dropped, other #if kept+flagged. Full legend: omit " @@ -3069,7 +3071,8 @@ inline std::string sliceLegendText( const SliceEmitOpts& opts ) { out = ""; + "flow_redundant=\"1\" (unseeded both=): adds no line beyond the flat inventory; seed at= for real gain. -->"; } } // H1: the residue clause, in BOTH dialects and as its own comment — opened `