Skip to content

docs(research): reuse-decline, specified precisely — and the two cross-context smells we do not measure - #312

Draft
joyful-ii-V-I wants to merge 4 commits into
mainfrom
lane/research-ai-smells
Draft

joyful-ii-V-I wants to merge 4 commits into
mainfrom
lane/research-ai-smells

Conversation

@joyful-ii-V-I

Copy link
Copy Markdown
Collaborator

An investigation note, not a feature and not a plan. No src/ change.

Three of --quality-delta's ten kinds — verbosity, duplication and reuse-decline — were shaped by Tsantalis, Zhu & Rigby, AI-generated code smells (arXiv:2605.02741), the design already recorded at docs/LINEAGE.md L84. Of the three, reuse-decline (new-clone-of-reused-helper) is the only one that is ours rather than a standard metric, so it is the only one worth putting in front of anyone to critique. Full write-up: docs/research/ai-smells-reuse-decline.md.

(Base note: measured against main at 755f9026, before 0.6.2; the diff still merges cleanly onto today's main, 15a20855, tag 0.6.2.)

Headline finding. This branch writes down precisely what reuse-decline computes, proposes two unmeasured smells from the same line of work, and measures all three on real trees — including two unflattering results. Cross-file share of clone groups is 57.1% in this repository and 66.7% pooled over a 90-repository corpus, but the per-repository median is only 37.3% across a 0%–94.5% spread, so the pooled figure is not a number to threshold on. Replaying the kind over 150 commits of this repository fires it exactly once, and that one row is a shell-gate harness clone that the sibling duplication kind exempts and reuse-decline does not — the same asymmetry the ack ledger corroborates: 43 of 62 accepted findings sit at fan-in 3 or 4, one notch above kReusedHelperMinFanin, a threshold that was reasoned rather than fitted.

A real src/ finding, written up and deliberately not fixed here (this lane's scope is docs/bench only): duplication's reporter skips all-test-script clone groups and demotes recognized idioms to minor; reuse-decline's reporter, reading the same clone vectors, does neither. It wants its own lane.

This is a first pass with a lot further to push. §4 of the document lists, in our own words, the six questions we would most like to be told we have wrong — the fan-in threshold first, then the two proposed-but-unmeasured smells (clone-group locality, and a retrospective "one rule living in N files" read). Corpus A (90 shallow-cloned checkouts) blocks any future history-based measurement on it, which is stated plainly rather than worked around.

Our working premise across this line of investigations: algorithmic, deterministic checks applied while an AI writes are the practical way to keep code sound at the speed AI now writes it — no label, no vendor claim, a check that either fires or it doesn't.

🤖 Generated with Claude Code

…misses

An investigation note on the AI-code-smell taxonomy row already recorded in
docs/LINEAGE.md (arXiv:2605.02741), written so an outside reader can
re-implement --quality-delta's new-clone-of-reused-helper kind or show that it
is wrong. No src/ change.

Part 1 specifies the kind: the two clone passes and their thresholds, the
member-set identity, the "existed at the baseline" precondition and why it had
to be added, the fan-in rule (measured: fan-in counts distinct callers, not
call sites), the two scope exemptions, and the fact that a row of this kind is
always preexisting-worse and always major, so it always gates. It states what
the kind cannot catch -- above all a re-implementation that is not a token
clone -- and its false-alarm channels. bench/aismells/reuse_decline_example.sh
builds three throwaway fixtures and prints the real output for each.

Part 2 proposes two smells the tool does not measure, both cross-context:
clone-group LOCALITY as a facet on --clones and on the duplication row, and a
retrospective "one rule in N files" read over the co-change miner's existing
history walk. Each carries its own failure modes and the measurement that
would have to come first.

Part 3 measures. Cross-file share of clone groups is 57.2% here, 66.7% pooled
over a 90-repository corpus -- but the per-repository median is 37.3% with a
0%-94.5% spread, so the pooled figure is not a threshold anyone should set.
Replaying the kind over 150 commits of this repository fires it ONCE, and that
one row is a shell-gate harness clone the sibling duplication kind exempts and
this one does not. The ack ledger says the same: 43 of 62 accepted findings sit
at fan-in 3 or 4, one notch above a threshold that was reasoned rather than
fitted. The fix-spread signal is reported as weak, including the arm where it
reverses sign between two corpora, and the arm that could not run at all
because every checkout in that corpus is a depth-1 clone.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Sep 21, 2026

Copy link
Copy Markdown

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true

Comment @coderabbitai help to get the list of available commands.

barefootski and others added 3 commits September 21, 2026 11:52
Eight research lanes each added their own wording for the same `docs/README.md`
table line, which is eight conflicts on one line the moment two of them land.
This is the shared text, byte-identical everywhere, so the same addition on two
branches merges clean. The entry count above the table is corrected with it:
the table has held twenty rows for some time while the sentence still said
sixteen.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…rnal number

- arXiv:2605.02741's first author is Zhu, not Tsantalis (Zhu, Tsantalis
  and Rigby).
- The 377-recorded / 292-usable / 169-tagged chain was being read as
  169-of-377; state each step so 88%/12% is clearly of the 169 tagged
  rows, with 123 (42.1% of the 292 usable) left untagged.
- Drop the printed 17-38% external cross-context figure: our two
  proxies for that rate disagree and neither has a written recipe, so
  we do not quote a number for it.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
§1.7, §3.1, §3.4 and the question list in §4 described duplication's idiom demotion and
all-test-script skip as missing from new-clone-of-reused-helper (reuse-decline), calling it an
oversight. That fix landed in 5a31662 ("fix(quality-delta): reuse-decline gets duplication's two
clone-group demotions") and shipped in v0.6.2 (5a31662 is an ancestor of tag v0.6.2, b6c68d8 on
15a2085). A follow-up change under review adds behavioral gate arms for both demotions and fixes
an idiom= disclosure gap the gate work exposed (kFacetAttrs had no row for this kind).

Kept the pre-fix 150-commit replay measurement (1 firing, the all-test-script false positive) as
history and added the post-fix replay result: 0 firings over the same window. Reframed the §4 ask
now that we made the demotions symmetric — is there an argument an idiom collision involving a
well-reused helper should be MORE serious, making the demotion the wrong default. Did not cite the
unmerged follow-up lane by name.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants