docs(research): a withdrawn naming-consistency rule, reconstructed as evidence — first pass, negative result - #317
Draft
joyful-ii-V-I wants to merge 4 commits into
Draft
joyful-ii-V-I wants to merge 4 commits into
joyful-ii-V-I wants to merge 4 commits into
Conversation
A first-pass research document, not a feature: docs/research/naming-consistency-withdrawal.md follows Wang, Zhang, Jiang, Tang, Li & Liu, "Deep Learning-Based Identification of Inconsistent Method Names: How Far Are We?" (EMSE 2025, doi:10.1007/s10664-024-10592-z, arXiv:2501.12617) — an evaluation-distribution collapse in DL-based naming detectors — and traces the same failure mode through our own withdrawn naming-body-mismatch rule (naminglens.h, shipped and pulled same day, commit a63a9f1/7eeb976f). The withdrawal is reconstructed as evidence rather than retold: a63a9f1 was rebuilt in an isolated worktree and rerun against ripwire's own src/, reproducing the original 159/217 (73%) naming-finding count exactly and pulling fresh worked examples (probabilityMass, isPublicApi, msetDiff, sha1Hex, skipOneTrivia) straight off the rebuilt binary's own output rather than repeating the four names in the withdrawal commit message. §3 generalizes the lesson against the paper's own finding and against ripwire's existing rename- history calibration instrument (--naming-calibration, §9.5): what a naming check needs to demonstrate before shipping — a realistic (uncurated) evaluation population, a stated base rate, a precision floor derived from that base rate, and directional validation. §4 measures a base-rate estimate honestly: bench/naming/sample_base_rate.py draws a reproducible random sample of 42 (name, body) pairs from real on-disk checkouts (CPython, Zulip, Alamofire, swift-nio, ripwire's own src/), 38 judged after excluding 2 framework-mandated hooks and 2 extraction artifacts, 0/38 judged inconsistent — reported as a rule-of-three ~8% upper bound, not a point estimate, with the sample's skew toward test functions and reviewed OSS stated as a limit. §5 inventories what ripwire ships today that touches names (the symbol index, 8 remaining naming-* lint rules, naming-uninformative, --naming-consistency, --naming-calibration, doc-comment BM25 weighting in --for) to state plainly that none of it is a semantic name-accuracy judgment. §6 asks for outside critique on the two open questions: whether any deterministic axis could clear a realistic precision floor, and whether a better-labelled dataset exists to tighten the base-rate estimate. No src/ changes. Nothing from any private correspondence appears in this document — citations are to the published paper only. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: trueComment |
ripwirepubliccheck arm 2 fired — `--repos-root` defaulted to one machine's home directory, which is both a leak and useless to anyone else running the sampler. It now reads `RIPWIRE_BENCH_ASSETS`, falling back to a relative `bench-assets`, so the flag documents the layout instead of one checkout of it. docs/README.md gains the canonical `research/` row arm 6b requires. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ming-consistency withdrawal Point the existing --naming-calibration probe at the five corpora bench/ensemblecal/ already tracks, instead of only ripwire's own history. Two corpora (ripwire, gameA) clear the 30-pair floor for the first time outside ripwire alone; pooled, the current naming-* rules almost never fire on genuine renames (3 of 846 rule-pair opportunities, all pointing at the chosen name, not the abandoned one). rustCLI and appleXR have no git history to mine, matching docs/EVALS.md's own "(no git)" rows; tree-sitter (vendored) stays below its own pairs floor. This corroborates the withdrawal's low-base-rate claim by an independent, larger-sample method; it cannot re-score naming-body-mismatch itself, which no longer exists to score. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Verified from git history: naming-body-mismatch was added by a63a9f1 (2026-08-05) and withdrawn by its direct child 7eeb976 the same day; neither commit is an ancestor of origin/main or any release tag, so the rule never reached a release. The opening paragraph said "withdrew before it shipped", contradicting section 2's own "shipped for one commit, measured, and pulled" — align it with the verified history so an outside reader (and the companion email) can match the same phrasing. Also disclose, in section 4, that the base-rate judging was done by a single unblinded reader with no per-pair record published — the note described the method and result but never stated who judged or how. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Follows Wang, Zhang, Jiang, Tang, Li & Liu, Deep Learning-Based Identification of Inconsistent Method Names: How Far Are We? (EMSE 2025, doi:10.1007/s10664-024-10592-z, arXiv:2501.12617): DL-based naming-inconsistency detectors collapse once scored against a realistic, unbalanced distribution instead of the balanced benchmarks earlier evaluations used. The row this traces through is our own withdrawn deterministic rule, already recorded at
docs/LINEAGE.mdL104 (MNire / the CIC name-versus-body proxy). Full write-up:docs/research/naming-consistency-withdrawal.md.This is a first pass, and the honest result is a withdrawal, not a feature. No
src/changes — one document plus a one-off sampling script behind its base-rate estimate. (Base note: measured againstmainat755f9026, before 0.6.2; the diff still merges cleanly onto today'smain,15a20855, tag 0.6.2.)Headline finding, our own negative result.
src/naminglens.hshipped a deterministic name-vs-body vocabulary-overlap rule (naming-body-mismatch) and pulled it the same day — measured on ripwire's own source it produced 159 of 217 (73%) of the naming lens's findings and flagged the best-named functions in the tree. This PR reconstructs that withdrawal as reproducible evidence (the old commit rebuilt and rerun, not just quoted), generalizes the lesson against the paper's own finding and against ripwire's existing rename-history calibration instrument, and measures an honest, capped base-rate estimate of genuine name/body inconsistency in real code: a seeded sample of 38 judged (name, body) pairs across four real OSS checkouts plus our ownsrc/, 0/38 judged genuinely inconsistent, reported as a rule-of-three ~95% upper bound (≈7.9%) and explicitly not a point estimate. It states plainly what ripwire ships today that touches names versus what it does not — none of it is a semantic name-accuracy judgment.This is a first pass with a lot further to push, and we would like critique on two open questions: whether any deterministic, non-learned axis could clear a realistic-population precision floor for this kind of check at all, and whether a better-labelled real-world dataset exists to tighten our base-rate estimate past a self-sampled, 38-function upper bound.
Our working premise across this line of investigations: algorithmic, deterministic checks applied while an AI writes are the practical way to keep code sound at the speed AI now writes it — no label, no vendor claim, a check that either fires or it doesn't. This withdrawal is exactly that premise holding itself to its own bar.
🤖 Generated with Claude Code