research(readability): construct-validity investigation for --readability - #313
Draft
joyful-ii-V-I wants to merge 10 commits into
Draft
joyful-ii-V-I wants to merge 10 commits into
joyful-ii-V-I wants to merge 10 commits into
Conversation
…lity Reads the code rather than assuming it: --readability is already disclosed as an ordering-only lens (never a grade, not one of --quality-delta's ten gating kinds, folded into --ensemble as a rank not a score). Designs a pairwise + rank-correlation human validation of that ordering claim, and runs the two proxies that need no human label: refactor-commit direction (80 commits, 484 function pairs, lens agrees with the commit's implied direction on only 30.2% of them) and self-consistency under meaning-preserving rewrites (95 mutation attempts across 80 sampled functions, 100% exact tie, as the formula predicts mathematically). Reports both numbers, including the unflattering one, and treats naminglens.h's withdrawn naming-body-mismatch rule as the template for what acting on a failed validation would look like. No src/ changes. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: trueComment |
…rch row ripwirepubliccheck arm 8 fired: the reference list cited an internal working document by name eight times, and that document is not in this repository, so every one of those citations is a dangling link for any reader outside it. The citations stay — they are where the claims come from — spelled as "the readability design note" instead of a filename nobody else can open. docs/README.md gains the canonical `research/` row arm 6b requires, byte-identical to the one on every other research lane. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…on before computing them Fixes the selection regex, the lens-name mask, the n target (100 directional pairs), and the decision bands (>=60% holds, <=40% inverted -> stop rule, between inconclusive) in the note before the harness exists or runs. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Proxy (c) (declared-readability commits, pre-registered in 867c84b) reached n=4 of the 100 target, all regex false positives: no verdict, the stop rule does not trigger. §3a regenerated at 413 pairs (38.3% right-direction); 91% of wrong-direction pairs are driven by the Halstead-volume term, and the sign of the token-count change predicts the lens's direction on 94.2% of pairs. No src/ change; a legend narrowing is proposed for owner sign-off only. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
bench/readability_refactor_pairs.py and bench/readability_declared_pairs.py walked `git log --all`, which sweeps every branch in this clone's shared .git (~291 worktree lanes at the time this was found) — so the published pair count moved whenever an unrelated lane was pushed, independent of the lens. A number that moves when an unrelated branch is pushed is not a measurement. Both scripts now default to walking exactly one immutable ref (the v0.6.2 tag, matching the scoring binary's own build commit), overridable with --ref; --until (added for the prior workaround) is kept for narrowing within whichever ref is walked. docs/research/readability-construct-validity.md's §3a headline is now this pinned run (412 pairs, 154 right-direction, 37.4%) with the original unpinned figure kept visible as a one-line "first recorded as" note; §3c documents the instrument defect and reports its proxy (c) and decomposition numbers on the same pinned population. The inversion headline still holds (37.4% < 40%), just less starkly than the unpinned 30.2% — every other citation of the old numbers in the note (70%, ~338, 94.2%, etc.) is updated to match. test/ripwirepubliccheck.sh: ALL PASS. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
joyful-ii-V-I
pushed a commit
that referenced
this pull request
Sep 22, 2026
…om shipped help The previous fix for the non-reproducing 484/30.2% figure landed shipped --help text that pointed a reader at `(lane/research-readability-validity @9aecbc96)` and carried an internal label, `MEASURED (t14-cleanup #8, revised)`. Neither resolves for a reader of a released binary, and the research lane may never merge — nothing that ships should depend on an unmerged branch to be understood. --help now states only what it needs to: the pinned v0.6.2 measurement (412 pairs, 96.0% token-count-sign agreement) and how to read a move, trimmed from a nine-line paragraph to six, with the secondary figures (37.4%, 91.1%, the retracted 484/30.2% history) left to docs/EVALS.md §8 where they already lived. README's matching parenthetical now points at EVALS §8 alone instead of the branch/commit. EVALS §8 itself keeps citing the research note's derivation, but by PR (#313, open) rather than by branch name and commit — the standard way to reference draft work that may not land. docs/COMMANDS.md regenerated via docs/docs_commands_build.py; test/printf_parity.manifest re-pinned (UPDATE_GOLDEN=1) — the verified diff moves exactly help_all. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The reference list called arXiv:2605.13280 the "Readability Spectrum" study -- a title the paper does not have. Its real title, from the arXiv export API, is "Characterizing Readability Issue Patterns and the Role of Prompt Design in LLM-Generated Code" (Ye, Ran, Xu, Zhou). Corrected all three occurrences (the "published results" intro, the design-note cross-reference, and the reference-list entry) and reworded each to state only what the paper's abstract supports: prompt design's overall role in generated-code readability is bounded. No sentence in the note now attributes to this paper a finding it does not contain. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Neither §2 nor §3 tests the claim an agent actually relies on when it uses --readability to pick a refactor target: does a low score predict that a function needs fixing later? Adds bench/readability_fixrate_validity.py and this document's new §4 protocol, fixed before any number exists: population at a pinned v0.6.2-history cutoff (never --all), exposure = z (readability) and ccx (the same cognitive-complexity metric --quality-delta's "complexity" kind already trusts) read at the cutoff, outcome = touched by a fix-shaped commit in a pinned multi-week follow-up window (function spans re-resolved per commit so line drift cannot misattribute a hunk), statistic = risk ratio with Wilson/Katz CIs, raw and size-stratified. Discloses in writing, before running it, that a lines-based size stratification cannot fully rule out the token-volume confound §3c already found behind z's direction. No src/ change. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
bench/readability_fixrate_validity.py run against v0.6.2 history: 2,868 functions at cutoff 4f5c310 (2026-09-03), a 1,138-commit / 556-fix-shaped follow-up window to v0.6.2 (2026-09-21). Raw: least-readable-quartile risk ratio 2.87 [2.57, 3.21] for later fixes, 2.25 [2.08, 2.44] for any later modification — both larger than the complexity control's 2.26 / 1.93 on the identical population. The pre-registered "most likely outcome" (signal vanishes once size is held constant) did not happen: the readability RR clears 1 in five of six lines-stratified rows, only touching it in one (FIXED/T2, lower bound 0.995). Complexity inverts (RR significantly below 1) in the same T2 band on both outcomes, unanticipated by this protocol and reported as measured. Verdict per the pre-registered bands: keep the ordering claim as an actionable signal on this corpus, not withdraw it — with the disclosed caveat that a lines-based stratification does not fully rule out the token-volume mechanism §3c already found, so a token-count stratification is named as the next test, not run here. No src/ change. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…ount The line-based size stratification in the prior commit did not hold the confound it set out to test: §3c already found z tracks the SIGN of the token-count change on 96.0% of this repo's own refactor pairs, not the line-count change, so a lines tercile does not hold z's own size axis constant. Adds a toks= (Halstead N, --readability's own count, the exact integer the lens's volume term is computed from) stratification beside the existing lines one on the identical population/outcomes. The two disagree, as intended: line-stratified, readability's RR excluded 1 in 5/6 rows; token-stratified (the decisive one), it excludes 1 in only 4/6 - both T2 (middle-third-by-tokens) rows now include 1 (1.06 [0.82,1.36] fixed, 1.05 [0.88,1.26] modified), where the line version had reported a borderline signal. Verdict restated against the pre-registered bands: not "keep exactly as-is" (requires no vanishing) and not "withdraw" (T1 and T3 both still clear RR>1 with CIs excluding 1) - this document's own §2 middle band, "keep but narrow": the later-fix signal is real at the size extremes and silent in the middle third, so the raw-population RR overstates what an agent should expect from a mid-sized function. Complexity's T2 inversion, checked for a file or symbol-kind concentration (max file share 5%, max keyword share 3.3% across ~40 files in both the line- and token-based T2 highest-ccx quartiles): no cheap explanation found, left as measured. Adds an external-validity line: every number in this section is this repository's own AI-authored, heavily-gated history, not a general claim about code. No src/ change; the pre-registration and first-results commits are untouched. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…hold size POST-HOC (prompted by looking at the tercile table, not pre-registered): the token terciles that showed a signal are also the widest internally (T1 7.2x, T3 42x); the silent one (T2) is the narrowest (2.6x). A tercile does not hold the lens's own size unit constant, so "survives at both extremes" is what a residual size effect predicts too, not only what an independent readability effect would. Adds decile-by-toks stratification (10 bands instead of 3) plus internal range per band to bench/readability_fixrate_validity.py, and a --load-tsv flag to re-stratify already-computed data in under a second instead of repeating the ~15-minute crawl. Result: 8 of 10 deciles (the genuinely narrow ones, 1.3x-2.9x internal range) show no effect on either outcome - CIs include 1, point estimates scattered both sides of 1. The two deciles that remain significant are the two with residual size range left inside them (D10 unambiguously at 16.3x; D09 flagged as not cleanly resolved at n=287/decile). A nearest-token-neighbour matched-control check was also attempted and found degenerate (46 distinct matches for 717 quartile members, top 2 matches covering 73.8%) - reported as discarded, not as evidence. Verdict restated against the pre-registered bands using deciles as decisive: withdraw, not "keep, narrowed" as the prior commit concluded. The raw and tercile-level separation is better explained as a residual size effect in the lens's own units than an independent later-fix signal - the same token-count mechanism §3c already found driving the lens's refactor-pair direction. The prior "keep, narrowed" verdict is superseded in this document, not rewritten. Pre-registration and both earlier results commits are untouched. No src/ change. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
joyful-ii-V-I
added a commit
that referenced
this pull request
Sep 23, 2026
Independent review (fable, $ORCH/reports/rv-flag-biggest-first.md, VERDICT READY-WITH-CHANGES) of 35ae521 confirmed the rename itself: stdout byte-identical between spellings across 17 shapes, alias mechanism matches deprecatedOrderFlag exactly, every named gate green, emitted names/schema untouched. Four findings, all text/test-only: - F1: CHANGELOG.md's [Unreleased] "Fixed" entry cited `docs/research/readability-construct-validity.md` §4, which is not in this tree (only in draft PR #313, OPEN) -- re-introducing exactly the kind of unmerged-branch citation train 16 removed from --help. Dropped the dangling path; the entry already closes with "Full derivation: docs/EVALS.md §8", matching how the base lane cites the same mechanism in cli.h ("derivation in docs/EVALS.md §8"). ripwirepubliccheck arm 6b (docs/ completeness) and arm 8 (no dangling references) both green. - F2: docs/LINEAGE.md:97 still stated "`--readability`, which emits P least-readable-first" as live behaviour -- the withdrawn ordering claim, in a row §6a's site survey happened not to cover. Reworded to name the current flag, state the actual order (largest Halstead volume/token-count/length first), and point at the withdrawal (docs/EVALS.md §8), matching this lane's own wording elsewhere. - F3 (deck artifacts stale): investigated, not fixed. present/README.md's rebuild recipe needs `node`/`npm` (pptxgenjs) and `soffice` (LibreOffice) for the PDF; none of the three is installed on this machine (checked PATH, Homebrew, /Applications). Per instruction, not faked -- present/ripwire-showcase.pptx and its PDF are left as they were (the alias keeps their --readability slide truthful in the meantime; deckclaimcheck reads them and stays green). - F4: readabilitycheck.sh arm H's `--json` stdout-parity assertion was vacuous -- `--json` is refused for this lens on both spellings (rc=1, 0 B stdout), so `cmp` of two empty files proved nothing. Replaced with three real assertions: the refusal's exit code matches on both spellings, stdout is empty on both, and the refusal SENTENCE on stderr is byte-identical once the old spelling's one-shot deprecation line is stripped off the top. cli.h ~1382 wording is untouched, per the orchestrator: train 17 is landing its own fix to the same help body and the train 18 builder owns that merge. Re-run and green: readmedriftcheck, deckcheck, deckclaimcheck, docdriftcheck, ripwirepubliccheck (arms 6b/8 specifically), readabilitycheck. No hung ASan children found in this worktree (none left running from the earlier session). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This is an investigation, not a behavior change — no
src/changes, docs + bench scripts only.Two published results motivate this round: LLM judges of code readability lean on surface features rather than the construct itself (CoReEval, arXiv:2510.16579), and prompt-side style constraints on LLM code generation help but plateau (Readability Spectrum, arXiv:2605.13280). Neither is folded into
docs/LINEAGE.mdas its own row — they are design influences on how we read--readability's claim, not implementations of either paper, so this PR does not invent a row for them. The feature they bear on does have a row:docs/LINEAGE.mdL97, Posnett/Hindle/Devanbu, which is where--readability's existing "ORDER, never a grade" disclosure comes from. Full write-up:docs/research/readability-construct-validity.md.(Base note: measured against
mainat755f9026, before 0.6.2; the diff still merges cleanly onto today'smain,15a20855, tag 0.6.2.)Headline finding, and it is unflattering. We designed a pairwise + rank-correlation validation of the ordering claim (not yet run — needs a blinded human study) and, ahead of that, ran two label-free proxies against the real binary. Refactor-commit direction: 80 refactor/simplify/cleanup commits, 484 matched before/after function pairs, and the lens agrees with the commit's own implied direction only 30.2% of the time — median Δz is negative, i.e. inverted more often than not. Self-consistency under rename/reorder mutations that should not change readability: 95 mutations, 100% exact tie, which the formula guarantees mathematically and so mostly certifies implementation correctness rather than construct validity. We also verified precisely what
--readabilitydoes and does not feed: it is not one of--quality-delta's ten gating kinds, and the one adjacent kind,verbosity, shares only one of its three formula inputs.We are not shipping a fix on the strength of one proxy measurement —
docs/LINEAGE.md's own withdrawnnaming-body-mismatchrule (§9.0 /src/naminglens.h) is used in the doc as the template for what acting on a failed validation would require, and this round stops at "validate first."This is a first pass with a lot further to push, and the doc closes with four concrete questions for readability researchers: whether the pairwise + rank-correlation design is right, whether the pass/fail bands are calibrated defensibly, whether commit-message mining is too noisy a label for the 30.2% number to stand on without a stricter filter, and whether Posnett 2011 is still the right deterministic formula to be defending in 2026.
Our working premise across this line of investigations: algorithmic, deterministic checks applied while an AI writes are the practical way to keep code sound at the speed AI now writes it — no label, no vendor claim, a check that either fires or it doesn't.
🤖 Generated with Claude Code