Skip to content

research(readability): construct-validity investigation for --readability - #313

Draft
joyful-ii-V-I wants to merge 10 commits into
mainfrom
lane/research-readability-validity
Draft

joyful-ii-V-I wants to merge 10 commits into
mainfrom
lane/research-readability-validity

Conversation

@joyful-ii-V-I

Copy link
Copy Markdown
Collaborator

This is an investigation, not a behavior change — no src/ changes, docs + bench scripts only.

Two published results motivate this round: LLM judges of code readability lean on surface features rather than the construct itself (CoReEval, arXiv:2510.16579), and prompt-side style constraints on LLM code generation help but plateau (Readability Spectrum, arXiv:2605.13280). Neither is folded into docs/LINEAGE.md as its own row — they are design influences on how we read --readability's claim, not implementations of either paper, so this PR does not invent a row for them. The feature they bear on does have a row: docs/LINEAGE.md L97, Posnett/Hindle/Devanbu, which is where --readability's existing "ORDER, never a grade" disclosure comes from. Full write-up: docs/research/readability-construct-validity.md.

(Base note: measured against main at 755f9026, before 0.6.2; the diff still merges cleanly onto today's main, 15a20855, tag 0.6.2.)

Headline finding, and it is unflattering. We designed a pairwise + rank-correlation validation of the ordering claim (not yet run — needs a blinded human study) and, ahead of that, ran two label-free proxies against the real binary. Refactor-commit direction: 80 refactor/simplify/cleanup commits, 484 matched before/after function pairs, and the lens agrees with the commit's own implied direction only 30.2% of the time — median Δz is negative, i.e. inverted more often than not. Self-consistency under rename/reorder mutations that should not change readability: 95 mutations, 100% exact tie, which the formula guarantees mathematically and so mostly certifies implementation correctness rather than construct validity. We also verified precisely what --readability does and does not feed: it is not one of --quality-delta's ten gating kinds, and the one adjacent kind, verbosity, shares only one of its three formula inputs.

We are not shipping a fix on the strength of one proxy measurement — docs/LINEAGE.md's own withdrawn naming-body-mismatch rule (§9.0 / src/naminglens.h) is used in the doc as the template for what acting on a failed validation would require, and this round stops at "validate first."

This is a first pass with a lot further to push, and the doc closes with four concrete questions for readability researchers: whether the pairwise + rank-correlation design is right, whether the pass/fail bands are calibrated defensibly, whether commit-message mining is too noisy a label for the 30.2% number to stand on without a stricter filter, and whether Posnett 2011 is still the right deterministic formula to be defending in 2026.

Our working premise across this line of investigations: algorithmic, deterministic checks applied while an AI writes are the practical way to keep code sound at the speed AI now writes it — no label, no vendor claim, a check that either fires or it doesn't.

🤖 Generated with Claude Code

…lity

Reads the code rather than assuming it: --readability is already disclosed as
an ordering-only lens (never a grade, not one of --quality-delta's ten gating
kinds, folded into --ensemble as a rank not a score). Designs a pairwise +
rank-correlation human validation of that ordering claim, and runs the two
proxies that need no human label: refactor-commit direction (80 commits, 484
function pairs, lens agrees with the commit's implied direction on only 30.2%
of them) and self-consistency under meaning-preserving rewrites (95 mutation
attempts across 80 sampled functions, 100% exact tie, as the formula predicts
mathematically). Reports both numbers, including the unflattering one, and
treats naminglens.h's withdrawn naming-body-mismatch rule as the template for
what acting on a failed validation would look like. No src/ changes.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Sep 21, 2026

Copy link
Copy Markdown

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true

Comment @coderabbitai help to get the list of available commands.

barefootski and others added 4 commits September 21, 2026 11:49
…rch row

ripwirepubliccheck arm 8 fired: the reference list cited an internal working
document by name eight times, and that document is not in this repository, so
every one of those citations is a dangling link for any reader outside it. The
citations stay — they are where the claims come from — spelled as "the
readability design note" instead of a filename nobody else can open.

docs/README.md gains the canonical `research/` row arm 6b requires, byte-identical
to the one on every other research lane.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…on before computing them

Fixes the selection regex, the lens-name mask, the n target (100 directional
pairs), and the decision bands (>=60% holds, <=40% inverted -> stop rule,
between inconclusive) in the note before the harness exists or runs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Proxy (c) (declared-readability commits, pre-registered in 867c84b) reached
n=4 of the 100 target, all regex false positives: no verdict, the stop rule
does not trigger. §3a regenerated at 413 pairs (38.3% right-direction); 91% of
wrong-direction pairs are driven by the Halstead-volume term, and the sign of
the token-count change predicts the lens's direction on 94.2% of pairs.
No src/ change; a legend narrowing is proposed for owner sign-off only.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
bench/readability_refactor_pairs.py and bench/readability_declared_pairs.py walked
`git log --all`, which sweeps every branch in this clone's shared .git (~291
worktree lanes at the time this was found) — so the published pair count moved
whenever an unrelated lane was pushed, independent of the lens. A number that
moves when an unrelated branch is pushed is not a measurement.

Both scripts now default to walking exactly one immutable ref (the v0.6.2 tag,
matching the scoring binary's own build commit), overridable with --ref; --until
(added for the prior workaround) is kept for narrowing within whichever ref is
walked. docs/research/readability-construct-validity.md's §3a headline is now
this pinned run (412 pairs, 154 right-direction, 37.4%) with the original
unpinned figure kept visible as a one-line "first recorded as" note; §3c documents
the instrument defect and reports its proxy (c) and decomposition numbers on the
same pinned population. The inversion headline still holds (37.4% < 40%), just
less starkly than the unpinned 30.2% — every other citation of the old numbers in
the note (70%, ~338, 94.2%, etc.) is updated to match.

test/ripwirepubliccheck.sh: ALL PASS.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
joyful-ii-V-I pushed a commit that referenced this pull request Sep 22, 2026
…om shipped help

The previous fix for the non-reproducing 484/30.2% figure landed shipped
--help text that pointed a reader at `(lane/research-readability-validity
@9aecbc96)` and carried an internal label, `MEASURED (t14-cleanup #8,
revised)`. Neither resolves for a reader of a released binary, and the
research lane may never merge — nothing that ships should depend on an
unmerged branch to be understood.

--help now states only what it needs to: the pinned v0.6.2 measurement
(412 pairs, 96.0% token-count-sign agreement) and how to read a move,
trimmed from a nine-line paragraph to six, with the secondary figures
(37.4%, 91.1%, the retracted 484/30.2% history) left to docs/EVALS.md §8
where they already lived. README's matching parenthetical now points at
EVALS §8 alone instead of the branch/commit. EVALS §8 itself keeps citing
the research note's derivation, but by PR (#313, open) rather than by
branch name and commit — the standard way to reference draft work that
may not land.

docs/COMMANDS.md regenerated via docs/docs_commands_build.py;
test/printf_parity.manifest re-pinned (UPDATE_GOLDEN=1) — the verified
diff moves exactly help_all.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
quaterniondrift and others added 5 commits September 22, 2026 09:57
The reference list called arXiv:2605.13280 the "Readability Spectrum"
study -- a title the paper does not have. Its real title, from the
arXiv export API, is "Characterizing Readability Issue Patterns and
the Role of Prompt Design in LLM-Generated Code" (Ye, Ran, Xu, Zhou).

Corrected all three occurrences (the "published results" intro, the
design-note cross-reference, and the reference-list entry) and
reworded each to state only what the paper's abstract supports:
prompt design's overall role in generated-code readability is
bounded. No sentence in the note now attributes to this paper a
finding it does not contain.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Neither §2 nor §3 tests the claim an agent actually relies on when it uses
--readability to pick a refactor target: does a low score predict that a
function needs fixing later? Adds bench/readability_fixrate_validity.py and
this document's new §4 protocol, fixed before any number exists: population
at a pinned v0.6.2-history cutoff (never --all), exposure = z (readability)
and ccx (the same cognitive-complexity metric --quality-delta's "complexity"
kind already trusts) read at the cutoff, outcome = touched by a fix-shaped
commit in a pinned multi-week follow-up window (function spans re-resolved
per commit so line drift cannot misattribute a hunk), statistic = risk ratio
with Wilson/Katz CIs, raw and size-stratified. Discloses in writing, before
running it, that a lines-based size stratification cannot fully rule out the
token-volume confound §3c already found behind z's direction. No src/ change.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
bench/readability_fixrate_validity.py run against v0.6.2 history: 2,868
functions at cutoff 4f5c310 (2026-09-03), a 1,138-commit / 556-fix-shaped
follow-up window to v0.6.2 (2026-09-21). Raw: least-readable-quartile risk
ratio 2.87 [2.57, 3.21] for later fixes, 2.25 [2.08, 2.44] for any later
modification — both larger than the complexity control's 2.26 / 1.93 on the
identical population. The pre-registered "most likely outcome" (signal
vanishes once size is held constant) did not happen: the readability RR
clears 1 in five of six lines-stratified rows, only touching it in one
(FIXED/T2, lower bound 0.995). Complexity inverts (RR significantly below 1)
in the same T2 band on both outcomes, unanticipated by this protocol and
reported as measured. Verdict per the pre-registered bands: keep the
ordering claim as an actionable signal on this corpus, not withdraw it — with
the disclosed caveat that a lines-based stratification does not fully rule
out the token-volume mechanism §3c already found, so a token-count
stratification is named as the next test, not run here. No src/ change.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…ount

The line-based size stratification in the prior commit did not hold the
confound it set out to test: §3c already found z tracks the SIGN of the
token-count change on 96.0% of this repo's own refactor pairs, not the
line-count change, so a lines tercile does not hold z's own size axis
constant. Adds a toks= (Halstead N, --readability's own count, the exact
integer the lens's volume term is computed from) stratification beside the
existing lines one on the identical population/outcomes.

The two disagree, as intended: line-stratified, readability's RR excluded 1
in 5/6 rows; token-stratified (the decisive one), it excludes 1 in only 4/6 -
both T2 (middle-third-by-tokens) rows now include 1 (1.06 [0.82,1.36] fixed,
1.05 [0.88,1.26] modified), where the line version had reported a borderline
signal. Verdict restated against the pre-registered bands: not "keep exactly
as-is" (requires no vanishing) and not "withdraw" (T1 and T3 both still
clear RR>1 with CIs excluding 1) - this document's own §2 middle band, "keep
but narrow": the later-fix signal is real at the size extremes and silent in
the middle third, so the raw-population RR overstates what an agent should
expect from a mid-sized function.

Complexity's T2 inversion, checked for a file or symbol-kind concentration
(max file share 5%, max keyword share 3.3% across ~40 files in both the line-
and token-based T2 highest-ccx quartiles): no cheap explanation found, left
as measured. Adds an external-validity line: every number in this section is
this repository's own AI-authored, heavily-gated history, not a general
claim about code. No src/ change; the pre-registration and first-results
commits are untouched.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…hold size

POST-HOC (prompted by looking at the tercile table, not pre-registered):
the token terciles that showed a signal are also the widest internally
(T1 7.2x, T3 42x); the silent one (T2) is the narrowest (2.6x). A tercile
does not hold the lens's own size unit constant, so "survives at both
extremes" is what a residual size effect predicts too, not only what an
independent readability effect would.

Adds decile-by-toks stratification (10 bands instead of 3) plus internal
range per band to bench/readability_fixrate_validity.py, and a --load-tsv
flag to re-stratify already-computed data in under a second instead of
repeating the ~15-minute crawl. Result: 8 of 10 deciles (the genuinely
narrow ones, 1.3x-2.9x internal range) show no effect on either outcome -
CIs include 1, point estimates scattered both sides of 1. The two deciles
that remain significant are the two with residual size range left inside
them (D10 unambiguously at 16.3x; D09 flagged as not cleanly resolved at
n=287/decile). A nearest-token-neighbour matched-control check was also
attempted and found degenerate (46 distinct matches for 717 quartile
members, top 2 matches covering 73.8%) - reported as discarded, not as
evidence.

Verdict restated against the pre-registered bands using deciles as
decisive: withdraw, not "keep, narrowed" as the prior commit concluded.
The raw and tercile-level separation is better explained as a residual
size effect in the lens's own units than an independent later-fix signal -
the same token-count mechanism §3c already found driving the lens's
refactor-pair direction. The prior "keep, narrowed" verdict is superseded
in this document, not rewritten. Pre-registration and both earlier results
commits are untouched. No src/ change.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
joyful-ii-V-I added a commit that referenced this pull request Sep 23, 2026
Independent review (fable, $ORCH/reports/rv-flag-biggest-first.md, VERDICT
READY-WITH-CHANGES) of 35ae521 confirmed the rename itself: stdout
byte-identical between spellings across 17 shapes, alias mechanism matches
deprecatedOrderFlag exactly, every named gate green, emitted names/schema
untouched. Four findings, all text/test-only:

- F1: CHANGELOG.md's [Unreleased] "Fixed" entry cited
  `docs/research/readability-construct-validity.md` §4, which is not in
  this tree (only in draft PR #313, OPEN) -- re-introducing exactly the
  kind of unmerged-branch citation train 16 removed from --help. Dropped
  the dangling path; the entry already closes with "Full derivation:
  docs/EVALS.md §8", matching how the base lane cites the same mechanism
  in cli.h ("derivation in docs/EVALS.md §8"). ripwirepubliccheck arm 6b
  (docs/ completeness) and arm 8 (no dangling references) both green.
- F2: docs/LINEAGE.md:97 still stated "`--readability`, which emits P
  least-readable-first" as live behaviour -- the withdrawn ordering claim,
  in a row §6a's site survey happened not to cover. Reworded to name the
  current flag, state the actual order (largest Halstead
  volume/token-count/length first), and point at the withdrawal
  (docs/EVALS.md §8), matching this lane's own wording elsewhere.
- F3 (deck artifacts stale): investigated, not fixed. present/README.md's
  rebuild recipe needs `node`/`npm` (pptxgenjs) and `soffice`
  (LibreOffice) for the PDF; none of the three is installed on this
  machine (checked PATH, Homebrew, /Applications). Per instruction, not
  faked -- present/ripwire-showcase.pptx and its PDF are left as they
  were (the alias keeps their --readability slide truthful in the
  meantime; deckclaimcheck reads them and stays green).
- F4: readabilitycheck.sh arm H's `--json` stdout-parity assertion was
  vacuous -- `--json` is refused for this lens on both spellings (rc=1,
  0 B stdout), so `cmp` of two empty files proved nothing. Replaced with
  three real assertions: the refusal's exit code matches on both
  spellings, stdout is empty on both, and the refusal SENTENCE on stderr
  is byte-identical once the old spelling's one-shot deprecation line is
  stripped off the top.

cli.h ~1382 wording is untouched, per the orchestrator: train 17 is
landing its own fix to the same help body and the train 18 builder owns
that merge.

Re-run and green: readmedriftcheck, deckcheck, deckclaimcheck,
docdriftcheck, ripwirepubliccheck (arms 6b/8 specifically), readabilitycheck.
No hung ASan children found in this worktree (none left running from the
earlier session).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants