Skip to content

docs(research): --adaptive on short/low-content queries — investigation - #311

Draft
joyful-ii-V-I wants to merge 4 commits into
mainfrom
lane/research-adaptive-shortquery
Draft

joyful-ii-V-I wants to merge 4 commits into
mainfrom
lane/research-adaptive-shortquery

Conversation

@joyful-ii-V-I

Copy link
Copy Markdown
Collaborator

Investigates Adaptive-k's (arXiv:2506.08479) stated short-query limitation against ripwire and a second, larger private C++/ObjC/Metal corpus. Adaptive-k is the design behind --adaptive; its docs/LINEAGE.md row is on main. Full write-up: docs/research/adaptive-short-query.md.

This is an investigation, not a behavior change — no src/ file on this branch is touched. (Base note: measured against main at 755f9026, before 0.6.2; the diff still merges cleanly onto today's main, 15a20855, tag 0.6.2.)

Headline finding. The paper's literal case (broad, contentless queries) is already protected — confidence="low" plus a byte-identical --adaptive no-op, measured on both corpora. A different, real failure was found instead: single common words that route name-exact onto a large homonym cluster (many unrelated symbols sharing one literal name) could get confidence="high" and a sharp cut driven by tie-break order rather than relevance, discarding 90%+ of an equally plausible pool (two reproduced examples: 231→6 and 148→10 on two different codebases). Query length/content does not separate the failing cases from the working ones — a homonym-pool-size statistic the code already computed but did not gate on does.

Where this stands today. That finding did not sit as a proposal — --adaptive on main now declines to narrow when a name-exact route's positive pool clears kAdaptiveHomonymPoolFloor (50) and the cut would land at or below twice the floor (see CHANGELOG.md, "--adaptive no longer narrows a large pool of identically-named symbols on tie-break order"). This PR is the investigation record behind that fix, kept as its own artefact with the full method, both measurement tables, the two alternative fixes it argues against, and a pre-registered measurement band for review.

This is a first pass with a lot further to push. The doc's closing section asks two questions we could not settle from the paper alone: whether homonym-tie-break cliffs are a known failure mode of largest-gap cutting outside symbol retrieval, and whether a principled pool-size/shape statistic exists to distinguish "sharp because relevant" from "sharp because tied" — we picked our threshold from twelve empirical points on two codebases, which is thin evidence for a shipped default. Review of that threshold, and of the rejected alternatives, is welcome.

Our working premise across this line of investigations: algorithmic, deterministic checks applied while an AI writes are the practical way to keep code sound at the speed AI now writes it — no label, no vendor claim, a check that either fires or it doesn't.

🤖 Generated with Claude Code

Measures Adaptive-k's stated short-query limitation against ripwire and canyonraid48. Finds
the paper's literal case (broad/contentless queries) already protected via confidence=low +
no-op, but a real, uncovered failure on name-exact routed common words with large homonym
clusters (update 231->6, run 148->10, both >=90% discarded under confidence=high). Proposes
a homonym-pool-gated decline-and-disclose fix and a pre-registered measurement for it.
Investigation only -- no src/ change.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Sep 21, 2026

Copy link
Copy Markdown

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true

Comment @coderabbitai help to get the list of available commands.

ripwirepubliccheck arms 1, 2, 5, 6b and 8 all fired on this branch. The
investigation used a second, private corpus and the artifact carried its
identity through: the tree name in 34 places, the absolute home path, its
file paths and type names in two listings, its symbol names as query text,
and 60 result rows naming files that are not in any public repository.

Every number is kept — the counts are the evidence and none of them
identifies anything. What is removed is identity: the corpus is now
`private-corpus`, its three corpus-specific queries are placeholders, the
served/discarded listings are `‹file A›`-style placeholders that preserve
the structure the argument rests on (same file twice; `sc=` absent on the
top six, present from rank 7), and the recorded JSON's `top5`/anchor fields
are redacted rather than renamed.

docs/README.md gains the `research/` row arm 6b requires. Its wording is
byte-identical across every research lane on purpose: eight lanes each
inventing their own sentence for the same table line is eight conflicts.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@joyful-ii-V-I

Copy link
Copy Markdown
Collaborator Author

ripwirepubliccheck fired five arms on the first push of this branch (1, 2, 5, 6b, 8). The cause is one thing: the investigation used a second, private corpus and this artifact carried its identity through — the tree name in 34 places, an absolute home path, its file paths and type names in two listings in the note, its symbol names used as query text, and 60 rows in the recorded JSON naming files that are in no public repository.

abcd59f6 fixes it forward. Every number is kept; the counts are the evidence and none of them identifies anything. What is gone is identity: the corpus is private-corpus, its three corpus-specific queries are placeholders, the served/discarded listings are ‹file A›-style placeholders that keep the structure the argument rests on (the same file twice; sc= absent on the top six and present from rank 7), and the JSON's top5/anchor fields are redacted rather than renamed. The gate is green on the new tip, all fifteen arms.

Two things for an owner call, neither of them blocking this draft:

  1. The first commit is still in this branch's history on a public remote. House rule is to fix leaks forward and never rewrite published history, which is what abcd59f6 does — but that leaves c04affc2 reachable until this branch is squashed or deleted. This PR is a draft and nothing depends on its history, so the options are: squash-merge whenever it lands, or close it and re-push the scrubbed tree as a fresh branch. I have not done either; deleting or rewriting a public ref is not mine to decide.

  2. The redaction is a stopgap, not the right final shape. A research note whose second corpus nobody else can obtain is not reproducible by a reader, and that is the point of publishing it. Before this leaves draft the second corpus should be a public one — the repo already measures against several — and the numbers re-run. The note's own "Next steps" section asks for a third corpus anyway; this makes it two public ones instead.

🤖 Generated with Claude Code

quaterniondrift and others added 2 commits September 21, 2026 13:07
….6.2

Read after the release, this note said "research only, no fix" and pinned the
binary at 0.6.1 — so anyone landing on the draft PR would conclude we found the
problem and did nothing. The fix landed through a separate lane and is in
v0.6.2: `kAdaptiveHomonymPoolFloor` (50), decline-and-serve-default-top-N, and
`confidence=` no longer claiming `high` on a cut it declined to trust.

Both statements are true at once, so both are stated: the branch still carries
no `src/` change, and the measurements below remain the pre-fix behaviour,
recorded as they were.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Adaptive-k (arXiv:2506.08479) does not state, in its abstract or its
Limitations section, that a largest-gap cut degrades on short or
low-content queries; its Limitations list summarisation tasks,
non-natural-language input, typographical sensitivity, and adversarial
chunks instead. Reword every sentence that attributed the short/
low-content case to the paper's "own stated limitation" or "literal
cases" so it reads as our own hypothesis about what would be hardest
for a largest-gap cut, not something the paper claims.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants