docs(research): --adaptive on short/low-content queries — investigation - #311
joyful-ii-V-I wants to merge 4 commits into
Conversation
Measures Adaptive-k's stated short-query limitation against ripwire and canyonraid48. Finds the paper's literal case (broad/contentless queries) already protected via confidence=low + no-op, but a real, uncovered failure on name-exact routed common words with large homonym clusters (update 231->6, run 148->10, both >=90% discarded under confidence=high). Proposes a homonym-pool-gated decline-and-disclose fix and a pre-registered measurement for it. Investigation only -- no src/ change. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: trueComment |
ripwirepubliccheck arms 1, 2, 5, 6b and 8 all fired on this branch. The investigation used a second, private corpus and the artifact carried its identity through: the tree name in 34 places, the absolute home path, its file paths and type names in two listings, its symbol names as query text, and 60 result rows naming files that are not in any public repository. Every number is kept — the counts are the evidence and none of them identifies anything. What is removed is identity: the corpus is now `private-corpus`, its three corpus-specific queries are placeholders, the served/discarded listings are `‹file A›`-style placeholders that preserve the structure the argument rests on (same file twice; `sc=` absent on the top six, present from rank 7), and the recorded JSON's `top5`/anchor fields are redacted rather than renamed. docs/README.md gains the `research/` row arm 6b requires. Its wording is byte-identical across every research lane on purpose: eight lanes each inventing their own sentence for the same table line is eight conflicts. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
Two things for an owner call, neither of them blocking this draft:
🤖 Generated with Claude Code |
….6.2 Read after the release, this note said "research only, no fix" and pinned the binary at 0.6.1 — so anyone landing on the draft PR would conclude we found the problem and did nothing. The fix landed through a separate lane and is in v0.6.2: `kAdaptiveHomonymPoolFloor` (50), decline-and-serve-default-top-N, and `confidence=` no longer claiming `high` on a cut it declined to trust. Both statements are true at once, so both are stated: the branch still carries no `src/` change, and the measurements below remain the pre-fix behaviour, recorded as they were. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Adaptive-k (arXiv:2506.08479) does not state, in its abstract or its Limitations section, that a largest-gap cut degrades on short or low-content queries; its Limitations list summarisation tasks, non-natural-language input, typographical sensitivity, and adversarial chunks instead. Reword every sentence that attributed the short/ low-content case to the paper's "own stated limitation" or "literal cases" so it reads as our own hypothesis about what would be hardest for a largest-gap cut, not something the paper claims. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Investigates Adaptive-k's (arXiv:2506.08479) stated short-query limitation against ripwire and a second, larger private C++/ObjC/Metal corpus. Adaptive-k is the design behind
--adaptive; itsdocs/LINEAGE.mdrow is onmain. Full write-up:docs/research/adaptive-short-query.md.This is an investigation, not a behavior change — no
src/file on this branch is touched. (Base note: measured againstmainat755f9026, before 0.6.2; the diff still merges cleanly onto today'smain,15a20855, tag 0.6.2.)Headline finding. The paper's literal case (broad, contentless queries) is already protected —
confidence="low"plus a byte-identical--adaptiveno-op, measured on both corpora. A different, real failure was found instead: single common words that route name-exact onto a large homonym cluster (many unrelated symbols sharing one literal name) could getconfidence="high"and a sharp cut driven by tie-break order rather than relevance, discarding 90%+ of an equally plausible pool (two reproduced examples: 231→6 and 148→10 on two different codebases). Query length/content does not separate the failing cases from the working ones — a homonym-pool-size statistic the code already computed but did not gate on does.Where this stands today. That finding did not sit as a proposal —
--adaptiveonmainnow declines to narrow when a name-exact route's positive pool clearskAdaptiveHomonymPoolFloor(50) and the cut would land at or below twice the floor (seeCHANGELOG.md, "--adaptiveno longer narrows a large pool of identically-named symbols on tie-break order"). This PR is the investigation record behind that fix, kept as its own artefact with the full method, both measurement tables, the two alternative fixes it argues against, and a pre-registered measurement band for review.This is a first pass with a lot further to push. The doc's closing section asks two questions we could not settle from the paper alone: whether homonym-tie-break cliffs are a known failure mode of largest-gap cutting outside symbol retrieval, and whether a principled pool-size/shape statistic exists to distinguish "sharp because relevant" from "sharp because tied" — we picked our threshold from twelve empirical points on two codebases, which is thin evidence for a shipped default. Review of that threshold, and of the rejected alternatives, is welcome.
Our working premise across this line of investigations: algorithmic, deterministic checks applied while an AI writes are the practical way to keep code sound at the speed AI now writes it — no label, no vendor claim, a check that either fires or it doesn't.
🤖 Generated with Claude Code