docs(research): is --for's ranking-confidence calibrated enough to abstain on? A first measurement, and what we'd need to go further - #315
Draft
joyful-ii-V-I wants to merge 7 commits into
Draft
joyful-ii-V-I wants to merge 7 commits into
joyful-ii-V-I wants to merge 7 commits into
Conversation
…d pre-register what abstention would take ARB (arXiv:2607.24882) names abstention as the unsolved retrieval axis; we shipped the disclosure half (confidence=/margin_pct= on every --for root) and two pre-registered rounds in EVALS.md then scored that signal against ARB's own unanswerable-query splits, both recorded negatives. Both rounds asked "does the signal know when there is no answer". Neither asked the question a CLI that must answer actually faces: does it know when ITS answer is wrong. This branch measures that, offline, on the LocBench held-out split (92 instances whose snapshots are already on disk; 214 more disclosed as an un-scored floor). The signal is directionally right and weakly discriminating, and unusable at its only operating point: it separates function-grain correctness (+0.375 [+0.132, +0.513], AUROC 0.622) but not file-grain (+0.134 [-0.081, +0.247]) and not the bundle as a whole (+0.012); warning iff confidence="low" fires on 74 of 92 answers and costs 60 false warnings to catch 14 misses. Structurally, margin_pct= is 0 for every "low" row, so there is no threshold to tune — the next round needs a fact, not a cut-point, which is where EVALS' round 2 also landed. The document carries the tables, the four abstention options with their costs against the honesty contract, the downstream agent experiment design (gated on a fire-rate ceiling, with its power arithmetic stated), a pre-registration with self-reject rules, and the two places outside help would most change what we build. bench/locbench/calibrate_confidence.py re-runs it. It fetches nothing, verifies the frozen dataset hash, imports the split and hit definitions from run_locbench.py and the registered AUROC/sweep from score_abstention_calibration.py so a LocBench number and an ARB number cannot disagree about the rule, refuses a --cache-dir inside the asset tree, and buckets every unscored instance under a named reason. No src/ change, no gate change, nothing published on any claim surface. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…utput byte-identical The first commit's own --quality-delta reported seven new-symbol regressions, all in this script: scored_instances (complexity 34, verbosity 91, params 7), markdown (23/77) and main (19/74). New-symbol findings never gate, but the debt is the author's, so it is paid in the same lane rather than left for a reader to find. scored_instances is now five named steps — load_rows (hash-verified, refuses rather than fetches), eligibility (every skip reason decided in ONE place, which is what makes the skip table a disclosure rather than a residual), instance_index, universe, grade and measure_instance — and its five knobs ride in a RunConfig instead of five positionals, which is how a cache dir and an asset dir get swapped at one call site and nowhere else. markdown splits into three section builders over a shared GRAINS list, so the markdown tables and the TSV lines emit the same four grains in the same order from one list rather than two. main hands its report to print_metric_lines and keeps argument handling. Proved neutral, not assumed: the full 92-instance held-out run reproduces the generated tables and every TSV metric line byte-identically, and an 8-instance run matches the pre-refactor per-instance rows exactly (wall clock excluded). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…helpers --quality-delta gated the previous commit with two duplication findings, both real and both mine: RunConfig::__init__ cloned bench/recalleval/run_recalleval.py's Label::__init__ (a hand-written __init__ whose whole body assigns its arguments), and RunConfig::run cloned bench/slice/run_slicerecall.py's run_ripwire. Introducing a clone while cleaning up complexity is the exact shape the check exists to catch, and it caught it on the commit that made it. RunConfig is now a collections.namedtuple — no __init__ body to clone, and the params finding on it goes with it. The invocation helper returns (stdout, returncode) instead of a CompletedProcess, which is both a different body and a better contract: every caller has to look at the return code to reach the output, and a run whose rc nobody read is how an empty answer becomes a measured zero. The full 92-instance held-out run reproduces the generated tables and every per-instance row byte-identically across all three revisions of this script. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…not a third copy of it The previous fix traded one clone for another: the replacement (stdout, rc) helper was a 46-token clone of bench/arb/run_arb.py's run_bin, which --quality-delta then gated. Writing a third private "how we run the binary" was the wrong move twice, and reuse was available both times — this script already imports its statistics from run_arb.py's sibling. It now imports run_bin itself. The coupling is the right one rather than a concession: a LocBench row and an ARB row are produced by literally the same call, so a LocBench number cannot drift from an ARB number through a private copy of the invocation. run_bin's 600 s ceiling becomes this harness's ceiling, and a TimeoutExpired is caught and bucketed as a NAMED skip reason (timeout, reported in the skip table) instead of ending a 92-instance run with a traceback. Re-run over the full held-out set: every table and every per-instance row byte-identical, timeout count 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ion in the measurement section Two facts the harness gained while its own quality findings were being paid off: the skip table now carries a named timeout bucket (0 on this run), and the binary is invoked through the ARB adapter's run_bin, which is worth stating where the commensurability claim is made rather than only in the script. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: trueComment |
Eight research lanes each added their own wording for the same `docs/README.md` table line, which is eight conflicts on one line the moment two of them land. This is the shared text, byte-identical everywhere, so the same addition on two branches merges clean. The entry count above the table is corrected with it: the table has held twenty rows for some time while the sentence still said sixteen. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…s, both FAIL Adds a dated addendum to the abstention note. served_syms, pre-registered at lane/served-syms-prereg (860b4df), scored AUROC 0.278 [0.177, 0.385] on func_hit at lane/served-syms-result (50c554e): at or below the refutation rung, no threshold meets the band, FAIL. The one pre-committed margin_bp re-score at lane/margin-rescore (9ffd362) scored 0.604 [0.484, 0.719]: FAIL, lane/for-margin-resolution closed. The reversed direction is not reported as a result; it would need its own pre-registration. No src/ change. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Agent Retrieval Bench (arXiv:2607.24882) reports that across four retrieval families no retriever tells its caller when its own ranking should not be trusted, and names abstention as the unsolved axis. Its
docs/LINEAGE.mdrow is at L89 — we already shipped the disclosure half: every--forroot carriesconfidence=/margin_pct=, derived from the same relevance-cliff statistic--adaptivecuts on, so the two can never disagree. Full write-up:docs/research/confidence-and-abstention.md.Disclosure is a first step, not a solution. This branch is the follow-up: is the signal we ship actually calibrated, and what would it take to go further than warning? It changes no code in
src/and adds no gate. (Base note: measured againstmainat755f9026, before 0.6.2; the diff still merges cleanly onto today'smain,15a20855, tag 0.6.2.)Headline finding. Two pre-registered
docs/EVALS.mdrounds already scored this signal against ARB's unanswerable-query splits and both closed as recorded negatives — they asked does the signal know there is no answer. A CLI that must answer faces a different question: does it know its own answer is wrong? Measured on a held-out LocBench slice (92 instances, repository-disjoint split), the signal is directionally right and weakly discriminating, and unusable at its only operating point: it separates function-grain correctness (+0.375 [+0.132, +0.513], AUROC 0.622) but not file-grain (+0.134 [−0.081, +0.247]) and not the bundle as a whole (+0.012). Warning iffconfidence="low"fires on 74 of 92 answers and costs 60 false warnings to catch 14 misses. Structurally,margin_pct=is 0 for every"low"row, so there is no threshold to tune — only a fact to add, which reachesdocs/EVALS.mdround 2's conclusion from the other direction.What else is in the document. The four ways a CLI could spell "abstain," each with its cost against this repository's honesty contract; a downstream agent experiment we cannot run here, designed in full with the power arithmetic and the results that would make us not ship; and a pre-registration with bands, a directional-refutation rung, and explicit self-reject rules, including a fire-rate ceiling today's signal fails by 3×.
This is a first pass with a lot further to push, and two places we would like outside help: calibration methodology —
--forreturns a set, "correct" has at least four grains and they disagree, so which grain should a single confidence attribute be a claim about, and is a two-valued label calibratable at all; and the abstention criterion — we pre-registered a 0.20 false-warn ceiling and can defend the direction but not the value, and we do not know of any measured evidence that agent behavior responds to a retrieval-confidence signal at all.Our working premise across this line of investigations: algorithmic, deterministic checks applied while an AI writes are the practical way to keep code sound at the speed AI now writes it — no label, no vendor claim, a check that either fires or it doesn't.
🤖 Generated with Claude Code