Skip to content

docs(research): is --for's ranking-confidence calibrated enough to abstain on? A first measurement, and what we'd need to go further - #315

Draft
joyful-ii-V-I wants to merge 7 commits into
mainfrom
lane/research-abstention
Draft

joyful-ii-V-I wants to merge 7 commits into
mainfrom
lane/research-abstention

Conversation

@joyful-ii-V-I

Copy link
Copy Markdown
Collaborator

Agent Retrieval Bench (arXiv:2607.24882) reports that across four retrieval families no retriever tells its caller when its own ranking should not be trusted, and names abstention as the unsolved axis. Its docs/LINEAGE.md row is at L89 — we already shipped the disclosure half: every --for root carries confidence=/margin_pct=, derived from the same relevance-cliff statistic --adaptive cuts on, so the two can never disagree. Full write-up: docs/research/confidence-and-abstention.md.

Disclosure is a first step, not a solution. This branch is the follow-up: is the signal we ship actually calibrated, and what would it take to go further than warning? It changes no code in src/ and adds no gate. (Base note: measured against main at 755f9026, before 0.6.2; the diff still merges cleanly onto today's main, 15a20855, tag 0.6.2.)

Headline finding. Two pre-registered docs/EVALS.md rounds already scored this signal against ARB's unanswerable-query splits and both closed as recorded negatives — they asked does the signal know there is no answer. A CLI that must answer faces a different question: does it know its own answer is wrong? Measured on a held-out LocBench slice (92 instances, repository-disjoint split), the signal is directionally right and weakly discriminating, and unusable at its only operating point: it separates function-grain correctness (+0.375 [+0.132, +0.513], AUROC 0.622) but not file-grain (+0.134 [−0.081, +0.247]) and not the bundle as a whole (+0.012). Warning iff confidence="low" fires on 74 of 92 answers and costs 60 false warnings to catch 14 misses. Structurally, margin_pct= is 0 for every "low" row, so there is no threshold to tune — only a fact to add, which reaches docs/EVALS.md round 2's conclusion from the other direction.

What else is in the document. The four ways a CLI could spell "abstain," each with its cost against this repository's honesty contract; a downstream agent experiment we cannot run here, designed in full with the power arithmetic and the results that would make us not ship; and a pre-registration with bands, a directional-refutation rung, and explicit self-reject rules, including a fire-rate ceiling today's signal fails by 3×.

This is a first pass with a lot further to push, and two places we would like outside help: calibration methodology — --for returns a set, "correct" has at least four grains and they disagree, so which grain should a single confidence attribute be a claim about, and is a two-valued label calibratable at all; and the abstention criterion — we pre-registered a 0.20 false-warn ceiling and can defend the direction but not the value, and we do not know of any measured evidence that agent behavior responds to a retrieval-confidence signal at all.

Our working premise across this line of investigations: algorithmic, deterministic checks applied while an AI writes are the practical way to keep code sound at the speed AI now writes it — no label, no vendor claim, a check that either fires or it doesn't.

🤖 Generated with Claude Code

barefootski and others added 5 commits September 20, 2026 14:43
…d pre-register what abstention would take

ARB (arXiv:2607.24882) names abstention as the unsolved retrieval axis; we shipped the
disclosure half (confidence=/margin_pct= on every --for root) and two pre-registered rounds
in EVALS.md then scored that signal against ARB's own unanswerable-query splits, both
recorded negatives. Both rounds asked "does the signal know when there is no answer". Neither
asked the question a CLI that must answer actually faces: does it know when ITS answer is
wrong.

This branch measures that, offline, on the LocBench held-out split (92 instances whose
snapshots are already on disk; 214 more disclosed as an un-scored floor). The signal is
directionally right and weakly discriminating, and unusable at its only operating point:
it separates function-grain correctness (+0.375 [+0.132, +0.513], AUROC 0.622) but not
file-grain (+0.134 [-0.081, +0.247]) and not the bundle as a whole (+0.012); warning iff
confidence="low" fires on 74 of 92 answers and costs 60 false warnings to catch 14 misses.
Structurally, margin_pct= is 0 for every "low" row, so there is no threshold to tune — the
next round needs a fact, not a cut-point, which is where EVALS' round 2 also landed.

The document carries the tables, the four abstention options with their costs against the
honesty contract, the downstream agent experiment design (gated on a fire-rate ceiling,
with its power arithmetic stated), a pre-registration with self-reject rules, and the two
places outside help would most change what we build.

bench/locbench/calibrate_confidence.py re-runs it. It fetches nothing, verifies the frozen
dataset hash, imports the split and hit definitions from run_locbench.py and the registered
AUROC/sweep from score_abstention_calibration.py so a LocBench number and an ARB number
cannot disagree about the rule, refuses a --cache-dir inside the asset tree, and buckets
every unscored instance under a named reason.

No src/ change, no gate change, nothing published on any claim surface.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…utput byte-identical

The first commit's own --quality-delta reported seven new-symbol regressions, all in this
script: scored_instances (complexity 34, verbosity 91, params 7), markdown (23/77) and main
(19/74). New-symbol findings never gate, but the debt is the author's, so it is paid in the
same lane rather than left for a reader to find.

scored_instances is now five named steps — load_rows (hash-verified, refuses rather than
fetches), eligibility (every skip reason decided in ONE place, which is what makes the skip
table a disclosure rather than a residual), instance_index, universe, grade and
measure_instance — and its five knobs ride in a RunConfig instead of five positionals, which
is how a cache dir and an asset dir get swapped at one call site and nowhere else. markdown
splits into three section builders over a shared GRAINS list, so the markdown tables and the
TSV lines emit the same four grains in the same order from one list rather than two. main
hands its report to print_metric_lines and keeps argument handling.

Proved neutral, not assumed: the full 92-instance held-out run reproduces the generated
tables and every TSV metric line byte-identically, and an 8-instance run matches the
pre-refactor per-instance rows exactly (wall clock excluded).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…helpers

--quality-delta gated the previous commit with two duplication findings, both real and both
mine: RunConfig::__init__ cloned bench/recalleval/run_recalleval.py's Label::__init__ (a
hand-written __init__ whose whole body assigns its arguments), and RunConfig::run cloned
bench/slice/run_slicerecall.py's run_ripwire. Introducing a clone while cleaning up
complexity is the exact shape the check exists to catch, and it caught it on the commit that
made it.

RunConfig is now a collections.namedtuple — no __init__ body to clone, and the params
finding on it goes with it. The invocation helper returns (stdout, returncode) instead of a
CompletedProcess, which is both a different body and a better contract: every caller has to
look at the return code to reach the output, and a run whose rc nobody read is how an empty
answer becomes a measured zero.

The full 92-instance held-out run reproduces the generated tables and every per-instance row
byte-identically across all three revisions of this script.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…not a third copy of it

The previous fix traded one clone for another: the replacement (stdout, rc) helper was a
46-token clone of bench/arb/run_arb.py's run_bin, which --quality-delta then gated. Writing
a third private "how we run the binary" was the wrong move twice, and reuse was available
both times — this script already imports its statistics from run_arb.py's sibling.

It now imports run_bin itself. The coupling is the right one rather than a concession: a
LocBench row and an ARB row are produced by literally the same call, so a LocBench number
cannot drift from an ARB number through a private copy of the invocation. run_bin's 600 s
ceiling becomes this harness's ceiling, and a TimeoutExpired is caught and bucketed as a
NAMED skip reason (timeout, reported in the skip table) instead of ending a 92-instance run
with a traceback.

Re-run over the full held-out set: every table and every per-instance row byte-identical,
timeout count 0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ion in the measurement section

Two facts the harness gained while its own quality findings were being paid off: the skip
table now carries a named timeout bucket (0 on this run), and the binary is invoked through
the ARB adapter's run_bin, which is worth stating where the commensurability claim is made
rather than only in the script.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Sep 21, 2026

Copy link
Copy Markdown

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true

Comment @coderabbitai help to get the list of available commands.

barefootski and others added 2 commits September 21, 2026 11:52
Eight research lanes each added their own wording for the same `docs/README.md`
table line, which is eight conflicts on one line the moment two of them land.
This is the shared text, byte-identical everywhere, so the same addition on two
branches merges clean. The entry count above the table is corrected with it:
the table has held twenty rows for some time while the sentence still said
sixteen.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…s, both FAIL

Adds a dated addendum to the abstention note. served_syms, pre-registered at
lane/served-syms-prereg (860b4df), scored AUROC 0.278 [0.177, 0.385] on
func_hit at lane/served-syms-result (50c554e): at or below the refutation
rung, no threshold meets the band, FAIL. The one pre-committed margin_bp
re-score at lane/margin-rescore (9ffd362) scored 0.604 [0.484, 0.719]: FAIL,
lane/for-margin-resolution closed. The reversed direction is not reported as a
result; it would need its own pre-registration. No src/ change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants