From 3d8c16fd30f85a7d0d1257c155c8821f80879861 Mon Sep 17 00:00:00 2001 From: Phil Leggetter Date: Thu, 1 Oct 2026 15:35:06 +0100 Subject: [PATCH 1/2] Record what the first regression run measured, and fix what it made stale MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The run that #86 said to take before leaving the schedule enabled: eighteen cells green at two attempts, fifty minutes of wall clock, alert job correctly skipped. Whole-run time including queueing — the cells run in parallel, so there is no per-cell figure in it, and cost is still unmeasured because nothing records dollars for an agent that reports tokens. That run is also the first evidence for the assumption the design rests on. "Every agent passing is the expected state" was unverified until today, and the only record that existed contradicted it: one failing row from 10 August against an experiment that no longer exists. Three statements the change made wrong. The workflow is active, not disabled — it was off for six hours this morning so the monthly could not fire mid-change, and saying "it is disabled" would send the next reader to enable something already enabled. A dispatched run defaults to two attempts now, not one. And the Status section said the weak model is two worse with skills than without; on 28 September it is four better, and across four September runs the delta reads -1, +1, +4, +4. That last paragraph has now been rewritten three times, each time towards less confidence, so it now carries the warning rather than a number: do not report the delta as a single figure, do not call a sign replicated until it has replicated, and read it per row because the disagreements move. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01MQzUoMAwEBJWpEGVvVzSjK --- .github/workflows/eval-refresh.yml | 12 ++++--- AGENTS.md | 51 ++++++++++++++++++------------ 2 files changed, 38 insertions(+), 25 deletions(-) diff --git a/.github/workflows/eval-refresh.yml b/.github/workflows/eval-refresh.yml index b40d931..d5dd73b 100644 --- a/.github/workflows/eval-refresh.yml +++ b/.github/workflows/eval-refresh.yml @@ -17,11 +17,13 @@ name: Refresh eval results # instead of going red where nobody looks. The 21 September failure sat # unnoticed until somebody asked about it four days later. # -# Cost and duration are not yet measured. The $5-and-ten-minutes figure that -# used to sit here came from `.plans/delivery-plan.md` at planning time, before -# any regression run existed; measured wall clock on the benchmark is about -# 3.4 minutes per cell serialised, which would put eighteen cells nearer an -# hour. Replace this sentence with a measurement after the first run. +# Measured on the first run, 1 October 2026 (run 36869979661): eighteen cells +# green at two attempts each, fifty minutes of wall clock. That is whole-run +# time including runner queueing, not per-cell — the cells run in parallel, so +# no per-cell figure can be read off it. The $5-and-ten-minutes estimate that +# used to sit here was a planning figure from before any regression run +# existed; cost is still unmeasured, because nothing in the harness records it +# for an agent that reports tokens rather than dollars. # # **Benchmark runs are dispatched against a bucket of work, not a calendar.** # Something changed what we measure, or something shipped to the product, and diff --git a/AGENTS.md b/AGENTS.md index 51380bc..107b7db 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -37,14 +37,19 @@ healthy benchmark rather than a flat one. The frontier agents pass nearly everything; the weak model is where most failures live, which is the floor working as intended. -The skills axis is the interesting result and it is not uniform. Claude gains -one scenario from skills, GPT-5.6 nets zero, and the weak model is **two worse -with skills than without** — on the most recent published run it loses four -scenarios and gains two. That direction is a finding about our documentation -rather than about the model, and there is a known mechanism: a skill that lists -example values is read as an exhaustive list, which once led a weak model to -conclude a supported provider was unsupported. Do not report the skills delta as -a single number; it has a different sign at different capability levels. +The skills axis has not settled, and this paragraph has been rewritten three times +in the direction of less confidence. It once said the weak model was **two worse** +with skills than without; on 28 September that model reads **four better** (18/19 +against 14/19), and across the four September runs the delta has read -1, +1, +4 +and +4. #2 was closed on 22 September with its premise withdrawn: the -3 that +started it is from the credit-outage day and no run since reproduced it. + +What survives is the warning rather than the number. **Do not report the skills +delta as a single figure**, do not report a sign as replicated until it has +replicated, and read it per row — the disagreements move between runs, so a total +hides which cells produced it. The one mechanism that is documented rather than +inferred stands: a skill that lists example values is read as an exhaustive list, +which once led a weak model to conclude a supported provider was unsupported. **The weak-model figure is −2 and was written here as −3 for eleven days.** Both −3 readings are from 13 August and no run since has reproduced them: eight of @@ -98,8 +103,7 @@ others would have caught. `eval-refresh` runs the **regression suite** weekly (Monday 06:00 UTC, every experiment) and nothing else on a schedule. Benchmark runs are dispatched -against a bucket of work — see Runs below. The workflow is disabled, so a -dispatch needs `gh workflow enable eval-refresh.yml` first. Check `gh secret list` against the workflow env +against a bucket of work — see Runs below. Check `gh secret list` against the workflow env rather than trusting any list written here: `OUTPOST_API_KEY` was documented as a secret before it existed, and the first full matrix run scored `outpost-001` as six agent failures because of it. @@ -173,11 +177,15 @@ state, so a failing check **fails the job** here — the opposite of the benchma where a failure is a score — and the run opens or updates an issue labelled `regression-alert`, because a red run in a tab nobody has open is not a notification. -Its cost and duration are **not yet measured**. The "$5 and ten minutes" figure that -circulated came from the delivery plan at planning time, before any regression run -existed, and the benchmark's measured 3.4 minutes per cell serialised would put -eighteen cells nearer an hour. Measure it after the first run rather than quoting the -estimate again. +**Measured on the first run**, 1 October 2026 ([run 36869979661](https://github.com/hookdeck/evals/actions/runs/36869979661)): eighteen cells +green at two attempts each, **fifty minutes** of wall clock. Whole-run time including +runner queueing — the cells run in parallel, so there is no per-cell figure in it. The +"$5 and ten minutes" that circulated was a planning estimate from before any regression +run existed. Cost is still unmeasured. + +That run is also the first evidence that every agent passing is the expected state here. +It was an assumption until then, and the only record that existed contradicted it: one +failing row from 10 August against an experiment that no longer exists. **A benchmark run is dispatched against a bucket of work, never a date.** Two buckets, and a run belongs to one of them: @@ -194,11 +202,13 @@ consuming, at about $185 a month plus $90-110 a matrix. Five weeklies had produc harness defect and a lot of variance data about a delta we already know we cannot measure precisely enough (#2). -Re-enable the workflow before dispatching: it is disabled, which blocks -`workflow_dispatch` as well as the cron. +The workflow was disabled between 1 October 07:30 and 13:30 UTC so the monthly matrix +could not fire while the schedule was being changed. It is active again. Disabling is +the way to stop a cron without a commit, and it blocks `workflow_dispatch` too: ```bash -gh workflow enable eval-refresh.yml +gh workflow disable eval-refresh.yml # stops the cron and dispatch +gh workflow enable eval-refresh.yml # both again ``` ## What to work on next @@ -252,8 +262,9 @@ Two rules that are the whole point of the file: - **Record the loops that failed.** A change that did not work is more informative than one that did, and omitting them makes the rest less believable. -Re-runs for a loop need at least three attempts. A dispatched run defaults to one and -cannot separate a fix from variance. +Re-runs for a loop need at least three attempts. A dispatched run defaults to two, +which is enough to stop one unlucky cell deciding a verdict and not enough to separate +a fix from variance — ask for three. ## Releases From 87ade407c36fb0a97c7abde4f9630aa134dd1ab3 Mon Sep 17 00:00:00 2001 From: Phil Leggetter Date: Mon, 5 Oct 2026 17:28:51 +0100 Subject: [PATCH 2/2] Correct it: seventeen of eighteen, not eighteen The 1 October dispatch was reported here as green on all eighteen cells. It was not. regression-filtering-001-regex-capability failed for claude-code-sonnet-5 across both attempts, and the run still concluded success because the step that fails a cell on a red check is gated on `github.event_name == 'schedule'`. So a dispatched run cannot verify what a scheduled run enforces, which is what that dispatch existed to do, and a green conclusion is not evidence about cells. I took the run's conclusion as the cell result and told the same thing to everyone downstream. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01MQzUoMAwEBJWpEGVvVzSjK --- AGENTS.md | 24 +++++++++++++++--------- 1 file changed, 15 insertions(+), 9 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 107b7db..13bb7c2 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -177,15 +177,21 @@ state, so a failing check **fails the job** here — the opposite of the benchma where a failure is a score — and the run opens or updates an issue labelled `regression-alert`, because a red run in a tab nobody has open is not a notification. -**Measured on the first run**, 1 October 2026 ([run 36869979661](https://github.com/hookdeck/evals/actions/runs/36869979661)): eighteen cells -green at two attempts each, **fifty minutes** of wall clock. Whole-run time including -runner queueing — the cells run in parallel, so there is no per-cell figure in it. The -"$5 and ten minutes" that circulated was a planning estimate from before any regression -run existed. Cost is still unmeasured. - -That run is also the first evidence that every agent passing is the expected state here. -It was an assumption until then, and the only record that existed contradicted it: one -failing row from 10 August against an experiment that no longer exists. +**Measured on the first run**, 1 October 2026 ([run 36869979661](https://github.com/hookdeck/evals/actions/runs/36869979661)): eighteen cells, +**seventeen passed**, fifty minutes of wall clock. Whole-run time including runner +queueing — the cells run in parallel, so there is no per-cell figure in it. The "$5 and +ten minutes" that circulated was a planning estimate from before any regression run +existed. Cost is still unmeasured. + +**That run was reported as green and was not.** `regression-filtering-001-regex-capability` +failed for `claude-code-sonnet-5` across both attempts, and the run still concluded +success, because the step that fails a cell on a red check is gated on +`github.event_name == 'schedule'`. A dispatched run therefore cannot verify what a +scheduled run enforces — which is precisely what that dispatch was for. Read the result +files, not the run's conclusion. + +So **"every agent passing is the expected state" remains an assumption**, and the +evidence is against it for one scenario and one arm. See #92. **A benchmark run is dispatched against a bucket of work, never a date.** Two buckets, and a run belongs to one of them: