Repository navigation
gc-ratchet has been RED on main since 2026-08-18 (19 days); every labelled PR inherits it, and the failure has grown from 19 to 43 cells #9829
Description
Activity
Bisected: the first accrual is
e2e6f4a1c perf(object): shrink common objects to 40 bytes (#8313)— 19 of 19 cells, byte-exactScoped to the first accrual as asked.
e2e6f4a1cis the immediate successor
of the last green commitdc4bcf287, so one measurement either side settles
it without a search.Method — no baseline reasoning involved. I measured
gc_ratchet.py measure
atdc4bcf287and ate2e6f4a1con the same host and diffed the raw counters,
rather than checking either against a pinned baseline. The gated counters are
documented as "semantic and transfer across machine classes", and the
measurement confirms it: spread across repeats is 0 % on every gated cell,
i.e. bit-identical.Result
cells moved by e2e6f4a1cacross all 14 probes69 of CI's 19 original regression cells, moved by it 19 / 19 of those, landing on CI's exact reported "current" value 19 / 19 cells e2e6f4a1cdid not move0 The local
dc4bcf287measurement also reproduces the pinned baseline value
exactly for 13 of the 19; the other 6 differ by 0.02–0.15 % (e.g. 13,201,144 vs
13,181,528), which is inside their documented bands — and for all 6 the post
value still matches CI exactly.Nineteen byte-exact matches on a different machine class from the CI runner is
about as firm as attribution gets.Representative rows (CI current == local
e2e6f4a1c, exactly):cell baseline after e2e6f4a1c12_large_live_set.freed_bytes76,315,456 63,732,664 −16.5 % 12_large_live_set.promoted_bytes25,395,472 21,201,208 −16.5 % 13_large_eden_survivors.freed_bytes97,086,288 86,649,800 −10.8 % 04_dead_after_deep_stack.freed_bytes83,880,432 76,532,640 −8.8 % 03_cross_gen_writes.freed_bytes11,304,368 9,240,008 −18.3 % What the direction means
Overwhelmingly a win that was never re-pinned, and the shape is exactly
what "objects got smaller" predicts:freed_bytesdown 9–18 % in nearly every probe — less garbage produced,
because each object is smaller.copied_bytes,promoted_bytes,heap_total_bytes,heap_used_bytes
mostly down (heap_total−5 % to −14 %).copied_objectsup 7–17 % in several probes. That is not retention: with
smaller objects more of them fit in the nursery before it fills, so more
survive per cycle. A schedule shift, which is why probes elsewhere in the
43 look like they "stopped collecting".
Two cells I am NOT calling a win, per the standing constraint
12_large_live_set.heap_used_bytes1,104,848 → 1,681,712 (+52.2 %) —
live heap up by half on the large-live-set probe, while everything else on
that probe went down. Smaller objects reducing live bytes is expected;
increasing them is not, and I do not have an explanation.14_grow_then_churn.copied_bytes+11.1 % andcopied_objects+10.8 % — the
only probe where copying rose in bytes as well as count.
Both need a reason from someone who knows #8313's layout change before any
re-pin. Everything else on this list I would be comfortable calling correct.Scope
The second accrual is untouched — the 24 cells that appeared between
2026-08-18 and today are still unattributed, and I did not chase them in this
pass. Whether they are one event or several is the next question.Re-pinning still needs per-cell reasons, and this comment supplies them for 17
of the 19; the two above are the ones to resolve first.Filed alongside: #9830, for the automation that would have surfaced this on
2026-08-18 instead of nineteen days later.Both unexplained cells resolved. Neither is a regression — and
12_large_live_setis a net 9.9 MB memory improvement, not a 52 % increase12_large_live_set.heap_used_bytes+52.2 %The decisive fact is what that counter is. Reading the probe:
heap_used_bytesisprocess.memoryUsage().heapUsedsampled after the
release phase and an explicitgc()— 63 of every 64LNodes dropped, ~10,938
kept. It is the probe's stranded-in-old-gen detector, not its memory figure.
The probe's memory figure isheap_total_bytes, and that went down 12.7 %.I isolated the release phase and ran it at both commits on one host:
dc4bcf287e2e6f4a1clive heap, all 700k nodes live 42,225,832 37,127,176 −12.1 % retained after dropping 63/64 + gc()1,646,176 2,221,720 +35.0 % bytes per kept node 150.5 203.1 +35.0 % So objects did NOT get bigger — they got 12 % smaller, exactly as #8313
intends. The commit removes the per-objectkeys_arraymirror, taking
ObjectHeader32 → 16 bytes and a two-slot object 56 → 48 → 40. The live set
shrinks accordingly; what rose is only what survives the release.Putting both of that probe's retention counters together:
heap_total_bytes 82,837,504 -> 72,351,744 -10,485,760 heap_used_bytes 1,104,848 -> 1,681,712 +576,864 net -9,908,896 bytes#8313 makes this probe use ~9.9 MB less memory. The cell that rose is a
~1 MB tail on a ~80 MB workload, and it rose because the objects are smaller:
more of them pack per page, so a cohort dropped uniformly (every 64th node
kept) leaves proportionally more partially-occupied pages that cannot be
reclaimed whole. That last clause is a hypothesis about the mechanism — the
measurement above establishes the facts (smaller objects, smaller live heap,
larger post-release residue) without depending on it.So this is neither of the two options in the brief. Not "accounting moved" and
not "objects got bigger": objects got smaller, and the release leaves more
behind. It is a real, small, and strictly-dominated effect — worth pinning
with that reason recorded, not worth calling a regression.I would still flag one thing for whoever owns #8313: the residue scales with
the drop pattern, not with the live set. A workload that frees most of a large
cohort and keeps a sparse remainder is the shape that pays it. cc does that, so
if anyone wants to chase footprint further, "post-release page occupancy after a
sparse survivor cohort" is a real lead — but it is a lead, not this commit's
defect, and it predates any question of re-pinning.14_grow_then_churncopied up in both terms — confirmed schedule, not sizeAsked to confirm rather than assume, and the average settles it in one line:
copied_objects 307 -> 340 (+10.75 %) copied_bytes 1,222,176 -> 1,357,872 (+11.10 %) average bytes per copied object 3,981.0 -> 3,993.7 (+0.32 %)The copied population's average size is unchanged (+0.32 %). These are
~4 KB objects — arrays, not the 40-byte objects #8313 touched — so their size
could not have moved, and 33 more of them were copied. That is the same
schedule shift that explains the other 17 cells: smaller objects mean more fit
in the nursery before it fills, so more survive per cycle. Confirmed, not
assumed.Verdict
All 19 of 19 first-accrual cells now have a stated reason:
- 17 — direct consequences of smaller objects (
freed_bytes,copied_bytes,
promoted_bytes,heap_total_bytesdown) or of the schedule shift that
follows (copied_objectsup). 12_large_live_set.heap_used_bytes— post-release residue up ~0.58 MB while
the same probe's total falls 10.5 MB; net −9.9 MB.14_grow_then_churn— schedule, confirmed by invariant average object size.
Nothing here needs filing as a regression, so the re-pin does not need to
carve anything out. A re-pin PR for these 19 cells can now be written with a
per-cell reason, as the gate's own docs require. I have not touched the 24
second-accrual cells and they must not be swept into the same re-pin.Method note, same as the bisect: every comparison is two commits measured
against each other on one host, never against a pinned baseline.- 17 — direct consequences of smaller objects (
Second accrual mapped: it is not one event — it is at least eight, and the cell set churns
Mapped without a single build, from evidence that already existed: each of the
59 failing scheduled runs since 08-18 carries its own regression rows, so
counting cells per run brackets every accrual to one six-hour window.The timeline
date head cells 08-18 7441e1f7319 first red — #8313 (bisected above) 08-20 526e0b50224 +5 08-24 f4a7559d632 +8 08-24 c76b4393028 −4 08-26 ead63646432 +4 08-28 6d10e8a1d38 +6 08-29 3b9e786d639 +1 09-01 b92c101b643 +4 09-01 3a9c4580142 −1 09-04 ef79a021148 +6 09-05 d36a1af0c43 −5 The count goes down as often as it needs to (32→28, 43→42, 48→43), so this
is not accumulation — it is several independent changes moving the same
counters, some partially offsetting.My earlier arithmetic was wrong, and here is the correction
I said "19 attributed, 24 remaining". That assumed the original 19 persisted.
They did not:- of the original 19, only 14 are still red today;
- 5 have since cleared —
02_survivor_promotion.copied_objects,
09_try_catch_roots.freed_bytes,11_collect_at_depth.freed_bytes,
14_grow_then_churn.copied_bytes, and
12_large_live_set.heap_used_bytes; - 29 accrued after 08-18 and are still red.
So today's 43 = 14 (from #8313) + 29 (later), not 19 + 24.
And that includes the cell I was asked to prioritise.
12_large_live_set.heap_used_bytesis not in today's failing set — it went
red at #8313 and has since come back into band. My analysis of why it moved
stands and produced a real fragmentation lead, but it was never blocking the
gate, and I should have checked that it was still failing before spending two
builds on it. Process lesson: confirm a cell is in the current failing set
before investigating it.The cell that now deserves that treatment instead is
10_store_receiver_across_alloc.heap_used_bytes, baseline 220,384 →
464,072, +110.57 % — today's largest direction-contradicting retention cell,
and it is in the accrued-29.The cell names identify the culprits — no blind bisect needed
For the three smallest windows the match is close to unambiguous:
window cells that appeared commits standout candidate 19→24 (8 commits) all five are 06_string_retention.*f1d231fa5..526e0b5025f9aa1f40 perf(codegen): retain string accumulators in concat chains (#8417)24→32 (4 commits) 04_dead_after_deep_stack,05_closure_capture,07_array_grow_evacuate,14_grow_then_churn— copy/promote across four probesc2da03439..f4a7559d606e1ab349 merge: land #8652/#8651/#8647/#8646/#8650 (#8657)(five PRs in one commit) or2382a9f15 perf(codegen): delete six instructions from the generic property-read hit path (#8665)39→43 (10 commits) 13_large_eden_survivors.{copied_bytes,copied_objects,heap_total_bytes,promoted_objects}+heap_usedon 09/112d6a05038..b92c101b63d1adbc4c gc: give the copied minor the liveness probes it never had (#9359)— a GC change to the copied minor, and the cells are copied/promoted on eden survivorsFull appeared/cleared lists for all eight transitions are in this session's
artefacts and can be pasted in if wanted.Worth naming: most of these windows are perf work — much of it this
campaign's. That is expected and not an accusation: perf changes move GC
counters, and this gate exists to make you justify them. It was dark, so nobody
was asked to.The judgement I am bringing back rather than making
- Confirming all eight by bisect would be ~30–50 builds. Confirming them from
the shortlists above is ~8 builds (one per window, testing the named
candidate). Both are real work; neither is a re-pin. - A re-pin covering all 43 needs a reason per cell, and 29 of them currently
have only a window and a candidate, not an attribution.
So: do you want the ~8 confirmations, or is the map itself — every cell
bracketed to a six-hour window with a named candidate — sufficient justification
for a re-pin that says "the gate was dark for nineteen days; here is what moved,
when, and by what"? I do not think I should choose that alone, and I have not
started either.Method unchanged: two commits measured against each other on one host, counter
definitions read before interpreting direction, structure printed before
trusting a diff.Classification done, and it stopped short of the pin — for a reason worth the delay
Split by what a wrong answer would cost, as instructed. The sweep for cells
whose direction contradicts their window's mechanism found four retention
cells rising, and confirming them needed zero builds — the gate's own
measurement artifact from main's scheduled run (33989581881,d36a1af0c)
carriesminor_cyclesandstep_cyclesper probe, which settles all four.The blocker:
10_store_receiver_across_allocis inert — filed as #9832metric baseline today minor_cycles1 0 copied_objects/copied_bytes/freed_bytes8,160 / 506,200 / 8,930,928 0 / 0 / 0 heap_used_bytes220,384 464,072 (+110.57 %) freed_bytes = 0rules out "a minor ran and found nothing" — no collection
ran at all. Every other probe still collects (minor_cycles1–22, ±1).That probe is the only cover for the stale-root-across-evacuation class
(#6970 / #9523), and its own header says dropping the allocating RHS "makes it
silently measure nothing". The +110.57 % is a consequence of nothing being
collected, not a retention regression — the same reading error that
12_large_live_set.heap_used_bytesinvited.So I did not write the pin. Pinning
minor_cycles = 0would permanently
bless an inert probe: a green gate that proves nothing, which is the exact
failure this suite's header warns about, and the exact thing you said would
destroy the gate's value. #9832 has the two possible orders and the decision I
should not make alone.The other three rising retention cells — explained, not regressions
13_large_eden_survivors.heap_total_bytes+3.51 % — the probe also shows
copied 0 → 377,569whilepromoted −28.7 %. That is a tenuring-policy
shift: objects that were promoted straight out are now copied within the
nursery first. Fewer promotions, slightly larger nursery. Explained by the
mechanism, dominated by improvements.07_array_grow_evacuate.heap_used_bytes+8.90 % /heap_total_bytes
+9.09 % withcopied_objects −21.86 %— same family, and itscopied_bytes
andfreed_bytescells already carry a documentedprobe_overridesexclusion
from GC Ratchet cannot gate 07_array_grow_evacuate: copied_bytes/freed_bytes are nondeterministic across traced runs #9790 for precisely this promotion-order/packing interaction.04_dead_after_deep_stack(+6720 % promoted_objects) and
05_closure_capture(0 → 3,390 promoted) — both are survivorship rising on
liveness probes, which is the shape that would hide a stale-root bug. But
both keep collecting (minor_cycles6→7 and 3→3), both showfreed_bytes
still falling (−12.6 %, −17.2 %), and the absolute magnitudes are tiny
(149 KB, 122 KB). Consistent with more objects surviving per cycle because
they are smaller — the same schedule shift as the other 17. I would not call
these regressions, and I would note them in the pin rather than measure
further.
Evidence grades, ready for whoever writes the pin
grade cells basis measured the original 14 still red, traced to #8313 two-commit bisect, 19/19 byte-exact (above) measured the 4 rising-retention cells minor_cycles/step_cyclesfrom CI's own artifactbracketed the remaining 25 six-hour window + named candidate per transition (map above) must not be pinned probe 10's 4 cells #9832 One thing the pin must say
The count went 43 → 48 → 43 in the last two days. This is not a ratchet
accumulating debt; it is several independent changes moving shared counters
with partial offsets. So the pin cannot be "the values as of today" without
saying that today is one sample of a moving set: it must record the SHA
(d36a1af0c), the run it came from (33989581881), and that the gate's failures
were unwatched from 2026-08-18 to 2026-09-06. Otherwise the next person reads it
as a considered baseline rather than the artefact of a dark gate.I am stopping here on this thread, per your steer — the retention-tail lead is
written up inHANDOFF_keystroke_retention_tail.mdand is where I would go
next.Probe repaired (#9833). The pin is now blocked on one more thing, and it is a tool bug (#9834)
#9833 fixes the inert probe:
minor_cycles0 → 9, sabotage-verified
(removing the allocating RHS returns the exact inert signature), and
heap_used_bytesreturns 464,072 → 244,648 against a baseline of 220,384 —
confirming that cell was never retention.The sweep that came with it is the bigger find: six of fourteen probes were
pinned atminor_cycles == 1, and one had already fallen to 0.01,02,
03,09,11are each one allocation win from going silently inert, while
this campaign lands allocation wins deliberately.gc_ratchet.pyrefuses to pin
minor_cycles < 1— it catches a probe that is already dead and blesses the
state that produces one. Detail in #9832.Why I have not committed a baseline
I measured 7 repeats on
main@d36a1af0c+ #9833 — 14 probes, no inert
probe, no correctness failure — andassemblerefused it:gc-ratchet error: 03_cross_gen_writes: wall_ms summary is inconsistent with its samples; 05_closure_capture: ... ; 11_collect_at_depth: ...The measurement is fine.
distribution()is not idempotent under its own
rounding: the storedsamplesare rounded to 6 dp, the recordedspread_pct
was computed from the raw ones, andvalidaterecomputes from the rounded ones.
Three probes differed by one unit in the last place. It is fatal rather than
deferrable, and it is nondeterministic — a different run hits a different
subset. Filed as #9834 with the one-line fix and the missing regression
test.So the pin needs, in order:
- fix(gc-ratchet): 10_store_receiver_across_alloc runs no minor collection — give it margin #9833 to land (a baseline cannot be pinned while a probe is inert — the
tool enforces this, so the 43-cell pin would have been rejected outright). - gc-ratchet: assemble/validate can reject a valid artifact — distribution() is not idempotent under its own rounding (blocks re-pinning) #9834 to land, or the pin will be refused at random.
- A measurement from a run that contains both. CI's own is the natural source;
my local GC counters have matched CI byte-for-byte throughout this
investigation, so either host will do for the gated cells.
The pin's content is ready
Grades, unchanged from the previous comment:
grade cells basis measured 14 original still red two-commit bisect, 19/19 byte-exact measured 4 rising-retention cells minor_cycles/step_cyclesfrom CI's artifactmeasured probe 10's 4 cells re-measured with the probe live (#9833) bracketed remaining 21 six-hour window + named candidate each And the sentence the pin must carry: the failing-cell count went 43 → 48 →
43 in the two days before it, so this is one sample of a moving set — pin
it with the SHA (d36a1af0c), the run (33989581881), and the fact that the
gate's failures were unwatched from 2026-08-18 to 2026-09-06.I am stopping here and moving to the retention tail
(HANDOFF_keystroke_retention_tail.md), per the standing steer. Happy to write
the pin the moment #9833 and #9834 are in.- fix(gc-ratchet): 10_store_receiver_across_alloc runs no minor collection — give it margin #9833 to land (a baseline cannot be pinned while a probe is inert — the
Diagnosed after
gc-ratchetandgc-root-dominancefailed identically on#9827, #9823, #9816 and #9807 — branches touching unrelated code. Unrelated
changes cannot produce identical failures, so the question was whether main
fails its own ratchet with no PR involved. It does.
1. Main fails its own ratchet
The workflow's main-line arm is a six-hourly schedule (
push: branches: [main]was removed in #7856 because it starved the queue). That arm is red:Run
33989581881(scheduled,main@d36a1af0c, no PR) fails with 43regression rows.
2. The PRs inherit it exactly — they contribute nothing
Extracting the regression rows from main's own scheduled run and from PR
#9823's run:
So the campaign's GC-adjacent work is not the cause, and re-pinning on any
of those PRs would be absorbing main's problem into a branch that did not
create it.
3. It started 2026-08-18 and has been red for 19 days
Every scheduled run since is red; the last green is
dc4bcf287:Of the last 100 scheduled runs (08-11 → 09-05): 15 success — all of them on
or before 08-17 — 64 failure, 16 cancelled.
4. It is NOT one regression: the failure has grown 19 -> 43 cells
The first red run (
7441e1f73) had 19 regression cells. Today's has 43.14 have been red the whole time; 29 were added since, including whole
probes that were clean on 08-18:
At least two accrual events, 19 days apart. That is the single most
important fact here, because it rules out a one-line re-pin.
Representative rows from today (main, scheduled):
10_store_receiver_across_alloc10_store_receiver_across_alloc10_store_receiver_across_alloc09_try_catch_roots13_large_eden_survivors12_large_live_setThe shape — byte counters moving together, and some probes going to zero
copied/promoted/freed — reads as probes whose collection schedule changed
(they stopped triggering a minor at all), not as a retention leak.
5. Candidate for the FIRST accrual, offered as a lead not a conclusion
Ten commits sit between the last green and the first red:
e2e6f4a1c— "shrink common objects to 40 bytes" — is the obviouscandidate: changing the common object size moves every byte-derived counter
and, by changing how fast the nursery fills, the collection schedule too, which
is exactly the failure shape. It is also plausibly a win that simply needs
the baseline re-pinned with that justification. I have not bisected it and I
am not asserting it. The second accrual is unattributed.
6. Why nobody saw it — the same class as #9774 and #7856
run-extended-tests, so almost every PR showsskippingand proves nothing.notifies nobody.
So main can drift past its own pinned baseline indefinitely, and the first PR
to ask for the gate inherits 19 days of accumulated red. A gate that has failed
64 of its last 100 runs while blocking nothing is dark in CLAUDE.md's sense —
a fourth way, alongside
continue-on-error, missing-from-required-contexts,cancelled and starved: failing but unwatched.
7. Do NOT re-pin to clear this
Re-pinning is how a real regression gets absorbed, and here it would absorb an
unknown number of them: 43 cells, at least two accrual events, 19 days and
~200 commits. Nobody can currently say which of those values are correct.
Note
423975f6f(Sep 5) touched the baseline file but changedprobe_overrides/rationalemetadata only — no one has re-pinned countervalues, so this red is not a hidden re-pin returning.
Suggested order of work
run on a single commit;
#8313first).demand: which counters moved, why the new values are correct, and what would
have failed under the old pin.
mainrun should openor update an issue automatically. That is the part that would have saved 19
days, and it is independent of whatever the counters turn out to mean.
Filed from the keystroke lane while unblocking #9807/#9823/#9828; happy to take
the bisection if it is wanted, but which counter values are correct is a
judgement I should not make alone.
https://claude.ai/code/session_014UZWia6L37DpA93VLtNK9m