Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
15 changes: 15 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -55,6 +55,21 @@ change measured timings are marked `bench:`.
index: skipped with a notice, because a gate that fails offline is a gate
people learn to skip.

- **`docs/LIMITATIONS.md` contradicted itself about the T4, and claimed a
measurement that never happened.** Line 147 said "Only 8.6 (A10G) and 7.5
(T4) have ever had a kernel measured on them here"; line 167, in the
section on results not porting, said "nothing has been measured on a T4".
The evidence agrees with the second: `docs/research-baseline.md` records
the tier-2 box as a `g5.xlarge` with an A10G, *chosen over the T4 by the
operator*, and `model-calibration.toml` names one device.

This is not a footnote. `README.md` points at this file twice — "read that
before trusting a result" — and the paragraph exists specifically to tell a
reader which rows of the model's device table are experience and which are
documented capacity. Someone deciding whether to act on a `--cc 7.5`
ranking was reading a sentence that said it had been validated on silicon.
Five of the six rows are capacity, not four.


- **`launchbound-runner` accepted a malformed `--budget-secs` in silence,
on the machine that costs money.** It parsed every value with
Expand Down
16 changes: 13 additions & 3 deletions docs/LIMITATIONS.md
Original file line number Diff line number Diff line change
Expand Up @@ -144,9 +144,19 @@ Two consequences worth stating:
and differ in SM count (L4 58 / L40 142, H100 SXM 132 / PCIe 114), and
the table picks one — the rows say which.

Only 8.6 (A10G) and 7.5 (T4) have ever had a kernel measured on them here.
The other four rows are documented capacity, not experience; the model's
Spearman correlations below were measured on the A10G alone.
**8.6 (A10G) is the only capability anything here has ever been measured
on.** The other five rows are documented capacity, not experience, and that
includes 7.5: `docs/research-baseline.md` records the tier-2 box as a
`g5.xlarge` with an A10G, "chosen over the T4 by the operator", and
`model-calibration.toml` names one device. A `--cc 7.5` ranking is the model
speaking about a part no kernel in this repository has run on, which is
exactly what "Results do not port" below means by a verdict that does not
transfer — the model's Spearman correlations were measured on the A10G
alone.

(This paragraph used to claim the T4 as well, twenty lines above the section
that says "nothing has been measured on a T4". Both cannot be true, and the
evidence in the repository is with the second one.)

## Measurement noise floor

Expand Down
Loading