From bf077e9a8dbdee3b6cc8c5870c946a6f149c0a5b Mon Sep 17 00:00:00 2001 From: Vyncint Ng Date: Tue, 22 Sep 2026 12:59:15 +0700 Subject: [PATCH] docs: the T4 has never run a kernel here docs/LIMITATIONS.md said both "Only 8.6 (A10G) and 7.5 (T4) have ever had a kernel measured on them here" and, twenty lines later, "nothing has been measured on a T4". The evidence is with the second: docs/research-baseline.md records the tier-2 box as a g5.xlarge with an A10G, chosen over the T4 by the operator, and model-calibration.toml names one device. Not a footnote. README points at this file twice -- read it before trusting a result -- and the paragraph exists to tell a reader which rows of the model's device table are experience and which are documented capacity. Someone deciding whether to act on a --cc 7.5 ranking was reading a sentence saying it had been validated on silicon. Five of the six rows are capacity, not four. Signed-off-by: Vyncint Ng --- CHANGELOG.md | 15 +++++++++++++++ docs/LIMITATIONS.md | 16 +++++++++++++--- 2 files changed, 28 insertions(+), 3 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 61651fe..40f7f2d 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -55,6 +55,21 @@ change measured timings are marked `bench:`. index: skipped with a notice, because a gate that fails offline is a gate people learn to skip. +- **`docs/LIMITATIONS.md` contradicted itself about the T4, and claimed a + measurement that never happened.** Line 147 said "Only 8.6 (A10G) and 7.5 + (T4) have ever had a kernel measured on them here"; line 167, in the + section on results not porting, said "nothing has been measured on a T4". + The evidence agrees with the second: `docs/research-baseline.md` records + the tier-2 box as a `g5.xlarge` with an A10G, *chosen over the T4 by the + operator*, and `model-calibration.toml` names one device. + + This is not a footnote. `README.md` points at this file twice — "read that + before trusting a result" — and the paragraph exists specifically to tell a + reader which rows of the model's device table are experience and which are + documented capacity. Someone deciding whether to act on a `--cc 7.5` + ranking was reading a sentence that said it had been validated on silicon. + Five of the six rows are capacity, not four. + - **`launchbound-runner` accepted a malformed `--budget-secs` in silence, on the machine that costs money.** It parsed every value with diff --git a/docs/LIMITATIONS.md b/docs/LIMITATIONS.md index 5ef235f..45de568 100644 --- a/docs/LIMITATIONS.md +++ b/docs/LIMITATIONS.md @@ -144,9 +144,19 @@ Two consequences worth stating: and differ in SM count (L4 58 / L40 142, H100 SXM 132 / PCIe 114), and the table picks one — the rows say which. -Only 8.6 (A10G) and 7.5 (T4) have ever had a kernel measured on them here. -The other four rows are documented capacity, not experience; the model's -Spearman correlations below were measured on the A10G alone. +**8.6 (A10G) is the only capability anything here has ever been measured +on.** The other five rows are documented capacity, not experience, and that +includes 7.5: `docs/research-baseline.md` records the tier-2 box as a +`g5.xlarge` with an A10G, "chosen over the T4 by the operator", and +`model-calibration.toml` names one device. A `--cc 7.5` ranking is the model +speaking about a part no kernel in this repository has run on, which is +exactly what "Results do not port" below means by a verdict that does not +transfer — the model's Spearman correlations were measured on the A10G +alone. + +(This paragraph used to claim the T4 as well, twenty lines above the section +that says "nothing has been measured on a T4". Both cannot be true, and the +evidence in the repository is with the second one.) ## Measurement noise floor