From 08ec81a5e043e514950711d7722b0f1a87ce622b Mon Sep 17 00:00:00 2001 From: Vyncint Ng Date: Tue, 22 Sep 2026 16:38:20 +0700 Subject: [PATCH] docs: the CUDA path, re-measured on an A10G at the current pin set The first time anything automated has run it. The headline result in the README was produced by hand on 2026-08-20 and had not been reproduced since, while the lockstep set moved twice and LIMITATIONS documents a way for the compile step to break that nobody would have heard about. block_x=32 tile=128 came in at 0.033792 ms, which is the 0.0338 ms the README quotes -- from a run a month earlier, on a different instance of the same part, under a different analyzer and a different cuda-oxide. The gate admitted 3 and refused 8 as it does on a laptop, all three admitted candidates compiled, all three measured, and every structural invariant held. Recorded with its full provenance rather than as a number: host, driver, CUDA, rustc, analyzer, compiler and both pins. One thing noted rather than smoothed over: at ~30-75 microseconds two of the three 95% intervals are zero-width and every median is a multiple of 1024 ns, which is the CUDA event timer's granularity showing through. It does not affect this ranking, and it means the "overlapping intervals are indistinguishable" rule is doing less work at this scale than the noise-floor section implies. Signed-off-by: Vyncint Ng --- CHANGELOG.md | 20 +++++++++++ docs/research-baseline.md | 71 +++++++++++++++++++++++++++++++++++++++ 2 files changed, 91 insertions(+) diff --git a/CHANGELOG.md b/CHANGELOG.md index c9169ab..4d1ab89 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -9,6 +9,26 @@ change measured timings are marked `bench:`. ## [Unreleased] +### Changed + +- **The CUDA path has been run, on silicon, at the current pin set** (#83) — + the first time anything automated has. An A10G (cc 8.6, driver 595.71.05, + CUDA 13.2) with `cargo-reconverge 0.7.0`, `cargo-oxide 0.2.1` and + cuda-oxide `b0f961df`: the gate admitted 3 and refused 8, all three + admitted candidates compiled, and all three measured. + + **`block_x=32 tile=128` came in at 0.033792 ms — the 0.0338 ms the README + quotes**, from a run a month earlier on a different instance of the same + part under a different analyzer and a different cuda-oxide. The provenance + and the full numbers are in + [`docs/research-baseline.md`](docs/research-baseline.md). + + One observation recorded rather than smoothed over: at these durations + (~30–75 µs) two of the three 95% intervals are zero-width and every median + is a multiple of 1024 ns, which is the CUDA event timer's granularity + showing through. It does not affect this ranking, and it is noted where the + noise floor is discussed. + ## [2.3.0] - 2026-09-22 The production-readiness release. It moves the lockstep set onto a diff --git a/docs/research-baseline.md b/docs/research-baseline.md index afd293f..817f8c2 100644 --- a/docs/research-baseline.md +++ b/docs/research-baseline.md @@ -230,3 +230,74 @@ changes nothing this corpus can observe, and that the compile path works at the new pin on a machine with LLVM 21 and no CUDA toolkit. It does not establish that PTX *runs*: `inspect` lowers, it does not execute, and no GPU was involved in any line of this section. The corpus is still six kernels. + +## Re-measured: nightly-2026-08-28, cuda-oxide `b0f961df`, reconverge 0.7.0 + +2026-09-22, and the first time anything automated has run this project's CUDA +path (#83). The headline result in the README was produced by hand on +2026-08-20 and had not been reproduced since; the lockstep set moved twice in +between, and `cargo check` under the reconverge driver is documented as not +evaluating all codegen-time consts, so a gate-clean candidate failing the +real compile was a live possibility nobody would have heard about. + +### Provenance + +| | | +|---|---| +| host | AWS `g5.xlarge`, us-east-2c, **NVIDIA A10G** (`sm_86`, cc 8.6), 15 GiB RAM | +| driver | 595.71.05 | +| CUDA | 13.2 (`cuda_13.2.r13.2/compiler.37434383_0`) | +| rustc | 1.100.0-nightly (e457a7b0d 2026-08-27) | +| analyzer | `cargo-reconverge 0.7.0` (from crates.io) | +| compiler | `cargo-oxide 0.2.1`, built from the pinned checkout | +| cuda-oxide | `b0f961df3af0ff140b3b006fa2b6750b71f43f62` | +| launchbound | `v2.3.0` | + +Run unattended from EC2 user-data: `ssm:StartSession` is denied to the role +this account uses, and `gpu-sg` has no inbound rules, so the serial console +was the only channel back. The instance terminates itself. + +### The gate, on silicon + +``` +prune (cc 8.6) ... + 3 admitted, 8 refused; compiling admitted specializations ... +``` + +The same 3 / 8 split as on a laptop and as recorded for `reduce-flip` +throughout this repository. **All three admitted candidates compiled** — the +`#[unroll]`-class hole under "cuda-oxide is alpha" in +[LIMITATIONS](LIMITATIONS.md) did not bite at this pin. + +### The measurement + +``` +3/3 measured ok, 3.2 GPU-seconds, budget not exhausted + +0.033792 ms [0.033792, 0.034816] n=99 block_x=32 tile=128 +0.045056 ms [0.045056, 0.045056] n=75 block_x=32 tile=256 +0.074752 ms [0.074752, 0.074752] n=99 block_x=32 tile=512 +``` + +**`block_x=32 tile=128` at 0.0338 ms is the figure the README quotes**, from +a run a month earlier on a different instance of the same part, under a +different analyzer and a different cuda-oxide. That the two agree to the +digit published is the strongest evidence this repository has that its +measurement path is stable — and it is the first time the claim has been +checked rather than carried forward. + +Every structural invariant held: the results declare `results.v1`, no +candidate appears that the plan did not contain, no admitted candidate failed +to run, and something was measured. + +### One thing to look at, not a defect here + +Two of the three intervals are **zero-width** (`[0.045056, 0.045056]`), and +every median is a multiple of 1024 ns. That is the CUDA event timer's +granularity showing through at these durations, not a claim of perfect +reproducibility: `reduce-flip` at this size runs for ~30–75 µs, which is only +tens of timer ticks. It does not affect the ranking — the three are far apart +— but "two configurations whose intervals overlap are reported +indistinguishable" is doing less work at this scale than the noise-floor +section above implies, and a kernel this short deserves a note rather than a +silent zero. The A10G numbers in that section were taken on longer sweeps.