ci: a dispatchable workflow that re-runs the CUDA path - #95
Merged
Merged
Conversation
Nothing automated had ever run it. There is no GPU job anywhere else here, cuda.rs has no test that executes, and the headline result in the README -- the A10G sweep where six refused configurations measured up to 3.00x faster -- was produced by hand on 2026-08-20 and has not been reproduced since. The lockstep pin set moved twice in between, and LIMITATIONS already records a known way for the compile step to break: cargo check under the reconverge driver does not evaluate all codegen-time consts, so a gate-clean candidate can still fail to build. Nothing would have said so. The split is the one the product already has. stage prunes and compiles every admitted specialization on ubuntu-latest and the box-side binary ships beside the plan, so the expensive machine needs a driver and nothing else -- no toolchain, no cuda-oxide checkout, no analyzer. Dispatch-only, because it costs money; with measure: false it stages a plan for a box driven by hand, which is the same command the workflow issues. It gates on the claims that port -- the schema, no candidate in the results absent from the plan, no admitted candidate that failed to run, something measured -- and records the numbers with their device, driver and plan capability beside them rather than asserting them. Timings are exactly what does not transfer between parts. It does not provision the machine. This repository holds no cloud credentials and adding one is a decision with a blast radius rather than a workflow detail. Not yet run: there is no GPU available to this change, so the workflow is written, linted and documented but unexercised. The issue stays open for the first green run and the provenance it records. Signed-off-by: Vyncint Ng <chivy.nguyen@manabie.com>
vyncint
force-pushed
the
ci/gpu-verification-workflow
branch
from
September 22, 2026 06:24
2178691 to
2d10ac1
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Nothing automated has ever run this project's CUDA path. There is no GPU job in any workflow,
crates/launchbound-bench/src/cuda.rshas no test that executes, and the headline result in the README — the A10G sweep where six refused configurations measured up to 3.00x faster — was produced by hand on 2026-08-20 and has not been reproduced since.Between then and now the lockstep pin set moved twice, cuda-oxide is three commits past the pin (#70), and
docs/LIMITATIONS.mdalready records a known way for the compile step to break:cargo checkunder the reconverge driver does not evaluate all codegen-time consts, so a gate-clean candidate can still fail to build. Nothing would have told us if it had.The shape
Two jobs, split the way
stageandlaunchbound-runnerwere designed to be split:stage(ubuntu-latest) prunes and compiles every admitted specialization, and builds the box-side binary beside the plan. So the expensive machine needs a CUDA driver and nothing else — no toolchain, no cuda-oxide checkout, no analyzer.measureruns on${{ inputs.runner }}, defaultgpu, records the device and driver before measuring anything, runs the sweep under an enforced--budget-secs, and uploadsresults.jsonon every exit path.workflow_dispatchonly. It never runs on a push or a schedule, because it costs money. With-f measure=falseit stages a plan and stops, which is the path for a box you drive by hand:the same command the workflow issues.
What it gates on, and what it refuses to
The claims that port: the results declare
results.v1; no candidate in the results is absent from the plan; no admitted candidate failed to run; something was measured. An admitted candidate that failed is reported by id rather than tolerated — each one names another codegen-time const the gate does not evaluate.Not the timings. Those are valid only for the GPU, driver and compiler in their provenance, and
sm_75andsm_86do not transfer (docs/LIMITATIONS.md, "Results do not port"). The job prints the five fastest with the device, driver and plan capability beside them, and uploads the results. Asserting a number here would be asserting exactly the thing this project documents as non-transferable.What it does not do, deliberately
It does not provision the machine. This repository holds no cloud credentials, and adding one is a decision with a blast radius rather than a workflow detail. The issue's sketch had the workflow creating and terminating an instance; that is the half that belongs to whoever owns the account.
Not verified
There is no GPU available to this change, so the workflow is written,
actionlint-clean andzizmor-clean, and unexercised. #83 asks for "one green run recorded, with its provenance, at the current pin set" and that part is not done — the issue should stay open until someone dispatches it against a real box.Refs #83