Skip to content

ci: a dispatchable workflow that re-runs the CUDA path - #95

Merged
vyncint merged 1 commit into
mainfrom
ci/gpu-verification-workflow
Sep 22, 2026
Merged

vyncint merged 1 commit into
mainfrom
ci/gpu-verification-workflow

Conversation

@vyncint

@vyncint vyncint commented Sep 22, 2026

Copy link
Copy Markdown
Owner

Nothing automated has ever run this project's CUDA path. There is no GPU job in any workflow, crates/launchbound-bench/src/cuda.rs has no test that executes, and the headline result in the README — the A10G sweep where six refused configurations measured up to 3.00x faster — was produced by hand on 2026-08-20 and has not been reproduced since.

Between then and now the lockstep pin set moved twice, cuda-oxide is three commits past the pin (#70), and docs/LIMITATIONS.md already records a known way for the compile step to break: cargo check under the reconverge driver does not evaluate all codegen-time consts, so a gate-clean candidate can still fail to build. Nothing would have told us if it had.

The shape

Two jobs, split the way stage and launchbound-runner were designed to be split:

  • stage (ubuntu-latest) prunes and compiles every admitted specialization, and builds the box-side binary beside the plan. So the expensive machine needs a CUDA driver and nothing else — no toolchain, no cuda-oxide checkout, no analyzer.
  • measure runs on ${{ inputs.runner }}, default gpu, records the device and driver before measuring anything, runs the sweep under an enforced --budget-secs, and uploads results.json on every exit path.

workflow_dispatch only. It never runs on a push or a schedule, because it costs money. With -f measure=false it stages a plan and stops, which is the path for a box you drive by hand:

./launchbound-runner --budget-secs 900 plan.json results.json

the same command the workflow issues.

What it gates on, and what it refuses to

The claims that port: the results declare results.v1; no candidate in the results is absent from the plan; no admitted candidate failed to run; something was measured. An admitted candidate that failed is reported by id rather than tolerated — each one names another codegen-time const the gate does not evaluate.

Not the timings. Those are valid only for the GPU, driver and compiler in their provenance, and sm_75 and sm_86 do not transfer (docs/LIMITATIONS.md, "Results do not port"). The job prints the five fastest with the device, driver and plan capability beside them, and uploads the results. Asserting a number here would be asserting exactly the thing this project documents as non-transferable.

What it does not do, deliberately

It does not provision the machine. This repository holds no cloud credentials, and adding one is a decision with a blast radius rather than a workflow detail. The issue's sketch had the workflow creating and terminating an instance; that is the half that belongs to whoever owns the account.

Not verified

There is no GPU available to this change, so the workflow is written, actionlint-clean and zizmor-clean, and unexercised. #83 asks for "one green run recorded, with its provenance, at the current pin set" and that part is not done — the issue should stay open until someone dispatches it against a real box.

Refs #83

Nothing automated had ever run it. There is no GPU job anywhere else
here, cuda.rs has no test that executes, and the headline result in the
README -- the A10G sweep where six refused configurations measured up to
3.00x faster -- was produced by hand on 2026-08-20 and has not been
reproduced since. The lockstep pin set moved twice in between, and
LIMITATIONS already records a known way for the compile step to break:
cargo check under the reconverge driver does not evaluate all
codegen-time consts, so a gate-clean candidate can still fail to build.
Nothing would have said so.

The split is the one the product already has. stage prunes and compiles
every admitted specialization on ubuntu-latest and the box-side binary
ships beside the plan, so the expensive machine needs a driver and
nothing else -- no toolchain, no cuda-oxide checkout, no analyzer.
Dispatch-only, because it costs money; with measure: false it stages a
plan for a box driven by hand, which is the same command the workflow
issues.

It gates on the claims that port -- the schema, no candidate in the
results absent from the plan, no admitted candidate that failed to run,
something measured -- and records the numbers with their device, driver
and plan capability beside them rather than asserting them. Timings are
exactly what does not transfer between parts.

It does not provision the machine. This repository holds no cloud
credentials and adding one is a decision with a blast radius rather than
a workflow detail.

Not yet run: there is no GPU available to this change, so the workflow is
written, linted and documented but unexercised. The issue stays open for
the first green run and the provenance it records.

Signed-off-by: Vyncint Ng <chivy.nguyen@manabie.com>
@vyncint
vyncint force-pushed the ci/gpu-verification-workflow branch from 2178691 to 2d10ac1 Compare September 22, 2026 06:24
@vyncint
vyncint merged commit 4620f18 into main Sep 22, 2026
11 checks passed
@vyncint
vyncint deleted the ci/gpu-verification-workflow branch September 22, 2026 06:33
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants