Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
35 changes: 35 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -36,6 +36,41 @@ jobs:
- uses: taiki-e/install-action@94c31af3204a9f15ab40b35ad084410b905bbc73 # v2.87.17
with:
tool: just,cargo-deny
# The Metal path, executed rather than merely compiled.
#
# `run_metal` is the function behind the README's "On Apple Silicon it
# measures on Metal", and until now nothing called it outside the CLI:
# its crate had one test, about rewriting a constant in a source
# string, and the macOS leg compiled the `cfg(target_os = "macos")`
# code and then ran that. The only evidence the path works was
# `runs/reduce-stable-metal`, one sweep on one machine, committed and
# never re-run.
#
# A smoke test, deliberately: it asserts that a dispatch happens and
# produces a readable `results.v1`, not that the numbers are anything.
# Timings from a shared runner are not measurements and this does not
# pretend otherwise -- `--budget 60s` bounds it, and `report` reading
# the run back is the assertion.
#
# `--cc` is required by the parser and ignored on this path, which has
# no gate to run at any capability.
- name: Measure one kernel on Metal
if: runner.os == 'macOS'
run: |
set -euo pipefail
launchbound() { cargo run -q -p launchbound-cli --bin launchbound -- "$@"; }
launchbound tune corpus/reduce-stable --cc 8.6 --backend metal \
--out "${RUNNER_TEMP}/metal-smoke" --budget 60s
launchbound report "${RUNNER_TEMP}/metal-smoke" --json > "${RUNNER_TEMP}/metal.json"
python3 - "${RUNNER_TEMP}/metal.json" <<'PY'
import json, sys
report = json.load(open(sys.argv[1]))
chosen = report.get("chosen")
assert chosen, "the Metal sweep chose nothing"
assert chosen["summary"]["median_ms"] > 0, chosen
print("metal ok:", chosen["config"], chosen["summary"]["median_ms"], "ms")
PY

- run: just ci
env:
# Every screen a failing termlens wait embeds is also written here,
Expand Down
56 changes: 56 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -29,6 +29,62 @@ change measured timings are marked `bench:`.
is pointed at reconverge, which this project ships as a component and does
not reimplement.

- **Tests for the part of this tool that decides what the number is.** 94
tests for 7,571 lines, and the distribution was the finding: the
best-tested crates were the ones that *present* a result or *refuse* a
candidate, and the ones that produce it were the thin ones. 123 now.

- **The exit-code contract**, which the README states and nothing checked.
`launchbound-cli` had three tests, all about `parse_budget`, and no
integration test at all. Exit 1 — "the fastest candidates were refused
and the chosen one is slower" — is the product's whole argument in one
integer, and it is the one a refactor can quietly turn into 0, because
the text on stdout is identical either way.

The recorded Metal sweep has no refused candidate (`gate: "none"`), so
asserting against it as-is would have been exactly the kind of test that
cannot fail. The test marks the **two** fastest candidates refused — two,
because their 95% intervals overlap each other and `rejected_faster`
requires the refused interval to sit wholly below the chosen one's, which
is the same "indistinguishable, never ranked" rule the ranking uses — and
then requires exit 1. Replacing the branch with `ExitCode::SUCCESS` fails
it; nothing else in the suite notices.

Also checked: `--allow-unsafe` without `--reason`, and with a blank one;
`--reason` without `--allow-unsafe`; that `tune` has no `--allow-unsafe`
at all; that a malformed `--cc` is refused before anything is spawned;
and that `--json` and the text report never disagree about the status.

- **Checkpoint and resume.** `run_plan` needs a device, but the two halves
that decide whether a resume is safe do not. A checkpoint misread as "no
checkpoint" silently re-measures candidates that cost GPU time, so each
untrustworthy shape — not JSON, `results.v2`, no `schema` at all — is now
required to fail with the path named. And the write-then-rename leaves no
`.json.tmp` behind.

- **The Metal path.** `launchbound-metal` had one test, about rewriting a
constant in a string. Six now: the no-gate notice is asserted as text
rather than trusted as a constant, the off-macOS stub must refuse rather
than return an empty sweep, a CUDA-only dimension with no MSL twin is
skipped rather than invented, and an unterminated constant is named
rather than silently truncating the source.

- **The model's ranking**, pinned against the real corpus kernel rather
than a synthetic space, as a relation rather than as numbers: the same
space ranks the same way twice, and sorting by cost is a total order with
nothing infinite in it. Every other test there checked a piece of the
cost function; none checked that the pieces compose into the same order.
Plus: an unknown capability must list the ones that exist.

- **CI executes a Metal dispatch.** The macOS leg compiled the
`cfg(target_os = "macos")` code and then ran a string-rewriting test;
`run_metal` — the function behind "On Apple Silicon it measures on Metal" —
had no caller outside the CLI, and the only evidence the path worked was
one committed sweep on one machine. A smoke step now tunes one corpus
kernel under a 60s budget and reads the run back. It asserts that a
dispatch happened and produced a readable `results.v1`, not that the
numbers mean anything: timings from a shared runner are not measurements.

### Fixed

- **The GitHub Release is created by the release workflow, and the semver
Expand Down
100 changes: 99 additions & 1 deletion crates/launchbound-bench/src/run.rs
Original file line number Diff line number Diff line change
Expand Up @@ -596,7 +596,105 @@ pub fn parse_budget(flag: &str, text: &str) -> Result<f64, String> {

#[cfg(test)]
mod budget_tests {
use super::parse_budget;
use super::{Results, parse_budget};

// ---- checkpoint and resume -------------------------------------------
//
// The checkpoint is the only record that a candidate was measured, and
// every one of them cost GPU time. `run_plan` cannot be exercised without
// a device, but the two halves that decide whether a resume is safe --
// reading the file and writing it atomically -- can be, and they are
// where a wrong answer is expensive: a checkpoint misread as "start over"
// silently re-measures, and one misread as valid resumes into nonsense.

fn scratch(name: &str) -> std::path::PathBuf {
let dir = std::path::Path::new(env!("CARGO_MANIFEST_DIR"))
.join("../../target/tmp")
.join(name);
let _ = std::fs::remove_dir_all(&dir);
std::fs::create_dir_all(&dir).expect("a scratch directory");
dir
}

fn sample_results() -> Results {
Results {
schema: "results.v1".into(),
kernel: "k".into(),
entry: "k".into(),
plan_cc: "8.6".into(),
device_name: "probe".into(),
device_cc: "8.6".into(),
driver_version: "0".into(),
candidates: Vec::new(),
total_gpu_seconds: 1.5,
strategy: Some("exhaustive".into()),
budget_exhausted: false,
}
}

#[test]
fn no_checkpoint_is_a_fresh_start_not_an_error() {
let dir = scratch("results-absent");
assert!(
Results::load(&dir.join("results.json"))
.expect("a missing file is not a failure")
.is_none()
);
}

#[test]
fn a_checkpoint_round_trips_and_leaves_no_temporary_behind() {
let dir = scratch("results-roundtrip");
let path = dir.join("results.json");
let written = sample_results();
written.checkpoint(&path).expect("checkpointing");

let read = Results::load(&path).expect("readable").expect("present");
assert_eq!(read.kernel, written.kernel);
assert_eq!(read.total_gpu_seconds, written.total_gpu_seconds);
assert_eq!(read.strategy, written.strategy);

// Write-then-rename: the `.json.tmp` must not survive, or the next
// reader of the directory finds two documents and no rule for which.
let leftovers: Vec<_> = std::fs::read_dir(&dir)
.expect("readable directory")
.filter_map(Result::ok)
.map(|e| e.file_name().to_string_lossy().into_owned())
.filter(|n| n.ends_with(".tmp"))
.collect();
assert!(leftovers.is_empty(), "left behind: {leftovers:?}");
}

#[test]
fn a_checkpoint_that_cannot_be_trusted_stops_the_run() {
// Each of these used to be indistinguishable from "no checkpoint",
// and resuming from that discards measurements that cost GPU time.
// The error must name the path, because the operator is looking at a
// run directory, not at a stack trace.
let dir = scratch("results-untrustworthy");
for (name, body, expected) in [
("not-json.json", "{ this is not json", "not JSON"),
(
"wrong-schema.json",
r#"{"schema":"results.v2","kernel":"k"}"#,
"unsupported results schema",
),
(
"no-schema.json",
r#"{"kernel":"k","candidates":[]}"#,
"no `schema` field",
),
] {
let path = dir.join(name);
std::fs::write(&path, body).expect("writing the probe");
let err = Results::load(&path).expect_err("must not be read as a fresh start");
assert!(err.contains(expected), "{name}: {err}");
assert!(
err.contains(&path.display().to_string()),
"{name}: the message must name the file: {err}"
);
}
}

#[test]
fn accepted_forms_are_seconds() {
Expand Down
Loading
Loading