Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 0 additions & 1 deletion .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -27,7 +27,6 @@ ENV/

# uv
.uv/
uv.lock

# IDE
.vscode/
Expand Down
1 change: 1 addition & 0 deletions docs/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -63,5 +63,6 @@ Once allocated, workers launch inside containers, discover each other through et
- [Profiling](profiling.md) - Performance analysis with torch/nsys
- [SGLang Router](sglang-router.md) - Alternative to Dynamo for PD disaggregation
- [Services](services.md) - Sidecars and standalone stores launched next to the job
- [Prepared Jobs](prepared-jobs.md) - Immutable single-point inputs, literal clients, durable ownership and recovery
- [Configuration Reference](config-reference.md) - Every recipe section, with the generated field tables in [Schema Reference](schema-reference.md)
- [Legacy (v1) layout](legacy-v1.md) - The old `backend:` recipe layout; `srtctl migrate` rewrites it
48 changes: 48 additions & 0 deletions docs/phase1-validation.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,48 @@
# Phase 1 native prerequisite validation

Implementation starts from `984180e5b8755aef85e9995048b5a16cb5336bce` in a separate
worktree. The original NVIDIA baseline remains clean. No Slurm allocation, GPU
run, workflow dispatch, push, or publication was performed by these checks.

The locked Python 3.12 source suite on macOS completed with **2453 passed,
2 skipped, 6 integration tests deselected, and 5 failures**. All five failures
were independently reproduced against the untouched baseline on the same host:

- Three `test_sa_bench_http_reuse_flag_reaches_warmup_and_formal` variants fail
under macOS Bash 3 while resolving the sourced profiling helper.
- `test_probe_cpu_captures_affinity_and_slurm_allocation` assumes Linux's
`os.sched_getaffinity`, absent on macOS.
- `test_apply_mock_emits_submission_json_and_spawns_worker` assumes a Linux CPU
affinity count and sees `effective_for_check: null` on macOS.

Commands, using environments provisioned from committed `uv.lock`:

```text
uv sync --frozen --no-editable --python 3.12
PYTHONPATH=src python -m pytest tests/ -q --disable-warnings
ruff check src/srtctl/
ruff format --check src/srtctl/
PYTHONPATH=src python -m srtctl.cli.submit schema-docs --check
git diff --check
```

Ruff, formatting, generated schema checks and diff checks pass. The new native
behavior suite has **26 passing tests**, including real recording child argv,
environment and cwd; strict preparation before scheduler access; atomic duplicate
claim races; a killed submitter followed by recovery in a fresh interpreter;
required-server exit zero; repeated TERM during bounded cleanup; active controller
precedence; comment-scoped cancellation of multiple IDs; and a real relative HF
blob symlink read through the emitted model argument and native mount map.

A noneditable installed wheel was built with
`uv sync --frozen --no-dev --no-editable`; its isolated `python -I -m
srtctl.cli.submit prepare` succeeded using an explicit source profile and full
interpreter/distribution/source verification. Native config, model staging,
profiling and router regression checks passed after the model-path amendment.

These checks qualify CPU behavior only. The consumer still needs PR-linked H100
full-duration curves, real c28 evaluation, Pyxis mount/device/environment checks,
remote step cancellation and writer closure, and any enabled telemetry. The
cluster's filesystem must support the atomic directory claims and fsync durability
used by the intent journal. The six deselected tests use a real AIPerf integration
dependency; they are separate from the consumer's model/hardware acceptance.
145 changes: 145 additions & 0 deletions docs/prepared-jobs.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,145 @@
# Prepared single-job execution

The prepared API is an opt-in path for a consumer that has already selected one
benchmark point. Existing `apply`, dependency Bash templates, and cluster wrappers
remain supported. It starts at NVIDIA revision `984180e5b8755aef85e9995048b5a16cb5336bce`.

Install this revision into a dedicated compute-compatible environment using the
committed dependency lock, for example `uv sync --frozen --no-dev --no-editable`.
The environment and pinned source checkout must be readable on the compute host.
This provisioning step precedes submission; batch startup performs no installation.
Set the explicit profile's `srtctl_root` to that source checkout. Installed package
resources must match its source bytes. Generated package version metadata is
covered by the installed distribution inventory, separately from source identity.

```text
python -I -m srtctl.cli.submit capabilities --json
python -I -m srtctl.cli.submit prepare --recipe point.yaml --profile site.yaml \
--output /shared/prepared/point --expected-nodes 1 \
--runtime-python /shared/runtime/bin/python --json
python -I -m srtctl.cli.submit submit-prepared --prepared-dir /shared/prepared/point \
--intent run-attempt-point --cluster site --journal-dir /shared/intents --json
python -I -m srtctl.cli.submit wait --receipt /shared/intents/site/<token>/receipt.json \
--timeout 28800 --poll 5 --json
```

All commands emit one JSON object. `prepare` returns `prepared_dir`,
`manifest_sha256`, `resources`, `output_root`, and `capabilities`. It rejects
missing/malformed profiles, duplicate YAML keys, nested sweeps, orchestration flag
overrides, and queued demand that disagrees with native physical placement.
The resulting directory contains resolved `config.yaml`, frozen `profile.yaml`,
`job.slurm`, and a manifest. Source/profile path expansion is frozen before
publication. The full inputs exist before `sbatch` can start the batch script.

The manifest binds interpreter bytes, architecture, installed distribution files,
native source/resources, lockfile, recipe/profile input hashes, and allocation
summary. Compute startup validates these before creating any serving step and
checks actual allocated node cardinality. Mutable installed files or source edits
require preparation of a new bundle. Preserve the qualified environment and
source for every in-flight intent.

Prepared single-node direct-vLLM jobs also bind readiness to the launched worker.
Startup rejects an occupied public port. A standard-library container wrapper
records the worker's Linux PID, start ticks, and PID/network namespaces before
replacing itself with the unchanged engine command. Before HTTP readiness and
again before launching the client, the host verifies that every listening socket
on the public port belongs to that process or one of its current descendants.
A foreign listener with the same served model name is rejected, including one
that appears after the initial free-port probe.

This narrow path requires shared host PID/network namespaces and readable Linux
`/proc` process/socket records; unavailable evidence fails closed. Linux CI tests
real sockets and exec/descendant ownership. The actual cluster must separately
qualify visibility through its Pyxis/enroot setup. The existing one-second client
poll loop and final success boundary recheck ownership, so a surviving parent
cannot hide a vanished or replaced serving descendant. Detected ownership loss
stops and reaps the client and fails execution. Polling cannot guarantee zero
packets after a takeover between checks; it never accepts observed ownership loss. Other
frontends and non-prepared jobs retain their existing launch behavior.

## Literal client hook

```yaml
benchmark:
type: custom
argv: [/shared/client/bin/python, -m, example.client, --policy, /ix/policy.json]
cwd: /ix
env: {EXPLICIT_SETTING: value}
env_unset: [INHERITED_SIMULATION_SETTING]
concurrencies: [28]
```

`argv` and the existing shell `command` are mutually exclusive. Arguments remain
literal across the native Slurm wrapper, including JSON, spaces, `$`, backticks,
and empty argument values. The selected container must expose the declared cwd
and executable through its mounts. The hook receives `SRT_ENDPOINT`,
`SRT_FRONTEND_HOST`, `SRT_FRONTEND_PORT`, `SRT_JOB_ID`, `SRT_LOG_DIR` (`/logs`),
`SRT_MODEL_NAME`, and existing logical-worker/metrics endpoint variables. Clients
cannot override or unset the runtime-owned `SRT_*` context.

For local Hugging Face snapshots with relative `../../blobs` shard links, mount
the complete cache root at the same canonical absolute path. Native
`RuntimeContext.worker_model_arg` preserves the snapshot path under that mount,
and vLLM uses it directly. This keeps blob links readable without copying model
weights. A snapshot-only `/model` mount does not preserve links outside it.

## Ownership, observation and recovery

The journal namespace is `(cluster, intent_id, generation=0)`. Atomic directory
claiming precedes `sbatch`; an existing claim never authorizes another allocation.
The CLI strips inherited scheduler option variables, requests `--no-requeue`, and
adds the intent token as the Slurm comment. Receipt states are `claimed`,
`accepted`, and `unknown`; accepted IDs remain available even if secondary
bookkeeping fails. Child stdout is an independent acceptance journal.

Before starting an interruptible submission, call `intent-path --intent ID
--cluster SITE --journal-dir DIR --json` to obtain its `receipt_path`. This pure
query creates nothing and lets a parent recover ownership even when interrupted
before submission stdout arrives. Preserve the recorded `SLURM_CONF` and
`SLURM_CONF_SERVER` context for recovery; changing it is rejected.

Use `reconcile --receipt PATH --json` after an unknown response or process crash.
It reads the independent journals and matching controller/accounting comments;
it never calls `sbatch`. Ambiguous or absent evidence remains unknown. Keep every
reported `accepted_ids` entry for operator resolution. A lost claim before its
receipt was published blocks resubmission and needs journal inspection; deleting
an unknown intent is not a retry procedure.

`cancel --receipt PATH --json` targets only the numeric job in the accepted
receipt. It reports a cancellation request, not completed cleanup. Follow it
with bounded `wait` and inspect the terminal state. Run these commands against
the same site's Slurm controller used to submit; the cluster field is the caller's
site namespace, not automatic cross-cluster routing.

After cancellation use `wait --until-terminal`; a requeue verdict can fail the
workload while the allocation is still active. For an ambiguous receipt with
multiple known IDs, use `cancel-known --receipt PATH --json`, followed by
`wait-known --receipt PATH --timeout 120 --poll 2 --json`. Each ID's comment is
verified independently. `wait-known` returns `state: closed, terminal: true`
only when every owned ID has terminal evidence. A foreign comment is never
authorized for cancellation.

Prepared mode opens scheduler logs under the precreated `output_root/native-logs`
directory, so Slurm does not need a job-ID-specific directory before executing
the batch. The job's `logs/sweep_<id>.log` points to its scheduler log.

Live `scontrol` state takes precedence over accounting. Accounting can establish
completion only after explicit controller absence, with exactly one original
generation record. Requeue, controller errors, missing accounting, or observation
timeout cannot become success. `COMPLETED/0:0` additionally requires matching
runtime completion evidence with successful execution, closed writers and node
restoration. Failed/cancelled runs retain diagnostics and never qualify as a
successful benchmark.

Critical server exit, including exit zero before the client finishes, fails the
workload. Signals request shutdown without acquiring cleanup locks; repeated
signals do not reenter cleanup. The registry uses one monotonic deadline across
all process tiers and reports incomplete cleanup conservatively. Required setup
scripts fail if absent. Host teardown failures remain a separate restoration
verdict and do not replace the original workload failure.

CPU tests cover real recording children and scheduler executables, concurrent
claims, a killed submitter followed by fresh-process reconciliation, and repeated
TERM during cleanup. Actual Pyxis device visibility, remote-step signal delivery,
H100 full curves, real evaluation, cancellation, and telemetry still require the
consumer's PR-linked cluster qualification; a CPU test is not that evidence.
3 changes: 3 additions & 0 deletions docs/schema-reference.md
Original file line number Diff line number Diff line change
Expand Up @@ -185,6 +185,9 @@ Benchmark configuration.
| `use_chat_template` | bool | `True` | Pass --use-chat-template to benchmark (default: true) |
| `reuse_http_connections` | bool | `False` | SA-Bench Dynamo adapter: reuse a benchmark-scoped HTTP connection pool. Opt-in to preserve the historical per-request ClientSession behavior. |
| `command` | str \| None | `None` | Custom benchmark hook. ``command`` is passed to ``bash -lc`` verbatim; srtctl does NOT substitute placeholders like ``{nginx_url}`` or ``{slurm_job_id}``. Render any parameters when generating the recipe. See srtctl.benchmarks.custom.CustomBenchmarkRunner for details. |
| `argv` | list[str] \| None | `None` | Literal executable and arguments for a custom client; mutually exclusive with command. No shell or placeholder expansion is performed. |
| `cwd` | str \| None | `None` | Container working directory for the custom client. |
| `env_unset` | list[str] | `[]` | Variables removed from the inherited client environment. |
| `container_image` | str \| None | `None` | |
| `env` | dict[str, str] | `{}` | |
| `aiperf_package` | str \| None | `None` | aiperf pip install spec (e.g., "aiperf>=0.7.0", "aiperf @ git+https://...@commit") If set, runs pip install <spec> before benchmarking. Upgrades if already installed. |
Expand Down
8 changes: 4 additions & 4 deletions src/srtctl/backends/vllm.py
Original file line number Diff line number Diff line change
Expand Up @@ -996,10 +996,8 @@ def build_worker_command(
else:
leader_ip = get_hostname_ip(endpoint_nodes[0])

# Determine model path: HF model ID or container mount path
# For HF models (hf:prefix), model_path contains the HF model ID (e.g., "facebook/opt-125m")
# For local models, model is mounted to /model in the container
model_arg = str(runtime.model_path) if runtime.is_hf_model else "/model"
# Honor native staging and same-path cache mounts (HF blob symlinks).
model_arg = runtime.worker_model_arg

# Get served model name from config or use model path name
served_model_name = self.get_served_model_name(runtime.model_path.name)
Expand Down Expand Up @@ -1466,6 +1464,8 @@ def _config_to_cli_args(config: dict[str, Any]) -> list[str]:
elif isinstance(value, list):
args.append(f"--{flag_name}")
args.extend(str(v) for v in value)
elif isinstance(value, dict):
args.extend([f"--{flag_name}", json.dumps(value, separators=(",", ":"), allow_nan=False)])
elif value is not None:
args.extend([f"--{flag_name}", str(value)])
return args
8 changes: 8 additions & 0 deletions src/srtctl/benchmarks/base.py
Original file line number Diff line number Diff line change
Expand Up @@ -95,6 +95,14 @@ def get_environment(self, config: SrtConfig, runtime: RuntimeContext) -> dict[st
"""Get benchmark-specific environment variables."""
return {}

def get_environment_unset(self, config: SrtConfig, runtime: RuntimeContext) -> list[str]:
"""Variables explicitly removed before executing the child."""
return list(config.benchmark.env_unset) if config.benchmark.type == "custom" else []

def get_working_directory(self, config: SrtConfig, runtime: RuntimeContext) -> str | None:
"""Working directory inside the selected container."""
return config.benchmark.cwd if config.benchmark.type == "custom" else None


class AIPerfBenchmarkRunner(BenchmarkRunner):
"""Base class for AIPerf-driven benchmarks.
Expand Down
28 changes: 24 additions & 4 deletions src/srtctl/benchmarks/custom.py
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,7 @@

from __future__ import annotations

import re
from pathlib import Path
from typing import ClassVar

Expand Down Expand Up @@ -45,7 +46,9 @@ class CustomBenchmarkRunner(BenchmarkRunner):
"""

# BenchmarkConfig fields this runner reads (beyond the shared ones); see benchmark_config_fields().
config_fields: ClassVar[frozenset[str]] = frozenset({"command", "container_image", "env"})
config_fields: ClassVar[frozenset[str]] = frozenset(
{"command", "argv", "cwd", "env_unset", "container_image", "env"}
)

@property
def name(self) -> str:
Expand All @@ -56,12 +59,29 @@ def script_path(self) -> str:
return "<custom command>"

def validate_config(self, config: SrtConfig) -> list[str]:
if config.benchmark.command:
return []
return ["benchmark.command is required for benchmark.type=custom"]
b = config.benchmark
errors = []
if bool(b.command) == bool(b.argv):
errors.append("Exactly one of benchmark.command or benchmark.argv is required for benchmark.type=custom")
if b.argv is not None and (not b.argv or any(not isinstance(a, str) or "\0" in a for a in b.argv)):
errors.append("benchmark.argv must be a nonempty list of NUL-free strings")
if b.argv and not b.argv[0]:
errors.append("benchmark.argv executable must not be empty")
if b.cwd is not None and (not b.cwd.startswith("/") or "\0" in b.cwd):
errors.append("benchmark.cwd must be an absolute container path")
for key in [*b.env, *b.env_unset]:
if not re.fullmatch(r"[A-Za-z_][A-Za-z0-9_]*", key):
errors.append(f"Invalid client environment variable: {key!r}")
if set(b.env) & set(b.env_unset):
errors.append("benchmark.env and benchmark.env_unset must be disjoint")
if any(key.startswith("SRT_") for key in [*b.env, *b.env_unset]):
errors.append("Runtime SRT_* client context cannot be overridden or removed")
return errors

def build_command(self, config: SrtConfig, runtime: RuntimeContext) -> list[str]:
del runtime
if config.benchmark.argv is not None:
return list(config.benchmark.argv)
assert config.benchmark.command is not None
return ["bash", "-lc", config.benchmark.command]

Expand Down
10 changes: 9 additions & 1 deletion src/srtctl/cli/do_sweep.py
Original file line number Diff line number Diff line change
Expand Up @@ -327,15 +327,18 @@ def _run_host_teardown(self) -> None:
-- a failed teardown must not overwrite the job's real exit code.
"""
setup = self.config.host_setup
self.restoration_success = True
if not setup.teardown or not getattr(self, "_host_setup_ran", False):
return
try:
failures = self._run_host_commands(setup.teardown, phase="teardown")
except Exception:
# Cleanup path: a teardown failure must never mask the job's result.
logger.exception("host_setup teardown raised; node state may need manual cleanup")
self.restoration_success = False
return
if failures:
self.restoration_success = False
logger.error(
"host_setup teardown failed on %s; those nodes may be left in a modified state",
", ".join(failures),
Expand Down Expand Up @@ -734,12 +737,17 @@ def run(self) -> int:

finally:
logger.info("Cleanup")
if stop_event.is_set() or registry.check_failures():
exit_code = exit_code or 1
registry.finalizing = True
# NOTE: finalize before registry.cleanup() so samples and manifest are durable.
exit_code = self.finalize_power_telemetry(exit_code, interrupted=stop_event.is_set())
exit_code = self.finalize_cpu_power_telemetry(exit_code, interrupted=stop_event.is_set())
exit_code = self.finalize_cpu_power_host_telemetry(exit_code, interrupted=stop_event.is_set())
stop_event.set()
registry.cleanup()
self.cleanup_complete = registry.cleanup() and self.benchmark_child_allows_window_mutation is not False
if not self.cleanup_complete:
exit_code = exit_code or 1
# After cleanup so the GPUs are idle before node state is reverted.
self._run_host_teardown()
if exit_code != 0:
Expand Down
Loading
Loading