User-facing CLIs (experiment, deploy) and stdlib-only drift guards that protect invariants between two on-disk sources of truth (CI vs. pyproject, Python vs. C++ constants).
| Script | Role |
|---|---|
experiment.py (make experiment) |
Click group with run, tune, compare, holdout-eval, importance, study, clean subcommands. Drives the orchestration layer end to end. importance recomputes a finished run's feature importance by re-running its frozen config: it backfills feature_importance.json in place when the re-fit reproduces the run's metrics, or saves a separate source_run-tagged run when training is non-deterministic and the re-fit diverges. |
deploy.py |
Click group with create, predict, list, show, signals subcommands. Live-inference layer over a frozen trained run; idempotent on --as-of. |
check_ci_deps.py |
Drift guard: every runtime dep in pyproject.toml appears in CI's python-test pip install line; every types-* / *-stubs dev dep appears in CI's lint-and-typecheck pip install line. Runs in CI as an early lint step. (The webapp + webapp-frontend jobs use pip install -e ".[webapp]" so their installs cannot drift from [webapp] extras.) |
check_constants_sync.py |
Drift guard: every numeric constant mirrored between src/core/constants.py and cpp/include/quant/core/types.hpp (trading-calendar counts, position limits) has the same value on both sides. Pairs to verify are listed in MIRROR_PAIRS in the script. |
dump_openapi.py (make webapp-openapi-snapshot) |
Boot FastAPI, write its OpenAPI 3.1 spec to webapp/frontend/openapi.snapshot.json (the committed contract that npm run gen:api reads). |
check_openapi_snapshot.py (make webapp-check-openapi-snapshot) |
Drift guard: re-build the OpenAPI spec, fail if it diverges from the committed snapshot. Tells the developer to rerun make webapp-openapi-snapshot and commit. |
check_webapp_schema_mirror.py (make webapp-check-schema-mirror) |
Drift guard: extract every Pydantic write-DTO mirrored as a zod schema (LoginRequest, UserCreate) into webapp/frontend/schema-mirror.snapshot.json; pair vitest test asserts the zod schema agrees on field names, types, min/max constraints. --write regenerates the snapshot. |
regen_stubs.py (make stubs) |
Regenerate quant_engine pybind11 stubs and apply ruff lint / format so the checked-in artefact passes make lint. |
regen_spy_fixture.py |
Refetch + normalize + validate tests/fixtures/SPY.parquet (SPY daily, 2018-01-01 -> 2024-12-31, auto_adjust=True). Run when the committed fixture goes stale. |
validate_cam_baskets.py |
Validate the CrossAssetMomentum baskets: reads each basket (primary + feature tickers + lags) from main_study.yaml and the universe profiles, then prints the contemporaneous return correlation (shared structure, not near-duplicate) and the peer-momentum lead-lag table from yfinance, with a pass / flag verdict per basket. Not run in CI. |
validate_pairs_candidates.py |
Screen candidate PairsTrading pairs before committing them to a study: for each economically linked candidate, fetch closes from yfinance, slice to the in-sample dev window (same holdout boundary the runner reserves), and report Engle-Granger cointegration, mean-reversion half-life, and whether the spread amplitude clears the normal round-turn cost wall, with a pass / flag verdict + ranked recommendation. Surfaced v_ma (the lone survivor of the twenty-candidate panel) for config/study/pairs_extension.yaml. Not run in CI. |
backfill_save_markers.py |
One-time migration: re-mark model save directories persisted before the .save_complete marker existed so they load again. Walks the store; for each run missing markers, writes them only if the strategy then loads (the completeness oracle) - a save that fails to load is reverted and reported. Model data is never modified. Idempotent; --dry-run lists without writing. |
plot_xgb_gain.py |
Render a tree strategy's native XGBoost-gain feature importance as a top-N horizontal bar chart, in the study report's permutation-bar style. Aggregates the xgb_gain entries already persisted per leg in feature_importance.json (mean across the strategy's legs) and writes <study-dir>/plots/feature_importance/<Strategy>_gain.{png,svg}. Generates the gain figure embedded in the results chapter. Not run in CI. |
| File | Role |
|---|---|
experiment.py |
experiment run / tune / compare / holdout-eval / study / clean. |
study.py |
experiment study run / report - sub-group registered under experiment.py's cli. |
deploy.py |
deploy create / predict / list / show / signals - thin click wrapper over src/orchestration/deployment.py. |
check_ci_deps.py |
Stdlib-only (no PyYAML) so it runs in CI before deps install. |
check_constants_sync.py |
Stdlib-only; text-parses both files via regex so it runs in the same early-CI lint step as check_ci_deps.py. |
dump_openapi.py |
Lazy-imports webapp.backend.app.main so --out consumers without webapp deps can still load the module's DEFAULT_SNAPSHOT_PATH constant. |
check_openapi_snapshot.py |
Reuses dump_openapi.build_openapi_spec() for the live spec; lazy webapp import keeps diff_against_snapshot testable without fastapi. |
check_webapp_schema_mirror.py |
Walks model_fields on Pydantic v2 models; --write regenerates the snapshot, default mode diffs. Frontend pair: webapp/frontend/tests/lib/schemas/mirror.test.ts. |
regen_stubs.py |
Wraps pybind11-stubgen + ruff check --fix + ruff format. |
regen_spy_fixture.py |
Wraps yfinance.download + DataNormalizer + validate_bars; not run in CI. |
validate_cam_baskets.py |
Reads CrossAssetMomentum baskets via load_study_spec + compose_leg_config; fetches returns with yfinance.download; correlation + lead-lag diagnostics. Not run in CI. |
validate_pairs_candidates.py |
Stdlib argparse CLI over a hard-coded candidate panel; reuses CointegrationTester.engle_granger, resolve_holdout_boundary, and the normal cost tier; AR(1) half-life + rolling-z-score trade-frequency proxy. Fetches closes with yfinance.download. Not run in CI. |
backfill_save_markers.py |
Provisional-mark -> load_strategy_from_run_dir certify -> keep-or-revert. Pairs with tests/unit/test_backfill_save_markers.py. |
plot_xgb_gain.py |
Click CLI; reuses read_aggregated_importance + the src/visualization/plots style constants and save_png_and_svg; filters the persisted importances to ImportanceMethod.XGB_GAIN. Not run in CI. |
| Command | Output | Notes |
|---|---|---|
run --config <yaml> [--feature-importance] |
experiment_results/runs/<experiment_id>/ |
Walk-forward -> manifest.json + fold_results.jsonl + metrics.json + optional strategy_state/. --feature-importance adds OOS per-fold feature importance (permutation + XGBoost gain) as feature_importance.json for feature-consuming strategies. |
tune --config <exp.yaml> --hpo-config <hpo.yaml> |
experiment_results/hpo/<study>/ |
Optuna study via StrategyTuner; resumable. |
compare --config <yaml> ... --out <name> [--reuse-runs <dirs>] |
experiment_results/comparisons/<out>/ |
N strategies on aligned data, ranked + pairwise-bootstrapped. With --reuse-runs <a,b,...> (one path per --config in matching order) the per-strategy walk-forward step is skipped and prior fold artifacts feed ranking + bootstrap. |
holdout-eval --run-dir <path> | --hpo-best <path> |
experiment_results/holdout_evals/<out>/ |
Refit on full dev, evaluate once on the reserved holdout - the honest one-shot OOS number. Sources are mutually exclusive; manifest cross-checks holdout_start + data_hash before fitting. |
study run --spec <yaml> |
<store_root>/<spec.output_dir>/ |
Cross-strategy x cross-universe sweep: tune -> run -> holdout-eval per leg, then per-universe cross-strategy compare. Resumable via study_state.json; per-leg failures isolated. |
study report --study-dir <path> |
<study_dir>/{tables,plots,manifest.json} |
Walk a completed study tree; emit master / per-universe / holdout rankings (.tex+.csv), dev-vs-holdout scatter, and per-universe equity-overlay / per-leg holdout-equity copies. Read-only with respect to the per-leg tree. |
clean [--store-root experiment_results] [--apply] [--keep <name>] |
<store-root>/ |
Remove ephemeral child directories under the store root (default: experiment_results/), plus a small allowlist of stray top-level sweep-tracking files (.sweep_pid, .sweep_started_at, .sweep_log_path, sweep_*.log). Preserves any directory (or stray-file name) given via --keep and every other top-level file; refuses to delete any directory containing git-tracked files. Default = dry-run; pass --apply to delete. |
Multi-ticker pairs configs route through experiment run with no
special flag. The builder dispatches on the strategy class's
is_pairs_strategy capability flag.
| Command | Output | Notes |
|---|---|---|
create --from-run <run_id> | --from-hpo <study_name> [--name X] [--warmup-bars N] |
experiment_results/deployments/<id>/ |
Pins a deployment to a completed run or HPO study (best trial). Sources are mutually exclusive. Auto-generates name as <ticker>-<strategy>-<train_end> (or ...-HPO-<study> for HPO). |
predict <deployment_id> [--as-of YYYY-MM-DD] |
one row appended to signals.jsonl |
Generates (or recalls) the signal for the latest complete bar through --as-of. Idempotent on the target bar_ts. Anti-leakage: refuses to act on a bar <= train_end. |
list |
one line per deployment | Sorted by id. |
show <deployment_id> |
manifest JSON | The typed Deployment for one id. |
signals <deployment_id> [--limit N] |
one JSON row per line | Most-recent first. |
A deployment is pinned to one trained run. Training a fresher model is
a separate concern: use quant experiment run, then create a new
deployment pointing at the resulting run.
| Flag | Applies to | Role |
|---|---|---|
--override key.path=value (repeatable) |
run, tune, compare |
Dotted-path mutation of the loaded YAML before pydantic re-validation. Value parsed with yaml.safe_load; intermediate keys must already exist (typo guard). On compare the same set applies to every --config. |
--publish-label <slug> |
run, compare, holdout-eval |
Stable LaTeX \caption + \label slug for the emitted tables. When unset the legacy volatile id (experiment_id / out_name / source_id) is used; when set the slug overrides it so thesis-prose \ref stays valid across reruns. Slug regex: starts with a letter, then letters / digits / _ / - / :. |
Every persistent CLI subcommand (run, tune, compare,
holdout-eval, study run, study report) tees its full Python-logger
stream to a timestamped file under
<store_root>/cli_logs/<command>_<UTC_YYYYMMDD_HHMMSS>_<pid>.log (or
<study_dir>/cli_logs/... for study report). The first line of every
invocation echoes the resolved log path, e.g.
running study from spec ... -> log: experiment_results/cli_logs/study_run_20260504_193000_12345.log.
The file contains the same lines tail -f would show on stdout (same
formatter), so a dropped terminal during a multi-day sweep doesn't lose
the diagnostic trail. Utility commands (clean) skip the file handler
since they're fast and produce no logger output worth persisting.
The same stream is also appended to one shared
<store_root>/cli_logs/combined.log, so the full history of a store
reads top-to-bottom in a single file while the per-invocation files stay
around for isolating one command. Each invocation opens with a
===== <command> <UTC_ts> pid=<pid> ===== banner in the combined file to
keep the boundaries legible.
- CI <-> pyproject. A new runtime / type-stub dep landing in
pyproject.tomlwithout a matching update to.github/workflows/ci.ymlfailscheck_ci_deps.pyin the same PR. Thewebapp+webapp-frontendjobs usepip install -e ".[webapp]", so their installs cannot drift from[webapp]extras and need no separate guard. - Python <-> C++ constants. Numeric scalars mirrored between
src/core/constants.pyandcpp/include/quant/core/types.hpp(trading-calendar counts, position limits) must agree value-for-value;check_constants_sync.pyflags any divergence so an annualization-factor edit on one side without the other fails CI lint. - OpenAPI snapshot <-> live FastAPI app.
check_openapi_snapshot.pyre-dumps the spec and fails if it diverges fromwebapp/frontend/openapi.snapshot.json. The committed snapshot is whatnpm run gen:apireads to emitsrc/api/generated/schema.ts. - Pydantic write-DTOs <-> zod form schemas.
check_webapp_schema_mirror.pyextracts the canonical Pydantic shape intowebapp/frontend/schema-mirror.snapshot.json; the paired vitest test asserts each zod schema's.shapeagrees on field names, types, and min/max constraints.
check_ci_deps.py and check_constants_sync.py are stdlib-only;
check_openapi_snapshot.py and check_webapp_schema_mirror.py need
[webapp] deps installed.
# Single-experiment run on a standalone strategy YAML
make experiment ARGS="run --config config/strategies/adaptive_bollinger.yaml"
# HPO study
make experiment ARGS="tune --config config/strategies/adaptive_bollinger.yaml \
--hpo-config config/hpo/adaptive_bollinger.yaml"
# Cross-strategy comparison
make experiment ARGS="compare \
--config config/strategies/adaptive_bollinger.yaml \
--config config/strategies/momentum_gatekeeper.yaml \
--out demo_comparison"experiment.pyis a thin click wrapper oversrc/orchestration/builder.py,src/orchestration/comparison.py, andsrc/orchestration/holdout_eval.py.deploy.pyis a thin click wrapper oversrc/orchestration/deployment.py.- The Makefile binds these CLIs to
make experiment/make stubs.