Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
76 changes: 76 additions & 0 deletions doc/api.rst
Original file line number Diff line number Diff line change
@@ -0,0 +1,76 @@
Python API
==========

Use ``cloudai.api`` to validate scenarios, run experiments, and read their results
from Python. The CLI and API share the same execution logic.

Run an experiment
-----------------

Pass configuration files as ``pathlib.Path`` objects:

.. code-block:: python

from pathlib import Path

import cloudai.api

scenario = Path("conf/common/test_scenario/sleep.toml")
system = Path("conf/common/system/example_slurm_cluster.toml")
tests_dir = Path("conf/common/test")

valid, errors = cloudai.api.validate_scenario(scenario, system, tests_dir=tests_dir)
if not valid:
raise ValueError(errors)

experiment = cloudai.api.run_experiment(scenario, system, tests_dir=tests_dir)
print(experiment.status, experiment.path)

Test ``path`` references resolve relative to the scenario file. Other relative
paths keep their usual meaning relative to the working directory.

``run_experiment`` installs workload prerequisites as needed and returns an
``Experiment`` model containing the saved results. Each invocation gets its own
result directory. Use ``experiment.model_dump()`` or ``experiment.model_dump_json()``
to serialize the result.

Optional keyword arguments:

* ``tests_dir`` and ``hook_dir`` select test definitions and hooks.
* ``output_dir`` overrides the system's results directory.
* ``dry_run=True`` generates workload commands without submitting jobs.
* ``single_sbatch=True`` runs a Slurm scenario in one allocation.
* ``enable_cache_without_check=True`` uses installed workload components without
checking them first, as with the CLI option.

Find and read results
---------------------

``list_experiments(system)`` returns ``(experiment_id, result_directory)`` pairs
from the system's ``output_path``, sorted by directory name. Directories without
``experiment.json`` are skipped. A missing results directory returns an empty list.
When using ``output_dir`` for a run, use that directory as ``output_path`` in the
system configuration passed to ``list_experiments``.

``get_experiment(exp)`` accepts a result directory or the ``experiment.json`` file
itself, as either ``str`` or ``Path``. It reads the saved snapshot without querying
the scheduler. See :doc:`reporting` for the result fields.

Errors
------

``validate_scenario`` returns ``(True, {})`` for valid inputs, or ``(False, errors)``
with ``system`` or ``scenario`` keys and error messages. It checks configuration and
execution constraints without installing workloads, creating a results directory,
or querying cluster availability.

``run_experiment`` raises exceptions for invalid configuration and setup errors.
Workload failures appear in the experiment and run statuses.

Missing result files raise ``FileNotFoundError``. Invalid experiment JSON raises
``pydantic.ValidationError``. Listing also raises for unreadable or invalid result
files rather than silently omitting them.

The API does not install signal handlers or invoke the CLI's logging configuration.
The legacy public functions in ``cloudai.cli.handlers`` remain available with
deprecation warnings; use ``cloudai.api`` for new integrations.
2 changes: 2 additions & 0 deletions doc/index.rst
Original file line number Diff line number Diff line change
Expand Up @@ -14,6 +14,7 @@ This document contains the following chapters:
Tutorial
USER_GUIDE
reporting
api
systems
workloads/index
DEV
Expand All @@ -22,6 +23,7 @@ This document contains the following chapters:
- :doc:`Tutorial`
- :doc:`USER_GUIDE`
- :doc:`reporting`
- :doc:`api`
- :doc:`systems`
- :doc:`workloads/index`
- :doc:`DEV`
Expand Down
11 changes: 9 additions & 2 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -153,14 +153,20 @@ root_package = "cloudai"
"cloudai.util",
"cloudai.cli",
"cloudai.handlers",
"cloudai.api",
"cloudai.report_generator",
]

[[tool.importlinter.contracts]]
name = "Report generator is leaf dependency"
type = "forbidden"
forbidden_modules = ["cloudai.workloads", "cloudai.cli", "cloudai.handlers"]
allow_indirect_imports = true # allow "from cloudai.core import ..."
forbidden_modules = [
"cloudai.workloads",
"cloudai.cli",
"cloudai.handlers",
"cloudai.api",
]
allow_indirect_imports = true # allow "from cloudai.core import ..."
source_modules = ["cloudai.report_generator"]

[[tool.importlinter.contracts]]
Expand All @@ -171,6 +177,7 @@ root_package = "cloudai"
"cloudai.workloads",
"cloudai.cli",
"cloudai.handlers",
"cloudai.api",
]
allow_indirect_imports = true
source_modules = ["cloudai.util"]
Expand Down
2 changes: 1 addition & 1 deletion src/cloudai/_core/base_runner.py
Original file line number Diff line number Diff line change
Expand Up @@ -161,7 +161,7 @@ def submit_test(self, tr: TestRun):
self.update_run_output(job)
except JobSubmissionError as e:
logging.error(e)
exit(1)
raise

def on_job_submit(self, tr: TestRun) -> None:
return
Expand Down
36 changes: 30 additions & 6 deletions src/cloudai/_core/runner.py
Original file line number Diff line number Diff line change
Expand Up @@ -16,6 +16,7 @@

import datetime
import logging
from pathlib import Path
from types import FrameType
from typing import Optional

Expand Down Expand Up @@ -44,18 +45,41 @@ class Runner:
test_scenario (TestScenario): The test scenario to be executed.
"""

def __init__(self, mode: str, system: System, test_scenario: TestScenario):
def __init__(
self,
mode: str,
system: System,
test_scenario: TestScenario,
*,
runner_class: type[BaseRunner] | None = None,
output_path: Path | None = None,
):
logging.info(f"Initializing Runner [{mode.upper()}] mode")
self.runner = self.create_runner(mode, system, test_scenario)

def create_runner(self, mode: str, system: System, test_scenario: TestScenario) -> BaseRunner:
if runner_class is None and output_path is None:
self.runner = self.create_runner(mode, system, test_scenario)
else:
self.runner = self.create_runner(
mode, system, test_scenario, runner_class=runner_class, output_path=output_path
)

def create_runner(
self,
mode: str,
system: System,
test_scenario: TestScenario,
*,
runner_class: type[BaseRunner] | None = None,
output_path: Path | None = None,
) -> BaseRunner:
"""
Dynamically create a runner instance based on the system's scheduler type.

Args:
mode (str): The operation mode ('dry-run', 'run').
system (System): The system configuration.
test_scenario (TestScenario): The test scenario to run.
runner_class: Runner override for this invocation.
output_path: Exact scenario result directory, when already allocated.

Returns:
BaseRunner: A runner instance suitable for the system.
Expand All @@ -69,13 +93,13 @@ def create_runner(self, mode: str, system: System, test_scenario: TestScenario)
msg = f"No runner registered for scheduler: {scheduler_type}"
logging.error(msg)
raise NotImplementedError(msg)
runner_class = registry.runners_map[scheduler_type]
runner_class = runner_class or registry.runners_map[scheduler_type]
logging.info(f"Creating {runner_class.__name__}")

if not system.output_path.exists():
system.output_path.mkdir()
current_time = datetime.datetime.now().strftime("%Y-%m-%d_%H-%M-%S")
results_root = system.output_path / f"{test_scenario.name}_{current_time}"
results_root = output_path or system.output_path / f"{test_scenario.name}_{current_time}"

return runner_class(mode, system, test_scenario, results_root)

Expand Down
96 changes: 96 additions & 0 deletions src/cloudai/api.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,96 @@
# SPDX-FileCopyrightText: NVIDIA CORPORATION & AFFILIATES
# Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.

from pathlib import Path

import cloudai.handlers
import cloudai.models.output


def run_experiment(
scenario: Path,
system: Path,
*,
tests_dir: Path | None = None,
hook_dir: Path | None = None,
output_dir: Path | None = None,
dry_run: bool = False,
single_sbatch: bool = False,
enable_cache_without_check: bool = False,
) -> cloudai.models.output.Experiment:
"""
Run a scenario and return its experiment snapshot.

Pass configuration files as Path. Relative test paths resolve relative to
the scenario file. Optional directories have the same meaning as in the CLI.
The call waits for execution to finish.

Configuration and setup errors raise exceptions. Workload failures are
recorded in the returned experiment's status. This function does not invoke
CLI logging configuration or install signal handlers.
"""
return cloudai.handlers.run_experiment(
scenario,
system,
tests_dir=tests_dir,
hook_dir=hook_dir,
output_dir=output_dir,
dry_run=dry_run,
single_sbatch=single_sbatch,
enable_cache_without_check=enable_cache_without_check,
)


def list_experiments(system: Path) -> list[tuple[str, Path]]:
"""
Return experiment IDs and local result directories under system.output_path.

Pass the system configuration file as Path. A missing results
directory returns an empty list. Unreadable or invalid experiment files raise
exceptions rather than returning an incomplete catalog.
"""
return cloudai.handlers.list_experiments(system)


def validate_scenario(
scenario: Path,
system: Path,
*,
tests_dir: Path | None = None,
hook_dir: Path | None = None,
single_sbatch: bool = False,
) -> tuple[bool, dict[str, str]]:
"""
Validate configuration files without installing or running jobs.

Return (True, {}) on success, otherwise (False, errors), with configuration
names as keys and error messages as values. This checks configuration and
execution constraints, not whether cluster resources are currently available.
"""
return cloudai.handlers.validate_scenario(
scenario, system, tests_dir=tests_dir, hook_dir=hook_dir, single_sbatch=single_sbatch
)


def get_experiment(exp: str | Path) -> cloudai.models.output.Experiment:
"""
Read experiment.json from a result directory or a JSON file path.

Both str and Path represent filesystem paths here. Missing files raise
FileNotFoundError; invalid JSON or schema raises pydantic.ValidationError.
This reads the saved snapshot without querying the scheduler.
"""
return cloudai.handlers.get_experiment(exp)
Loading
Loading