nemo_gym.orchestration.executors.slurm

View as Markdown

Module Contents

Classes

NameDescription
SlurmExecutorSlurm executor for Pyxis-enabled clusters (https://github.com/NVIDIA/pyxis).

Functions

NameDescription
_parse_sbatch_resultsBenchmark name to (job_id, error), exactly one of which is set.
_sbatch_commandOne sbatch that reports its own benchmark, exit status and output.
_validate_benchmark_namesFail before anything is staged or copied.
_validate_mounts-

Data

_JOB_ID_RE

_MARKER

_VALID_BENCHMARK_NAME

API

class nemo_gym.orchestration.executors.slurm.SlurmExecutor()

Bases: BaseExecutor

Slurm executor for Pyxis-enabled clusters (https://github.com/NVIDIA/pyxis).

Every service and the driver are launched via srun --container-image so they run inside the container specified in their config. Health checks run as plain bash inside the sbatch script (no container needed — they just poll HTTP).

nemo_gym.orchestration.executors.slurm.SlurmExecutor._build_record(
cluster: str,
gym_job_id: str,
now: datetime.datetime,
remote_run_dir: pathlib.Path,
benchmark_names: list[str],
output: str
nemo_gym.orchestration.executors.slurm.SlurmExecutor._dry_run(
remote_run_dir: pathlib.Path
) -> None
nemo_gym.orchestration.executors.slurm.SlurmExecutor._stage(
remote_run_dir: pathlib.Path,
staging: pathlib.Path
) -> pathlib.Path
nemo_gym.orchestration.executors.slurm.SlurmExecutor.run(
dry_run: bool = False
nemo_gym.orchestration.executors.slurm._parse_sbatch_results(
output: str
) -> dict[str, tuple[str | None, str | None]]

Benchmark name to (job_id, error), exactly one of which is set.

nemo_gym.orchestration.executors.slurm._sbatch_command(
benchmark: str,
script: pathlib.Path
) -> str

One sbatch that reports its own benchmark, exit status and output.

rc is captured immediately: $? after the tr pipeline would be tr’s status, which is zero however badly sbatch failed. tr flattens sbatch’s multi-line error messages so the whole result stays on one marker line.

nemo_gym.orchestration.executors.slurm._validate_benchmark_names(
benchmarks: list[str]
) -> None

Fail before anything is staged or copied.

_sbatch_command raises this same check, but only once submission is already underway — after staging, connecting, mount validation and the rsync. Checking here first means a bad name fails before any of that, and fails on --dry-run too, instead of the dry run printing a clean script listing for a benchmark that could never actually be submitted.

nemo_gym.orchestration.executors.slurm._validate_mounts(
) -> None
nemo_gym.orchestration.executors.slurm._JOB_ID_RE = re.compile('^\\d+(;\\S+)?$')
nemo_gym.orchestration.executors.slurm._MARKER = '__GYM_JOB:'
nemo_gym.orchestration.executors.slurm._VALID_BENCHMARK_NAME = re.compile('^[A-Za-z0-9._-]+$')