> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/gym/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/gym/_mcp/server.

# nemo_gym.orchestration.executors.slurm

## Module Contents

### Classes

| Name                                                                     | Description                                                                                                     |
| ------------------------------------------------------------------------ | --------------------------------------------------------------------------------------------------------------- |
| [`SlurmExecutor`](#nemo_gym-orchestration-executors-slurm-SlurmExecutor) | Slurm executor for Pyxis-enabled clusters ([https://github.com/NVIDIA/pyxis](https://github.com/NVIDIA/pyxis)). |

### Functions

| Name                                                                                             | Description                                                          |
| ------------------------------------------------------------------------------------------------ | -------------------------------------------------------------------- |
| [`_parse_sbatch_results`](#nemo_gym-orchestration-executors-slurm-_parse_sbatch_results)         | Benchmark name to `(job_id, error)`, exactly one of which is set.    |
| [`_sbatch_command`](#nemo_gym-orchestration-executors-slurm-_sbatch_command)                     | One `sbatch` that reports its own benchmark, exit status and output. |
| [`_validate_benchmark_names`](#nemo_gym-orchestration-executors-slurm-_validate_benchmark_names) | Fail before anything is staged or copied.                            |
| [`_validate_mounts`](#nemo_gym-orchestration-executors-slurm-_validate_mounts)                   | -                                                                    |

### Data

[`_JOB_ID_RE`](#nemo_gym-orchestration-executors-slurm-_JOB_ID_RE)

[`_MARKER`](#nemo_gym-orchestration-executors-slurm-_MARKER)

[`_VALID_BENCHMARK_NAME`](#nemo_gym-orchestration-executors-slurm-_VALID_BENCHMARK_NAME)

### API

```python
class nemo_gym.orchestration.executors.slurm.SlurmExecutor()
```

**Bases:** [BaseExecutor](/nemo/gym/nemo-gym/nemo_gym/orchestration/executors/base#nemo_gym-orchestration-executors-base-BaseExecutor)

Slurm executor for Pyxis-enabled clusters ([https://github.com/NVIDIA/pyxis](https://github.com/NVIDIA/pyxis)).

Every service and the driver are launched via `srun --container-image` so they
run inside the container specified in their config. Health checks run as plain
bash inside the sbatch script (no container needed — they just poll HTTP).

```python
nemo_gym.orchestration.executors.slurm.SlurmExecutor._build_record(
    cluster: str,
    compute: nemo_gym.orchestration.api.SlurmComputeConfig,
    gym_job_id: str,
    now: datetime.datetime,
    remote_run_dir: pathlib.Path,
    benchmark_names: list[str],
    output: str
) -> nemo_gym.orchestration.jobs.SubmissionRecord
```

```python
nemo_gym.orchestration.executors.slurm.SlurmExecutor._dry_run(
    config: nemo_gym.orchestration.api.SubmitConfig,
    compute: nemo_gym.orchestration.api.SlurmComputeConfig,
    remote_run_dir: pathlib.Path
) -> None
```

```python
nemo_gym.orchestration.executors.slurm.SlurmExecutor._stage(
    config: nemo_gym.orchestration.api.SubmitConfig,
    compute: nemo_gym.orchestration.api.SlurmComputeConfig,
    remote_run_dir: pathlib.Path,
    staging: pathlib.Path
) -> pathlib.Path
```

```python
nemo_gym.orchestration.executors.slurm.SlurmExecutor.run(
    config: nemo_gym.orchestration.api.SubmitConfig,
    dry_run: bool = False
) -> nemo_gym.orchestration.jobs.SubmissionRecord | None
```

```python
nemo_gym.orchestration.executors.slurm._parse_sbatch_results(
    output: str
) -> dict[str, tuple[str | None, str | None]]
```

Benchmark name to `(job_id, error)`, exactly one of which is set.

```python
nemo_gym.orchestration.executors.slurm._sbatch_command(
    benchmark: str,
    script: pathlib.Path
) -> str
```

One `sbatch` that reports its own benchmark, exit status and output.

`rc` is captured immediately: `$?` after the `tr` pipeline would be *tr's*
status, which is zero however badly sbatch failed. `tr` flattens sbatch's
multi-line error messages so the whole result stays on one marker line.

```python
nemo_gym.orchestration.executors.slurm._validate_benchmark_names(
    benchmarks: list[str]
) -> None
```

Fail before anything is staged or copied.

`_sbatch_command` raises this same check, but only once submission is
already underway — after staging, connecting, mount validation and the
rsync. Checking here first means a bad name fails before any of that, and
fails on `--dry-run` too, instead of the dry run printing a clean script
listing for a benchmark that could never actually be submitted.

```python
nemo_gym.orchestration.executors.slurm._validate_mounts(
    config: nemo_gym.orchestration.api.SubmitConfig,
    conn: nemo_gym.orchestration.executors.connection.Connection
) -> None
```

```python
nemo_gym.orchestration.executors.slurm._JOB_ID_RE = re.compile('^\\d+(;\\S+)?$')
```

```python
nemo_gym.orchestration.executors.slurm._MARKER = '__GYM_JOB:'
```

```python
nemo_gym.orchestration.executors.slurm._VALID_BENCHMARK_NAME = re.compile('^[A-Za-z0-9._-]+$')
```