nemo_gym.orchestration.executors.slurm
nemo_gym.orchestration.executors.slurm
Module Contents
Classes
Functions
Data
API
Bases: BaseExecutor
Slurm executor for Pyxis-enabled clusters (https://github.com/NVIDIA/pyxis).
Every service and the driver are launched via srun --container-image so they
run inside the container specified in their config. Health checks run as plain
bash inside the sbatch script (no container needed — they just poll HTTP).
Benchmark name to (job_id, error), exactly one of which is set.
One sbatch that reports its own benchmark, exit status and output.
rc is captured immediately: $? after the tr pipeline would be tr’s
status, which is zero however badly sbatch failed. tr flattens sbatch’s
multi-line error messages so the whole result stays on one marker line.
Fail before anything is staged or copied.
_sbatch_command raises this same check, but only once submission is
already underway — after staging, connecting, mount validation and the
rsync. Checking here first means a bad name fails before any of that, and
fails on --dry-run too, instead of the dry run printing a clean script
listing for a benchmark that could never actually be submitted.