nemo_gym.orchestration.executors.slurm_script
nemo_gym.orchestration.executors.slurm_script
Module Contents
Functions
Data
_RAY_SERVE_GATEWAY_SOURCE_PATH
API
ray start for one node of a ray service: the head, or with worker a node joining it.
address_var names the env var holding the head’s host:port (see ray_head_address_var).
The node the driver runs on in a multi-node job: the policy’s first node.
The driver reaches the policy on localhost. A multi-node policy serves its API from node 0, and a pinned one from its pool’s first node.
Escape text for safe embedding inside a double-quoted bash string (”…”).
Where this service answers its health probe: the node it runs on, as _service_nodelist places it.
Whether the script declares the allocation’s host array to place services with.
The one value every pool agrees on for attribute.
A plain sbatch job takes a single —partition/—ntasks-per-node/—gpus-per-node for the whole allocation, so pools that disagree cannot both be honoured. Slurm would silently apply whichever directive came last; say so instead.
Each pool’s (first node index, node count) within the allocation.
Pools are laid out contiguously in declaration order, which is the order the single #SBATCH —nodes total is built from, so pool i owns the nodes after every pool before it.
How the batch script, which runs on the allocation’s first node, reaches node.
Each spanned pool’s worker nodes as (first index, count): every node but the head.
The head is the driver’s node, which is always the first node of some pool.
The collector’s srun step. Started before the model services so the scrape covers their startup.
Runs on exactly one node, the batch host. That is the first node of the allocation, which is
also where the Ray prelude puts the head of a multi-node vLLM service and where the driver runs,
so localhost:<port> reaches the API server and the OTLP endpoints from the same node. In a
container the job directory is mounted for the config and the local otel/*.jsonl output; on
the node it is simply there, so no mounts or workdir are passed.
Run after the driver: one more scrape interval so the final counters are seen, then a graceful stop, keeping the driver’s exit code.
The TERM goes to the collector process itself, matched by its unique --config path. Sent to
srun instead, TERM makes Slurm kill the step outright and INT is treated as a console
interrupt; neither reaches the collector, so its final batch would be lost.
One set of directives for the whole allocation, not one per pool.
Pools divide an allocation between services (see _pool_offsets); they are not separate Slurm requests. Emitting —nodes per pool made every pool but the last a no-op, so a two-pool job asked for one pool’s nodes while the rest of the executor sized itself on the sum.
Declare the allocation’s hosts, and export each node pool’s, for —nodelist.
Export every ray service’s head address, and its worker hosts, from the nodes Slurm gave it.
The head runs beside the driver. Workers and the driver join it without knowing, when the config is written, which host the job will land on.
A ray service’s srun steps: the head beside the driver, then one step per spanned pool’s workers.
Every step uses the service’s one container, env and mounts.
Return an ‘env K=V …’ prefix string (trailing space) scoped to a single command, or ” if empty.
A runtime:VAR value (see resolve_env_dict in api.py) is emitted as an unquoted K=$VAR
shell reference instead of a literal, so it’s resolved from the job’s own environment when
the command actually runs on the compute node, rather than baked in at script-generation time.
What a service’s —nodelist names, as the shell variable it expands.
A pinned service names its pool. In a multi-node job an unpinned service that fits on one node joins the driver, which reaches it on localhost; otherwise Slurm could start it on any node.
How many nodes this service actually runs on.
A service pinned to a pool sees only that pool, so a single-node pool inside a ten-node job is a single-node deployment and must not be built as a multi-node Ray one.
Enable capture for submitted evaluations unless explicitly disabled.
Auto-derive model_call_capture_dir from this benchmark’s own real output directory when observability is on and the caller didn’t set one.
Hydra interpolation resolves before remote_bench_dir exists (it’s computed here, in build_sbatch_script, well after SubmitConfig validation), so there’s no way for a YAML value to reference it — this has to happen in Python, once the real path is known. An explicit model_call_capture_dir in run always wins over this default.
The env var build_sbatch_script exports with a node pool’s comma-separated hosts.
The env var build_sbatch_script exports with a ray service’s head host:port.
The env var build_sbatch_script exports with a ray service’s worker hosts in one pool.