> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/gym/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/gym/_mcp/server.

# nemo_gym.orchestration.api

## Module Contents

### Classes

| Name                                                                           | Description                                                                                 |
| ------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------- |
| [`BaseComputeConfig`](#nemo_gym-orchestration-api-BaseComputeConfig)           | -                                                                                           |
| [`BaseModelServiceConfig`](#nemo_gym-orchestration-api-BaseModelServiceConfig) | Base for services that serve a model and can be wired as the policy model.                  |
| [`BaseServiceConfig`](#nemo_gym-orchestration-api-BaseServiceConfig)           | -                                                                                           |
| [`BenchmarkRunConfig`](#nemo_gym-orchestration-api-BenchmarkRunConfig)         | -                                                                                           |
| [`DriverConfig`](#nemo_gym-orchestration-api-DriverConfig)                     | -                                                                                           |
| [`GymInstallConfig`](#nemo_gym-orchestration-api-GymInstallConfig)             | -                                                                                           |
| [`HealthCheckConfig`](#nemo_gym-orchestration-api-HealthCheckConfig)           | -                                                                                           |
| [`JobConfig`](#nemo_gym-orchestration-api-JobConfig)                           | -                                                                                           |
| [`NodePool`](#nemo_gym-orchestration-api-NodePool)                             | -                                                                                           |
| [`OtelConfig`](#nemo_gym-orchestration-api-OtelConfig)                         | An OpenTelemetry collector beside every benchmark job: scrapes each model service's         |
| [`RayServiceConfig`](#nemo_gym-orchestration-api-RayServiceConfig)             | -                                                                                           |
| [`ResumeConfig`](#nemo_gym-orchestration-api-ResumeConfig)                     | -                                                                                           |
| [`SlurmComputeConfig`](#nemo_gym-orchestration-api-SlurmComputeConfig)         | -                                                                                           |
| [`SubmitConfig`](#nemo_gym-orchestration-api-SubmitConfig)                     | -                                                                                           |
| [`VllmPDServiceConfig`](#nemo_gym-orchestration-api-VllmPDServiceConfig)       | Prefill/decode disaggregated vLLM: two tiers behind a vllm-router.                          |
| [`VllmPDTierConfig`](#nemo_gym-orchestration-api-VllmPDTierConfig)             | One tier of a `vllm_pd` service as deployed. Built by the service, never written by a user. |
| [`VllmServiceConfig`](#nemo_gym-orchestration-api-VllmServiceConfig)           | -                                                                                           |
| [`_StrictModel`](#nemo_gym-orchestration-api-_StrictModel)                     | -                                                                                           |

### Functions

| Name                                                                     | Description                                                                                      |
| ------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------ |
| [`effective_ray_serve`](#nemo_gym-orchestration-api-effective_ray_serve) | Whether the Ray Serve gateway manages this service's instances/routing instead of vLLM's own DP. |
| [`resolve_env_dict`](#nemo_gym-orchestration-api-resolve_env_dict)       | Resolve `lit:`/`host:`/`runtime:` prefixes on `env` values. Every value must use one             |

### Data

[`ComputeConfig`](#nemo_gym-orchestration-api-ComputeConfig)

[`RUNTIME_ENV_PREFIX`](#nemo_gym-orchestration-api-RUNTIME_ENV_PREFIX)

[`ServiceConfig`](#nemo_gym-orchestration-api-ServiceConfig)

[`_DECODE_NIXL_PORT_OFFSET`](#nemo_gym-orchestration-api-_DECODE_NIXL_PORT_OFFSET)

[`_ENV_VAR_NAME_RE`](#nemo_gym-orchestration-api-_ENV_VAR_NAME_RE)

[`_PD_SERVICE_ONLY_FIELDS`](#nemo_gym-orchestration-api-_PD_SERVICE_ONLY_FIELDS)

### API

```python
class nemo_gym.orchestration.api.BaseComputeConfig()
```

**Bases:** [\_StrictModel](#nemo_gym-orchestration-api-_StrictModel)

```python
class nemo_gym.orchestration.api.BaseModelServiceConfig()
```

**Bases:** [BaseServiceConfig](#nemo_gym-orchestration-api-BaseServiceConfig)

Base for services that serve a model and can be wired as the policy model.

**`model`** `str`

---

**`port`** `int = 8000`

---

**`served_model_name`** `str | None = None`

---

```python
class nemo_gym.orchestration.api.BaseServiceConfig()
```

**Bases:** [\_StrictModel](#nemo_gym-orchestration-api-_StrictModel)

**`container`** `str`

---

**`env`** `dict[str, str] = {}`

---

**`health_check`** `HealthCheckConfig | None = None`

---

**`mounts`** `list[str] = []`

---

**`node_pool`** `str | None = None`

---

**`placement`** `str | None = None`

---

**`pre_command`** `str = ''`

---

```python
nemo_gym.orchestration.api.BaseServiceConfig._resolve_env_prefixes(
    v: dict[str, str]
) -> dict[str, str]
```

classmethod

```python
class nemo_gym.orchestration.api.BenchmarkRunConfig()
```

**Bases:** [\_StrictModel](#nemo_gym-orchestration-api-_StrictModel)

**`command`** `str | None = None`

---

**`prepare`** `dict[str, Any] = {}`

---

**`resumable`** `bool | ResumeConfig = False`

---

**`resume_config`** `ResumeConfig | None`

---

**`run`** `dict[str, Any] = {}`

---

```python
nemo_gym.orchestration.api.BenchmarkRunConfig._validate_command() -> nemo_gym.orchestration.api.BenchmarkRunConfig
```

```python
class nemo_gym.orchestration.api.DriverConfig()
```

**Bases:** [\_StrictModel](#nemo_gym-orchestration-api-_StrictModel)

**`benchmarks`** `dict[str, BenchmarkRunConfig]`

---

**`container`** `str = 'python:3.12'`

---

**`env`** `dict[str, str] = {}`

---

**`gym_install`** `GymInstallConfig | None = None`

---

**`mounts`** `list[str] = []`

---

**`policy_host`** `str = 'localhost'`

---

**`policy_model`** `str | None = None`

---

**`policy_model_type`** `str = 'openai_model'`

---

```python
nemo_gym.orchestration.api.DriverConfig._resolve_env_prefixes(
    v: dict[str, str]
) -> dict[str, str]
```

classmethod

```python
class nemo_gym.orchestration.api.GymInstallConfig()
```

**Bases:** [\_StrictModel](#nemo_gym-orchestration-api-_StrictModel)

**`ref`** `str`

---

**`repo`** `str = 'https://github.com/NVIDIA-NeMo/gym'`

---

```python
class nemo_gym.orchestration.api.HealthCheckConfig()
```

**Bases:** [\_StrictModel](#nemo_gym-orchestration-api-_StrictModel)

**`path`** `str = '/health'`

---

**`port`** `int | None = None`

---

**`timeout_seconds`** `int = 60`

---

```python
class nemo_gym.orchestration.api.JobConfig()
```

**Bases:** [\_StrictModel](#nemo_gym-orchestration-api-_StrictModel)

**`output_path`** `str`

---

```python
class nemo_gym.orchestration.api.NodePool()
```

**Bases:** [\_StrictModel](#nemo_gym-orchestration-api-_StrictModel)

**`extra_args`** `dict[str, str] = {}`

---

**`gpus_per_node`** `int | None = None`

---

**`nodes`** `int = 1`

---

**`ntasks_per_node`** `int = 1`

---

**`partition`** `str`

---

```python
class nemo_gym.orchestration.api.OtelConfig()
```

**Bases:** [\_StrictModel](#nemo_gym-orchestration-api-_StrictModel)

An OpenTelemetry collector beside every benchmark job: scrapes each model service's
Prometheus `/metrics`, receives OTLP from the job's own processes on :4317/:4318, and ships
both to an OTLP/HTTP backend while keeping a copy under `&lt;job dir&gt;/otel/`. On by default, so
a run is observable unless it opts out; `endpoint` and `service_name` come from the
deployment's own config (a cluster fragment, typically) and are required while enabled.

**`binary`** `str = 'otelcol-contrib'`

---

**`component`** `str = 'gym-vllm'`

---

**`container`** `str | None = None`

---

**`enabled`** `bool = True`

---

**`endpoint`** `str | None = None`

---

**`gpu_metrics_port`** `int | None = 9400`

---

**`gym_logs`** `bool = True`

---

**`gym_span_groups`** `str = 'default,verify'`

---

**`gym_telemetry`** `bool = True`

---

**`health_check_timeout_seconds`** `int = 300`

---

**`node_metrics_port`** `int | None = 9100`

---

**`scrape_interval_seconds`** `int = 15`

---

**`service_name`** `str | None = None`

---

**`token_env`** `str = 'OTEL_TOKEN'`

---

```python
nemo_gym.orchestration.api.OtelConfig._validate_positive(
    v: int
) -> int
```

classmethod

```python
nemo_gym.orchestration.api.OtelConfig._validate_token_env(
    v: str
) -> str
```

classmethod

```python
class nemo_gym.orchestration.api.RayServiceConfig()
```

**Bases:** [BaseServiceConfig](#nemo_gym-orchestration-api-BaseServiceConfig)

**`extra_args`** `str = ''`

---

**`node_pools`** `list[str] = []`

---

**`num_cpus`** `int | None = None`

---

**`num_gpus`** `int | None = None`

---

**`port`** `int = 6379`

---

**`resources`** `dict[str, dict[str, float]] = {}`

---

**`type`** `Literal['ray']`

---

```python
nemo_gym.orchestration.api.RayServiceConfig._validate_shape() -> nemo_gym.orchestration.api.RayServiceConfig
```

```python
class nemo_gym.orchestration.api.ResumeConfig()
```

**Bases:** [\_StrictModel](#nemo_gym-orchestration-api-_StrictModel)

**`max_retries`** `int = 3`

---

**`max_walltime`** `str | None = None`

---

```python
class nemo_gym.orchestration.api.SlurmComputeConfig()
```

**Bases:** [BaseComputeConfig](#nemo_gym-orchestration-api-BaseComputeConfig)

**`account`** `str`

---

**`extra_args`** `dict[str, str] = {}`

---

**`hostname`** `str | None = None`

---

**`node_pools`** `dict[str, NodePool] = {}`

---

**`type`** `Literal['slurm']`

---

**`walltime`** `str | None = None`

---

```python
class nemo_gym.orchestration.api.SubmitConfig()
```

**Bases:** [\_StrictModel](#nemo_gym-orchestration-api-_StrictModel)

**`compute`** `dict[str, ComputeConfig]`

---

**`deployed_services`** `dict[str, ServiceConfig]`

`services` with each vllm\_pd service expanded into its two tiers and its router.

---

**`driver`** `DriverConfig`

---

**`job`** `JobConfig`

---

**`otel`** `OtelConfig = OtelConfig()`

---

**`services`** `dict[str, ServiceConfig]`

---

```python
nemo_gym.orchestration.api.SubmitConfig._resolve_and_validate_placements() -> nemo_gym.orchestration.api.SubmitConfig
```

```python
nemo_gym.orchestration.api.SubmitConfig._validate_pd_services(
    compute: nemo_gym.orchestration.api.ComputeConfig
) -> None
```

```python
nemo_gym.orchestration.api.SubmitConfig._validate_vllm_gpu_footprint(
    service_name: str,
    service: nemo_gym.orchestration.api.VllmServiceConfig,
    total_nodes: int,
    node_pools: dict[str, nemo_gym.orchestration.api.NodePool],
    gpus_per_node_values: list[int],
    is_ray_serve: bool
) -> None
```

```python
class nemo_gym.orchestration.api.VllmPDServiceConfig()
```

**Bases:** [BaseModelServiceConfig](#nemo_gym-orchestration-api-BaseModelServiceConfig)

Prefill/decode disaggregated vLLM: two tiers behind a vllm-router.

Deployed as three services named `&lt;name&gt;-prefill`, `&lt;name&gt;-decode` and `&lt;name&gt;`
(the router); `driver.policy_model` names the router.

**`_tiers`** `dict[str, VllmPDTierConfig] = PrivateAttr(default_factory=dict)`

---

**`decode`** `VllmServiceConfig`

---

**`decode_policy`** `str = 'cache_aware'`

---

**`intra_node_data_parallel_size`** `int = 1`

---

**`kv_connector`** `str = 'NixlConnector'`

---

**`kv_load_failure_policy`** `str = 'fail'`

---

**`log_level`** `str = 'error'`

---

**`nixl_side_channel_port`** `int = 5600`

---

**`prefill`** `VllmServiceConfig`

---

**`prefill_policy`** `str = 'cache_aware'`

---

**`request_timeout_secs`** `int = 86400`

---

**`server_per_node`** `bool = False`

---

**`type`** `Literal['vllm_pd']`

---

```python
nemo_gym.orchestration.api.VllmPDServiceConfig._build_tiers() -> nemo_gym.orchestration.api.VllmPDServiceConfig
```

```python
nemo_gym.orchestration.api.VllmPDServiceConfig._fill_tier_defaults(
    data: typing.Any
) -> typing.Any
```

classmethod

```python
nemo_gym.orchestration.api.VllmPDServiceConfig.tiers(
    name: str
) -> dict[str, nemo_gym.orchestration.api.VllmPDTierConfig]
```

```python
class nemo_gym.orchestration.api.VllmPDTierConfig()
```

**Bases:** [VllmServiceConfig](#nemo_gym-orchestration-api-VllmServiceConfig)

One tier of a `vllm_pd` service as deployed. Built by the service, never written by a user.

**`kv_transfer_config`** `dict[str, str]`

---

**`nixl_side_channel_port`** `int`

---

**`server_per_node`** `bool`

---

```python
class nemo_gym.orchestration.api.VllmServiceConfig()
```

**Bases:** [BaseModelServiceConfig](#nemo_gym-orchestration-api-BaseModelServiceConfig)

**`data_parallel_rpc_port`** `int = 13345`

---

**`extra_args`** `str = ''`

---

**`number_of_instances`** `int = 1`

---

**`pipeline_parallel_size`** `int = 1`

---

**`tensor_parallel_size`** `int = 1`

---

**`trust_remote_code`** `bool = False`

---

**`type`** `Literal['vllm']`

---

**`use_ray_serve`** `bool = False`

---

```python
nemo_gym.orchestration.api.VllmServiceConfig._default_health_check() -> nemo_gym.orchestration.api.VllmServiceConfig
```

```python
nemo_gym.orchestration.api.VllmServiceConfig._validate_number_of_instances(
    v: int
) -> int
```

classmethod

```python
class nemo_gym.orchestration.api._StrictModel()
```

**Bases:** `BaseModel`

**`model_config`** `= ConfigDict(extra='forbid')`

---

```python
nemo_gym.orchestration.api.effective_ray_serve(
    service: nemo_gym.orchestration.api.VllmServiceConfig,
    total_nodes: int,
    gpus_per_node_values: list[int]
) -> bool
```

Whether the Ray Serve gateway manages this service's instances/routing instead of vLLM's own DP.

```python
nemo_gym.orchestration.api.resolve_env_dict(
    env: dict[str, str]
) -> dict[str, str]
```

Resolve `lit:`/`host:`/`runtime:` prefixes on `env` values. Every value must use one
of these prefixes; a missing or misspelled prefix raises rather than being guessed at.

* `lit:VALUE` -> literal VALUE.
* `host:VAR` -> read from os.environ\[VAR] on the machine running `gym eval submit`;
  raises if VAR isn't set there.
* `runtime:VAR` -> left unresolved; canonicalized to `runtime:VAR` for executors to
  pick up and reference from the job's own environment at run time.

```python
nemo_gym.orchestration.api.ComputeConfig = Annotated[Annotated[SlurmComputeConfig, Tag('slurm')], Discriminator('type')]
```

```python
nemo_gym.orchestration.api.RUNTIME_ENV_PREFIX = 'runtime:'
```

```python
nemo_gym.orchestration.api.ServiceConfig = Annotated[Annotated[VllmServiceConfig, Tag('vllm')] | Annotated[RayServiceConfig...
```

```python
nemo_gym.orchestration.api._DECODE_NIXL_PORT_OFFSET = 1
```

```python
nemo_gym.orchestration.api._ENV_VAR_NAME_RE = re.compile('^[A-Za-z_][A-Za-z0-9_]*$')
```

```python
nemo_gym.orchestration.api._PD_SERVICE_ONLY_FIELDS = ('server_per_node', 'nixl_side_channel_port')
```