> This page is for version 26.07 · v1.3.0.
> For other versions, use one of these documentation indexes:
> - Latest · v1.4.0 (26.09) (default): https://docs.nvidia.com/nemo/curator/latest/llms.txt
> - Main · preview: https://docs.nvidia.com/nemo/curator/main/llms.txt
> - 26.09 · v1.4.0: https://docs.nvidia.com/nemo/curator/v26.09/llms.txt
> - 26.07 · v1.3.0: https://docs.nvidia.com/nemo/curator/v26.07/llms.txt
> - 26.04 · v1.2.0: https://docs.nvidia.com/nemo/curator/v26.04/llms.txt
> - 26.02 · v1.1.0: https://docs.nvidia.com/nemo/curator/v26.02/llms.txt
> - 25.09 · v1.0.0: https://docs.nvidia.com/nemo/curator/v25.09/llms.txt

> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/curator/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/curator/_mcp/server.

# nemo_curator.utils.vllm_utils

Shared vLLM setup utilities.

These helpers centralise the boilerplate that every vLLM-based inference stage
needs: finding a free port, initialising an `vllm.LLM` engine with
automatic port-collision retry, and resolving an HuggingFace model ID to a
local snapshot path.

They were extracted from the Nemotron-Parse inference stage, which was the
first stage in NeMo Curator to be tested at scale (320x H100).  Future stages
that use vLLM (video, text, audio) should import from here rather than
duplicating this logic.  See GitHub issue #1720 for the roadmap to wire these
utilities into other modalities.

## Module Contents

### Functions

| Name                                                                                      | Description                                                              |
| ----------------------------------------------------------------------------------------- | ------------------------------------------------------------------------ |
| [`_exception_chain_text`](#nemo_curator-utils-vllm_utils-_exception_chain_text)           | -                                                                        |
| [`_is_engine_startup_failure`](#nemo_curator-utils-vllm_utils-_is_engine_startup_failure) | Return True if `exc` looks like a retryable vLLM engine-startup failure. |
| [`create_vllm_llm`](#nemo_curator-utils-vllm_utils-create_vllm_llm)                       | Create a `vllm.LLM` instance with automatic port-collision retry.        |
| [`create_vllm_llm_with_retry`](#nemo_curator-utils-vllm_utils-create_vllm_llm_with_retry) | Create a vLLM engine with port-collision retries and no added defaults.  |
| [`merge_vllm_kwargs`](#nemo_curator-utils-vllm_utils-merge_vllm_kwargs)                   | Merge caller-owned vLLM arguments with user-provided engine options.     |
| [`pick_free_port`](#nemo_curator-utils-vllm_utils-pick_free_port)                         | Return a free TCP port on the local machine.                             |
| [`resolve_local_model_path`](#nemo_curator-utils-vllm_utils-resolve_local_model_path)     | Resolve an HF model ID to a local snapshot path.                         |
| [`validate_vllm_kwargs`](#nemo_curator-utils-vllm_utils-validate_vllm_kwargs)             | Reject vLLM kwargs that override arguments owned by the caller.          |

### Data

[`_ENGINE_STARTUP_FAILURE_MARKERS`](#nemo_curator-utils-vllm_utils-_ENGINE_STARTUP_FAILURE_MARKERS)

[`_NON_RETRYABLE_MARKERS`](#nemo_curator-utils-vllm_utils-_NON_RETRYABLE_MARKERS)

### API

```python
nemo_curator.utils.vllm_utils._exception_chain_text(
    exc: BaseException
) -> str
```

```python
nemo_curator.utils.vllm_utils._is_engine_startup_failure(
    exc: BaseException
) -> bool
```

Return True if `exc` looks like a retryable vLLM engine-startup failure.

```python
nemo_curator.utils.vllm_utils.create_vllm_llm(
    model_path: str,
    max_num_seqs: int = 64,
    enforce_eager: bool = False,
    dtype: str = 'bfloat16',
    trust_remote_code: bool = True,
    limit_mm_per_prompt: dict | None = None,
    max_port_retries: int = 3,
    extra_engine_kwargs: object = {}
) -> 'vllm.LLM'
```

Create a `vllm.LLM` instance with automatic port-collision retry.

vLLM selects a MASTER\_PORT for the distributed backend at startup.  On a
busy node the chosen port may already be in use, causing an
`EADDRINUSE` `RuntimeError`.  This helper picks a fresh free port on
each attempt so that transient collisions are handled transparently.

**Parameters:**

**`model_path`** `str`

Local path or HuggingFace model ID to load.

---

**`max_num_seqs`** `int` — default: 64

Maximum number of sequences vLLM processes concurrently.

---

**`enforce_eager`** `bool` — default: False

Disable CUDA graph capture (slower but uses less memory).

---

**`dtype`** `str` — default: 'bfloat16'

Model weight dtype passed to vLLM (e.g. `"bfloat16"`).

---

**`trust_remote_code`** `bool` — default: True

Whether to trust remote code in the model repository.

---

**`limit_mm_per_prompt`** `dict | None` — default: None

Multimodal token limits per prompt (e.g. `&#123;"image": 1&#125;`).
Defaults to `&#123;"image": 1&#125;` when `None`.

---

**`max_port_retries`** `int` — default: 3

Number of port-pick attempts before re-raising the error.

---

**`extra_engine_kwargs`** `object` — default: \{}

Additional keyword arguments forwarded verbatim to `vllm.LLM`
(e.g. `gpu_memory_utilization`, `max_num_batched_tokens`). Keys here
override the explicit defaults above when they collide.

---

```python
nemo_curator.utils.vllm_utils.create_vllm_llm_with_retry(
    max_port_retries: int = 3,
    engine_kwargs: object = {}
) -> 'vllm.LLM'
```

Create a vLLM engine with port-collision retries and no added defaults.

```python
nemo_curator.utils.vllm_utils.merge_vllm_kwargs(
    vllm_kwargs: collections.abc.Mapping[str, object],
    owned_kwargs: collections.abc.Mapping[str, object],
    owner_description: str
) -> dict[str, object]
```

Merge caller-owned vLLM arguments with user-provided engine options.

Callers describe the keyword arguments they own by passing the actual
`owned_kwargs` mapping. This keeps collision handling generic while the
owner-specific values remain next to the call that constructs the engine.

```python
nemo_curator.utils.vllm_utils.pick_free_port() -> int
```

Return a free TCP port on the local machine.

```python
nemo_curator.utils.vllm_utils.resolve_local_model_path(
    model_path: str
) -> str
```

Resolve an HF model ID to a local snapshot path.

Uses `local_files_only=True` so that workers on compute nodes never
attempt to reach the internet.  The model must be pre-downloaded (e.g.
via `huggingface-cli download`) before submitting the job.

**Parameters:**

**`model_path`** `str`

HuggingFace model ID or an already-local path.  If the path is
already a local directory it is returned unchanged.

---

```python
nemo_curator.utils.vllm_utils.validate_vllm_kwargs(
    vllm_kwargs: collections.abc.Mapping[str, object],
    reserved_keys: collections.abc.Collection[str],
    owner_description: str
) -> None
```

Reject vLLM kwargs that override arguments owned by the caller.

```python
nemo_curator.utils.vllm_utils._ENGINE_STARTUP_FAILURE_MARKERS = ('eaddrinuse', 'address already in use', 'engine core initialization failed', 'e...
```

```python
nemo_curator.utils.vllm_utils._NON_RETRYABLE_MARKERS = ('out of memory', 'cudaerrormemoryallocation', 'device-side assert', 'invalid co...
```