nemo_curator.utils.vllm_utils

View as Markdown

Shared vLLM setup utilities.

These helpers centralise the boilerplate that every vLLM-based inference stage needs: finding a free port, initialising an vllm.LLM engine with automatic port-collision retry, and resolving an HuggingFace model ID to a local snapshot path.

They were extracted from the Nemotron-Parse inference stage, which was the first stage in NeMo Curator to be tested at scale (320x H100). Future stages that use vLLM (video, text, audio) should import from here rather than duplicating this logic. See GitHub issue #1720 for the roadmap to wire these utilities into other modalities.

Module Contents

Functions

NameDescription
_exception_chain_text-
_is_engine_startup_failureReturn True if exc looks like a retryable vLLM engine-startup failure.
create_vllm_llmCreate a vllm.LLM instance with automatic port-collision retry.
create_vllm_llm_with_retryCreate a vLLM engine with port-collision retries and no added defaults.
merge_vllm_kwargsMerge caller-owned vLLM arguments with user-provided engine options.
pick_free_portReturn a free TCP port on the local machine.
resolve_local_model_pathResolve an HF model ID to a local snapshot path.
validate_vllm_kwargsReject vLLM kwargs that override arguments owned by the caller.

Data

_ENGINE_STARTUP_FAILURE_MARKERS

_NON_RETRYABLE_MARKERS

API

nemo_curator.utils.vllm_utils._exception_chain_text(
exc: BaseException
) -> str
nemo_curator.utils.vllm_utils._is_engine_startup_failure(
exc: BaseException
) -> bool

Return True if exc looks like a retryable vLLM engine-startup failure.

nemo_curator.utils.vllm_utils.create_vllm_llm(
model_path: str,
max_num_seqs: int = 64,
enforce_eager: bool = False,
dtype: str = 'bfloat16',
trust_remote_code: bool = True,
limit_mm_per_prompt: dict | None = None,
max_port_retries: int = 3,
extra_engine_kwargs: object = {}
) -> 'vllm.LLM'

Create a vllm.LLM instance with automatic port-collision retry.

vLLM selects a MASTER_PORT for the distributed backend at startup. On a busy node the chosen port may already be in use, causing an EADDRINUSE RuntimeError. This helper picks a fresh free port on each attempt so that transient collisions are handled transparently.

Parameters:

model_path
str

Local path or HuggingFace model ID to load.

max_num_seqs
intDefaults to 64

Maximum number of sequences vLLM processes concurrently.

enforce_eager
boolDefaults to False

Disable CUDA graph capture (slower but uses less memory).

dtype
strDefaults to 'bfloat16'

Model weight dtype passed to vLLM (e.g. "bfloat16").

trust_remote_code
boolDefaults to True

Whether to trust remote code in the model repository.

limit_mm_per_prompt
dict | NoneDefaults to None

Multimodal token limits per prompt (e.g. {"image": 1}). Defaults to {"image": 1} when None.

max_port_retries
intDefaults to 3

Number of port-pick attempts before re-raising the error.

extra_engine_kwargs
objectDefaults to {}

Additional keyword arguments forwarded verbatim to vllm.LLM (e.g. gpu_memory_utilization, max_num_batched_tokens). Keys here override the explicit defaults above when they collide.

nemo_curator.utils.vllm_utils.create_vllm_llm_with_retry(
max_port_retries: int = 3,
engine_kwargs: object = {}
) -> 'vllm.LLM'

Create a vLLM engine with port-collision retries and no added defaults.

nemo_curator.utils.vllm_utils.merge_vllm_kwargs(
vllm_kwargs: collections.abc.Mapping[str, object],
owned_kwargs: collections.abc.Mapping[str, object],
owner_description: str
) -> dict[str, object]

Merge caller-owned vLLM arguments with user-provided engine options.

Callers describe the keyword arguments they own by passing the actual owned_kwargs mapping. This keeps collision handling generic while the owner-specific values remain next to the call that constructs the engine.

nemo_curator.utils.vllm_utils.pick_free_port() -> int

Return a free TCP port on the local machine.

nemo_curator.utils.vllm_utils.resolve_local_model_path(
model_path: str
) -> str

Resolve an HF model ID to a local snapshot path.

Uses local_files_only=True so that workers on compute nodes never attempt to reach the internet. The model must be pre-downloaded (e.g. via huggingface-cli download) before submitting the job.

Parameters:

model_path
str

HuggingFace model ID or an already-local path. If the path is already a local directory it is returned unchanged.

nemo_curator.utils.vllm_utils.validate_vllm_kwargs(
vllm_kwargs: collections.abc.Mapping[str, object],
reserved_keys: collections.abc.Collection[str],
owner_description: str
) -> None

Reject vLLM kwargs that override arguments owned by the caller.

nemo_curator.utils.vllm_utils._ENGINE_STARTUP_FAILURE_MARKERS = ('eaddrinuse', 'address already in use', 'engine core initialization failed', 'e...
nemo_curator.utils.vllm_utils._NON_RETRYABLE_MARKERS = ('out of memory', 'cudaerrormemoryallocation', 'device-side assert', 'invalid co...