Curate TextSynthetic Data

Inference Server

View as Markdown

InferenceServer serves one or more local models behind an OpenAI-compatible endpoint on your Ray cluster. Choose a typed configuration for Ray Serve or NVIDIA Dynamo. Both backends support this guide’s serving capabilities and use the same client endpoint and lifecycle API. Benchmark your workload to select a backend.

The backend configurations are:

BackendModel ConfigurationServer Configuration
Ray ServeRayServeModelConfigRayServeServerConfig (default)
NVIDIA DynamoDynamoVLLMModelConfigDynamoServerConfig

Ray Serve uses vLLM from the base inference_server environment. Dynamo uses Ray to create an actor environment with the matching public Dynamo release and its vLLM extra, so its first startup takes longer while Ray resolves and installs the environment. Ray Serve deployments can also be accessed through Ray Serve handles instead of the OpenAI-compatible endpoint; benchmark the approach for your workloads.

The old InferenceModelConfig class was removed. Migrate to the backend-specific types in Migrate from InferenceModelConfig.

Prerequisites

Install the inference_server extra with the public CUDA 12.9 PyTorch and vLLM indexes:

uv pip install \
--torch-backend cu129 \
--extra-index-url https://wheels.vllm.ai/0.22.0/cu129 \
"nemo-curator[inference_server]"

The command keeps PyPI as the default index. The additional indexes provide CUDA 12.9 builds of PyTorch and vLLM. Other dependencies use PyPI.

This extra supports local GPU serving on x86_64 Linux and aarch64 Linux. It does not install vLLM, ai-dynamo, or NIXL on macOS. The current repository Dockerfile installs etcd and nats-server for Dynamo. For a source environment or other image, run the Dynamo dependency installation script or configure existing endpoints with DynamoServerConfig. Both services must be reachable from every Ray node that can run Dynamo actors:

from nemo_curator.core.serve import DynamoServerConfig
DynamoServerConfig(
etcd_endpoint="http://etcd.example.com:2379",
nats_url="nats://nats.example.com:4222",
)

The current dependency stack uses:

  • CUDA 12.9.1 in images built with the repository Dockerfile’s defaults.
  • PyTorch 2.11.
  • Ray Serve 2.57.0 or later.
  • vLLM 0.22.0 with CUDA 12.9.
  • A public ai-dynamo release 1.3.1 or later and NIXL 0.10.0 or later for the Dynamo backend.

Start Ray or connect to an existing cluster before you create the server. The examples use RayClient; you can also use an externally managed cluster.

Ray Serve Quickstart

Ray Serve is the default backend. Omitting backend= is equivalent to passing RayServeServerConfig().

from openai import OpenAI
from nemo_curator.core.client import RayClient
from nemo_curator.core.serve import InferenceServer, RayServeModelConfig
ray_client = RayClient(num_cpus=8, num_gpus=1)
ray_client.start()
model = RayServeModelConfig(
model_identifier="HuggingFaceTB/SmolLM2-135M-Instruct", # pragma: allowlist secret
deployment_config={
"autoscaling_config": {
"min_replicas": 1,
"max_replicas": 1,
},
},
engine_kwargs={
"tensor_parallel_size": 1,
"max_model_len": 2048,
},
)
with InferenceServer(models=[model]) as server:
client = OpenAI(base_url=server.endpoint, api_key="unused")
response = client.chat.completions.create(
model="HuggingFaceTB/SmolLM2-135M-Instruct", # pragma: allowlist secret
messages=[{"role": "user", "content": "Say hello in one word."}],
max_tokens=16,
)
print(response.choices[0].message.content)

The server waits for every configured model to appear at /v1/models before start() returns. server.endpoint contains the correct host and port for the selected backend.

Shared Server API

InferenceServer accepts a list of model configurations and one matching server configuration.

ParameterTypeDefaultDescription
modelslist[BaseModelConfig]RequiredModels to serve. Every item must use the same concrete configuration type.
backendBaseServerConfigRayServeServerConfig()Backend and server-level configuration.
namestr"default"Server name used for Ray Serve applications or Dynamo actor and placement-group prefixes.
portint8000Preferred frontend port. NeMo Curator selects a free port when necessary.
health_check_timeout_sint300Maximum time to wait for all models to register.
verboseboolFalsePreserve detailed backend and request logs when True.

Only one InferenceServer can be active in a Python process at a time. Stop the current server before starting another.

Use a context manager for automatic cleanup:

with InferenceServer(models=[model]) as server:
print(server.endpoint)

For explicit lifecycle control:

server = InferenceServer(models=[model])
server.start()
try:
print(server.endpoint)
finally:
server.stop()

Ray Serve Configuration

RayServeModelConfig

ParameterTypeDefaultDescription
model_identifierstrRequiredHugging Face model ID or local model path.
model_namestr | NoneNoneName exposed through the API. Defaults to model_identifier.
runtime_envdict{}Ray runtime environment merged into the model deployment.
deployment_configdict{}Ray Serve deployment settings, including autoscaling.
engine_kwargsdict{}vLLM engine settings such as tensor parallelism and model length.

Use model_name when weights come from a local path but clients should use a stable API name:

model = RayServeModelConfig(
model_identifier="/models/gemma-3-27b-it",
model_name="google/gemma-3-27b-it",
deployment_config={
"autoscaling_config": {"min_replicas": 1, "max_replicas": 2},
},
engine_kwargs={"tensor_parallel_size": 4},
)

Multiple Ray Serve Models

Each model can use a distinct deployment and runtime environment:

from nemo_curator.core.serve import InferenceServer, RayServeModelConfig
models = [
RayServeModelConfig(
model_identifier="HuggingFaceTB/SmolLM2-135M-Instruct", # pragma: allowlist secret
model_name="writer",
deployment_config={
"autoscaling_config": {"min_replicas": 1, "max_replicas": 2},
},
engine_kwargs={"tensor_parallel_size": 1},
),
RayServeModelConfig(
model_identifier="HuggingFaceTB/SmolLM-135M-Instruct", # pragma: allowlist secret
model_name="reviewer",
deployment_config={
"autoscaling_config": {"min_replicas": 1, "max_replicas": 1},
},
engine_kwargs={"tensor_parallel_size": 1},
),
]
server = InferenceServer(models=models)
server.start()

Clients select writer or reviewer in the OpenAI request’s model field.

Dynamo Aggregated Serving

Aggregated mode runs prefill and decode in the same vLLM worker. num_replicas creates static replicas; Dynamo does not use Ray Serve autoscaling configuration.

from nemo_curator.core.serve import (
DynamoServerConfig,
DynamoVLLMModelConfig,
InferenceServer,
)
model = DynamoVLLMModelConfig(
model_identifier="HuggingFaceTB/SmolLM2-135M-Instruct", # pragma: allowlist secret
mode="aggregated",
num_replicas=2,
engine_kwargs={
"tensor_parallel_size": 1,
"max_model_len": 2048,
},
)
server = InferenceServer(
models=[model],
backend=DynamoServerConfig(),
health_check_timeout_s=600,
)
server.start()

For aggregated models, a tensor-parallel replica can span multiple nodes when the tensor-parallel size divides evenly across available nodes. NeMo Curator prefers a single-node placement and otherwise uses equal GPU bundles on distinct nodes.

Dynamo Disaggregated Serving

Disaggregated mode runs independent prefill and decode workers. Configure both roles explicitly.

from nemo_curator.core.serve import (
DynamoRoleConfig,
DynamoServerConfig,
DynamoVLLMModelConfig,
InferenceServer,
)
model = DynamoVLLMModelConfig(
model_identifier="HuggingFaceTB/SmolLM2-135M-Instruct", # pragma: allowlist secret
mode="disagg",
engine_kwargs={
"max_model_len": 2048,
"tensor_parallel_size": 1,
},
prefill=DynamoRoleConfig(
num_replicas=2,
engine_kwargs={"tensor_parallel_size": 2},
),
decode=DynamoRoleConfig(
num_replicas=1,
engine_kwargs={"tensor_parallel_size": 1},
),
)
server = InferenceServer(
models=[model],
backend=DynamoServerConfig(),
health_check_timeout_s=600,
)
server.start()

Role-level engine_kwargs shallow-merge over the model-level values. In this example, prefill uses tensor parallelism of 2 and decode uses 1; both inherit max_model_len.

Each disaggregated role’s tensor-parallel group must fit on a single node. Multi-node tensor parallelism is supported for aggregated replicas, not for a disaggregated prefill or decode worker.

Disaggregated serving uses the NIXL connector for KV transfer by default. kv_transfer_config and kv_events_config are managed by NeMo Curator and are not constructor parameters.

Dynamo Configuration Reference

DynamoVLLMModelConfig

ParameterTypeDefaultDescription
model_identifierstrRequiredHugging Face model ID or local model path.
model_namestr | NoneNoneAPI-facing name; defaults to model_identifier.
runtime_envdict{}Packages, environment variables, and other Ray runtime settings for workers.
engine_kwargsdict{}Base vLLM engine settings.
num_replicasint1Static replica count for aggregated mode. Must be at least 1.
mode"aggregated" | "disagg""aggregated"Serving topology.
prefillDynamoRoleConfig | NoneNonePrefill replicas and overrides for disaggregated mode.
decodeDynamoRoleConfig | NoneNoneDecode replicas and overrides for disaggregated mode.
dynamo_kwargsdict{}Additional worker CLI options, translated from snake case to kebab case.

All models in one server must have unique model_name values. Dynamo also rejects names that sanitize to the same component slug.

DynamoRoleConfig

ParameterTypeDefaultDescription
num_replicasint1Number of workers for this role. Must be at least 0; DynamoVLLMModelConfig requires at least one prefill and one decode replica in disaggregated mode.
engine_kwargsdict{}Role-level vLLM settings merged over the model settings.

DynamoServerConfig

ParameterTypeDefaultDescription
etcd_endpointstr | NoneNoneExisting etcd endpoint. When omitted, NeMo Curator starts etcd.
nats_urlstr | NoneNoneExisting NATS endpoint. When omitted, NeMo Curator starts NATS.
namespacestrDynamo defaultDynamo discovery namespace.
request_planestrDynamo defaultRequest transport, for example "tcp".
event_planestrDynamo defaultEvent transport.
routerDynamoRouterConfigDefault configRouting mode and frontend options.
subprocess_envdict[str, str]{}Environment variables propagated to etcd/NATS-aware workers and the frontend.

NeMo Curator’s own resolved ETCD_ENDPOINTS and NATS_SERVER values always take precedence over any same-named keys in subprocess_env, so setting them there to point workers at a different etcd/NATS instance is silently ignored. Use etcd_endpoint and nats_url instead.

Dynamo Routing

from nemo_curator.core.serve import DynamoRouterConfig, DynamoServerConfig
backend = DynamoServerConfig(
router=DynamoRouterConfig(
mode="kv",
kv_events=False,
),
)
modeBehavior
NoneAuto-select KV routing when any model is disaggregated; otherwise let Dynamo use round-robin.
"round_robin"Rotate requests across replicas.
"random"Select a replica randomly.
"kv"Route using KV-cache affinity.
"direct"Use Dynamo’s direct-routing mode.

kv_events is valid only with KV routing. When you explicitly set mode="kv", kv_events=False uses approximate tree-based tracking. When mode=None auto-selects KV routing for a disaggregated model, NeMo Curator also enables exact event-backed routing even though the configured default is False. If a publishing role explicitly enables vLLM’s hybrid KV-cache manager, automatic routing instead keeps events disabled; explicitly requesting event-backed routing with that configuration raises an error.

Additional router_kwargs are forwarded to the Dynamo frontend. Do not put router_mode or router_kv_events in that dictionary; use the typed fields instead. Boolean false values are emitted as --no-* flags.

Multimodal workers need their model-specific settings, such as limit_mm_per_prompt in engine_kwargs and enable_multimodal in dynamo_kwargs. Use Dynamo’s default frontend processing path.

The default frontend path avoids the slower compatibility processor used by older examples.

Runtime Environments and Subprocess Variables

Both model types inherit runtime_env from BaseModelConfig:

model = DynamoVLLMModelConfig(
model_identifier="my-org/my-model",
runtime_env={
"uv": {"packages": ["my-model-plugin==1.2.0"]},
"env_vars": {"HF_HOME": "/models/hf-cache"},
},
)

NeMo Curator merges pip or uv package lists and environment variables instead of replacing the backend’s required packages. Dynamo automatically adds its tested ai-dynamo[vllm] actor environment and pins the actor’s Ray version to the cluster version. On a new cluster node, the first actor can take several minutes to create this cached environment.

Use DynamoServerConfig.subprocess_env for variables that must reach the Dynamo frontend and worker subprocesses:

backend = DynamoServerConfig(
subprocess_env={"DYN_TCP_REQUEST_TIMEOUT": "180"},
)

Choosing runtime_env vs. subprocess_env

A Dynamo model’s actor has its own isolated Python virtualenv, cloned from the environment that launched the driver process and then installed on top of additively (see Per-Stage Runtime Environments for this Ray behavior generally). Inside that actor, Dynamo then launches a worker subprocess that runs vLLM. DynamoVLLMModelConfig.runtime_env is per-model — it doesn’t reach other models’ worker actors. DynamoServerConfig.subprocess_env is server-wide — it reaches every configured model’s worker plus the frontend. Use the mechanism that matches both what you’re changing and who needs it:

You need toUseWhy
Install or pin a Python package before the worker starts (a plugin, a newer library version)runtime_env on the model configOnly runtime_env builds the isolated actor virtualenv; a package that isn’t installed there can’t be imported
Set an env var only one model’s worker needs (an engine feature flag, a model-specific cache path)runtime_env["env_vars"] on that model configDoesn’t reach other models’ workers; setting it on subprocess_env instead would leak it onto every one of them
Set a value every model’s worker (and the frontend) should read at startup (a transport timeout, a shared cache directory)subprocess_env on DynamoServerConfigApplied server-wide, without touching any virtualenv
Make a local module importable everywhere without installing it as a package (for example, a compatibility shim all workers need)PYTHONPATH in subprocess_envPYTHONPATH extends sys.path at process start; it doesn’t require a package install, so it belongs with the other server-wide subprocess settings

A model’s runtime_env is not isolated from the shared Dynamo frontend actor: NeMo Curator merges every configured model’s runtime_env into the frontend’s virtualenv build. env_vars merge cleanly (on a conflicting key, the last model in your models list wins on the frontend), but uv/pip package lists are concatenated, not reconciled — if two models pin incompatible versions of the same package, both requirements are installed together and can fail to resolve, which would prevent the frontend (and the whole server) from starting. Keep model-specific package pins compatible across every model in the same server, or move genuinely conflicting models to separate InferenceServer instances.

Do not export an installer-relevant or import-relevant variable in the shell that starts your driver process. The driver’s shell environment does not propagate into the isolated actor virtualenv, and if the variable does reach Ray itself rather than only the worker subprocess, it can cause Ray to import something unexpected and stall cluster or actor startup. Scope the variable with runtime_env or subprocess_env instead.

backend = DynamoServerConfig(
subprocess_env={"PYTHONPATH": "/abs/path/to/compat-shim-dir"},
)

Resource Placement

Before starting Dynamo, NeMo Curator checks the total requested GPUs and fails early when the cluster is too small.

  • Aggregated GPU count is num_replicas * tensor_parallel_size per model.
  • Disaggregated GPU count is the sum of each role’s num_replicas * tensor_parallel_size.
  • Aggregated tensor-parallel groups prefer one node, then use an equal multi-node split.
  • Each disaggregated role must fit on one node.
  • Set CURATOR_IGNORE_RAY_HEAD_NODE=1 to keep model placement off the Ray head node when worker nodes are labeled for that policy.

Dynamo creates detached, named placement groups and actors. On startup, it removes stale resources with the same server-name prefix. On shutdown, it reacquires actor handles by name, terminates the subprocess groups, and removes the placement groups.

HAProxy Ingress

Ray Serve’s Linux dependencies include the bundled ray-haproxy binary. Local Ray cluster initialization enables HAProxy ingress when that package or an explicit RAY_SERVE_HAPROXY_BINARY_PATH is available and assigns free metrics and statistics ports. Otherwise, Ray Serve uses its default Python proxy.

No additional proxy package installation or HAProxy configuration is required for InferenceServer, which enables HAProxy when starting Ray Serve. To confirm that a local cluster discovered the packaged binary, check startup logs for Ray Serve HAProxy ingress enabled.

Use with NeMo Curator Clients

Point any OpenAI-compatible client at server.endpoint:

from nemo_curator.models.client.openai_client import AsyncOpenAIClient
client = AsyncOpenAIClient(
base_url=server.endpoint,
api_key="unused",
max_concurrent_requests=10,
)

For Dynamo, the endpoint host is the node running the frontend placement-group bundle. Use server.endpoint instead of assuming localhost.

Run Pipelines While Serving

Ray Data and Ray Actor Pool executors honor Ray’s GPU accounting and schedule GPU stages away from GPUs held by either inference backend. Xenna manages GPU assignment independently and is rejected when an active InferenceServer and a GPU pipeline stage would conflict.

from nemo_curator.backends.ray_data import RayDataExecutor
with InferenceServer(models=[model]) as server:
results = pipeline.run(executor=RayDataExecutor())

CPU-only pipelines can use any executor while the server is active.

For Nemotron-Parse PDF pipelines, use the Dynamo backend with NemotronParseHTTPClientStage. Start with one model replica per GPU, a fixed HTTP stage pool of 4 * num_gpus workers, and 32 concurrent requests per HTTP worker. This starting point was validated on 8 H100 GPUs. Because request concurrency is hardware- and corpus-dependent, benchmark 64 on the target workload and keep it only if throughput improves without request failures or out-of-memory errors. The Nemotron-Parse PDF guide provides the exact install and run commands.

Migrate from InferenceModelConfig

Before:

from nemo_curator.core.serve import InferenceModelConfig, InferenceServer
model = InferenceModelConfig(
model_identifier="google/gemma-3-27b-it",
deployment_config={
"autoscaling_config": {"min_replicas": 1, "max_replicas": 1},
},
engine_kwargs={"tensor_parallel_size": 4},
)
server = InferenceServer(models=[model])

After, using Ray Serve:

from nemo_curator.core.serve import InferenceServer, RayServeModelConfig
model = RayServeModelConfig(
model_identifier="google/gemma-3-27b-it",
deployment_config={
"autoscaling_config": {"min_replicas": 1, "max_replicas": 1},
},
engine_kwargs={"tensor_parallel_size": 4},
)
server = InferenceServer(models=[model])

Or, using static Dynamo replicas:

from nemo_curator.core.serve import (
DynamoServerConfig,
DynamoVLLMModelConfig,
InferenceServer,
)
model = DynamoVLLMModelConfig(
model_identifier="google/gemma-3-27b-it",
num_replicas=1,
engine_kwargs={"tensor_parallel_size": 4},
)
server = InferenceServer(models=[model], backend=DynamoServerConfig())

Do not pass Ray Serve’s deployment_config to DynamoVLLMModelConfig. Use num_replicas or the disaggregated role counts instead.

Troubleshooting

Model Does Not Become Healthy

  • Increase health_check_timeout_s for large model downloads or slow actor-environment creation.
  • Check the model name returned by /v1/models; local paths often need an explicit model_name alias.
  • Confirm the requested tensor-parallel groups fit the cluster topology.
  • For Dynamo, inspect subprocess logs under the Ray session’s nemo_curator_dynamo_<id> directory.
  • If the whole server (including other models) fails to start, check whether two models’ runtime_env package pins conflict — see Choosing runtime_env vs. subprocess_env.

Dynamo Cannot Start etcd or NATS

For setup options, refer to Prerequisites. When using existing services, set etcd_endpoint and nats_url on DynamoServerConfig; both endpoints must be reachable from all Ray nodes that may run Dynamo actors.

Local Multi-GPU Startup Hangs

On PCIe systems without peer-to-peer GPU access, restart Ray or the Python kernel and either set NCCL_P2P_DISABLE=1 before starting the cluster or reduce tensor_parallel_size to 1.

Runtime Environment Times Out

Confirm that cluster nodes can reach the package index and use compatible Python, CUDA, and Ray versions. Dynamo actor setup times out after 600 seconds. If runtime installation is too slow or unavailable, build a custom image with the required additional packages or prepare another environment with those dependencies.

General Debugging Tips for Dynamo Startup and Compatibility Issues

  • Identify which environment failed. An error during actor creation, before any worker subprocess log appears, points to a runtime_env / virtualenv problem. An error inside the worker’s own logs, after the actor already exists, points to something in the already-built virtualenv or subprocess_env. See Choosing runtime_env vs. subprocess_env.
  • Remember that installing packages via runtime_env can disturb a pin already cloned into the actor virtualenv. The actor virtualenv starts as a clone of the driver’s, but any package runtime_env installs on top resolves its own dependencies — which can silently pull in a different version of something the driver already had pinned (its Ray version, an excluded incompatible build). Supply that pin or exclusion again through the model’s runtime_env rather than assuming the clone protects it.
  • Rule out GPU memory contention before suspecting a compatibility issue. A startup failure that looks like a version mismatch can also come from a competing process leaving too little free GPU memory for vLLM’s requested gpu_memory_utilization. Check with nvidia-smi first.
  • Reproduce with the smallest possible configuration: one model, one replica, and enforce_eager in engine_kwargs if graph capture is a suspect. Confirm a fix there before scaling back up to your full topology.
  • Smoke-test after any runtime_env, subprocess_env, model, or engine-kwarg change, with one replica and one request, before trusting a full run. A model appearing at /v1/models confirms registration, not that generation itself will succeed.

Example: A CUTLASS or QuACK Version-Mismatch Error

vLLM JIT-compiles some fused/quantized kernels using CUTLASS (NVIDIA’s GPU kernel template library, via its Python cutlass.cute interface) and, for some kernels, QuACK, which is built on that same interface. If the installed CUTLASS/QuACK build doesn’t match the rest of the stack, the failure usually doesn’t look like a version error — it shows up at kernel warm-up as a plain AttributeError naming a missing cutlass.cute.core symbol, for example:

AttributeError: module 'cutlass.cute.core' has no attribute 'ThrMma'

This is easy to mistake for a model or vLLM bug, since nothing in the message mentions CUDA or CUTLASS versions. If you hit something like this:

  • Check the CUDA tag on any newly added or newly resolved package against the cu129 baseline that NeMo Curator’s Dynamo stack targets (vllm==0.22.0+cu129). A stray cu13-tagged wheel is one checkable cause.
  • If the tags all check out, the installed CUTLASS/QuACK build may simply be too old for your GPU architecture — try a newer ai-dynamo[vllm] pin.

Next Steps