Inference Server
InferenceServer serves one or more local models behind an OpenAI-compatible endpoint on your Ray cluster. Choose a typed configuration for Ray Serve or NVIDIA Dynamo. Both backends support this guide’s serving capabilities and use the same client endpoint and lifecycle API. Benchmark your workload to select a backend.
The backend configurations are:
Ray Serve uses vLLM from the base inference_server environment. Dynamo uses
Ray to create an actor environment with the matching public Dynamo release and
its vLLM extra, so its first startup takes longer while Ray resolves and
installs the environment. Ray Serve deployments can also be accessed through
Ray Serve handles instead of the OpenAI-compatible endpoint; benchmark the
approach for your workloads.
The old InferenceModelConfig class was removed. Migrate to the backend-specific types in Migrate from InferenceModelConfig.
Prerequisites
Install the inference_server extra with the public CUDA 12.9 PyTorch and vLLM indexes:
The command keeps PyPI as the default index. The additional indexes provide CUDA 12.9 builds of PyTorch and vLLM. Other dependencies use PyPI.
This extra supports local GPU serving on x86_64 Linux and aarch64 Linux. It does
not install vLLM, ai-dynamo, or NIXL on macOS. The current repository Dockerfile
installs etcd and nats-server for Dynamo. For a source environment or other
image, run the Dynamo dependency installation script
or configure existing endpoints with DynamoServerConfig. Both services must
be reachable from every Ray node that can run Dynamo actors:
The current dependency stack uses:
- CUDA 12.9.1 in images built with the repository Dockerfile’s defaults.
- PyTorch 2.11.
- Ray Serve 2.57.0 or later.
- vLLM 0.22.0 with CUDA 12.9.
- A public ai-dynamo release 1.3.1 or later and NIXL 0.10.0 or later for the Dynamo backend.
Start Ray or connect to an existing cluster before you create the server. The examples use RayClient; you can also use an externally managed cluster.
Ray Serve Quickstart
Ray Serve is the default backend. Omitting backend= is equivalent to passing RayServeServerConfig().
The server waits for every configured model to appear at /v1/models before start() returns. server.endpoint contains the correct host and port for the selected backend.
Shared Server API
InferenceServer accepts a list of model configurations and one matching server configuration.
Only one InferenceServer can be active in a Python process at a time. Stop the current server before starting another.
Use a context manager for automatic cleanup:
For explicit lifecycle control:
Ray Serve Configuration
RayServeModelConfig
Use model_name when weights come from a local path but clients should use a stable API name:
Multiple Ray Serve Models
Each model can use a distinct deployment and runtime environment:
Clients select writer or reviewer in the OpenAI request’s model field.
Dynamo Aggregated Serving
Aggregated mode runs prefill and decode in the same vLLM worker. num_replicas creates static replicas; Dynamo does not use Ray Serve autoscaling configuration.
For aggregated models, a tensor-parallel replica can span multiple nodes when the tensor-parallel size divides evenly across available nodes. NeMo Curator prefers a single-node placement and otherwise uses equal GPU bundles on distinct nodes.
Dynamo Disaggregated Serving
Disaggregated mode runs independent prefill and decode workers. Configure both roles explicitly.
Role-level engine_kwargs shallow-merge over the model-level values. In this example, prefill uses tensor parallelism of 2 and decode uses 1; both inherit max_model_len.
Each disaggregated role’s tensor-parallel group must fit on a single node. Multi-node tensor parallelism is supported for aggregated replicas, not for a disaggregated prefill or decode worker.
Disaggregated serving uses the NIXL connector for KV transfer by default. kv_transfer_config and kv_events_config are managed by NeMo Curator and are not constructor parameters.
Dynamo Configuration Reference
DynamoVLLMModelConfig
All models in one server must have unique model_name values. Dynamo also rejects names that sanitize to the same component slug.
DynamoRoleConfig
DynamoServerConfig
NeMo Curator’s own resolved ETCD_ENDPOINTS and NATS_SERVER values always take precedence over any same-named keys in subprocess_env, so setting them there to point workers at a different etcd/NATS instance is silently ignored. Use etcd_endpoint and nats_url instead.
Dynamo Routing
kv_events is valid only with KV routing. When you explicitly set mode="kv", kv_events=False uses approximate tree-based tracking. When mode=None auto-selects KV routing for a disaggregated model, NeMo Curator also enables exact event-backed routing even though the configured default is False. If a publishing role explicitly enables vLLM’s hybrid KV-cache manager, automatic routing instead keeps events disabled; explicitly requesting event-backed routing with that configuration raises an error.
Additional router_kwargs are forwarded to the Dynamo frontend. Do not put router_mode or router_kv_events in that dictionary; use the typed fields instead. Boolean false values are emitted as --no-* flags.
Multimodal workers need their model-specific settings, such as
limit_mm_per_prompt in engine_kwargs and enable_multimodal in
dynamo_kwargs. Use Dynamo’s default frontend processing path.
The default frontend path avoids the slower compatibility processor used by older examples.
Runtime Environments and Subprocess Variables
Both model types inherit runtime_env from BaseModelConfig:
NeMo Curator merges pip or uv package lists and environment variables instead of replacing the backend’s required packages. Dynamo automatically adds its tested ai-dynamo[vllm] actor environment and pins the actor’s Ray version to the cluster version. On a new cluster node, the first actor can take several minutes to create this cached environment.
Use DynamoServerConfig.subprocess_env for variables that must reach the Dynamo frontend and worker subprocesses:
Choosing runtime_env vs. subprocess_env
A Dynamo model’s actor has its own isolated Python virtualenv, cloned from the environment that launched the driver process and then installed on top of additively (see Per-Stage Runtime Environments for this Ray behavior generally). Inside that actor, Dynamo then launches a worker subprocess that runs vLLM. DynamoVLLMModelConfig.runtime_env is per-model — it doesn’t reach other models’ worker actors. DynamoServerConfig.subprocess_env is server-wide — it reaches every configured model’s worker plus the frontend. Use the mechanism that matches both what you’re changing and who needs it:
A model’s runtime_env is not isolated from the shared Dynamo frontend actor: NeMo Curator merges every configured model’s runtime_env into the frontend’s virtualenv build. env_vars merge cleanly (on a conflicting key, the last model in your models list wins on the frontend), but uv/pip package lists are concatenated, not reconciled — if two models pin incompatible versions of the same package, both requirements are installed together and can fail to resolve, which would prevent the frontend (and the whole server) from starting. Keep model-specific package pins compatible across every model in the same server, or move genuinely conflicting models to separate InferenceServer instances.
Do not export an installer-relevant or import-relevant variable in the shell that starts your driver process. The driver’s shell environment does not propagate into the isolated actor virtualenv, and if the variable does reach Ray itself rather than only the worker subprocess, it can cause Ray to import something unexpected and stall cluster or actor startup. Scope the variable with runtime_env or subprocess_env instead.
Resource Placement
Before starting Dynamo, NeMo Curator checks the total requested GPUs and fails early when the cluster is too small.
- Aggregated GPU count is
num_replicas * tensor_parallel_sizeper model. - Disaggregated GPU count is the sum of each role’s
num_replicas * tensor_parallel_size. - Aggregated tensor-parallel groups prefer one node, then use an equal multi-node split.
- Each disaggregated role must fit on one node.
- Set
CURATOR_IGNORE_RAY_HEAD_NODE=1to keep model placement off the Ray head node when worker nodes are labeled for that policy.
Dynamo creates detached, named placement groups and actors. On startup, it removes stale resources with the same server-name prefix. On shutdown, it reacquires actor handles by name, terminates the subprocess groups, and removes the placement groups.
HAProxy Ingress
Ray Serve’s Linux dependencies include the bundled ray-haproxy binary. Local Ray cluster initialization enables HAProxy ingress when that package or an explicit RAY_SERVE_HAPROXY_BINARY_PATH is available and assigns free metrics and statistics ports. Otherwise, Ray Serve uses its default Python proxy.
No additional proxy package installation or HAProxy configuration is required for InferenceServer, which enables HAProxy when starting Ray Serve. To confirm that a local cluster discovered the packaged binary, check startup logs for Ray Serve HAProxy ingress enabled.
Use with NeMo Curator Clients
Point any OpenAI-compatible client at server.endpoint:
For Dynamo, the endpoint host is the node running the frontend placement-group bundle. Use server.endpoint instead of assuming localhost.
Run Pipelines While Serving
Ray Data and Ray Actor Pool executors honor Ray’s GPU accounting and schedule GPU stages away from GPUs held by either inference backend. Xenna manages GPU assignment independently and is rejected when an active InferenceServer and a GPU pipeline stage would conflict.
CPU-only pipelines can use any executor while the server is active.
For Nemotron-Parse PDF pipelines, use the Dynamo backend with
NemotronParseHTTPClientStage. Start with one model replica per GPU, a fixed
HTTP stage pool of 4 * num_gpus workers, and 32 concurrent requests per HTTP
worker. This starting point was validated on 8 H100 GPUs. Because request
concurrency is hardware- and corpus-dependent, benchmark 64 on the target
workload and keep it only if throughput improves without request failures or
out-of-memory errors. The
Nemotron-Parse PDF guide provides
the exact install and run commands.
Migrate from InferenceModelConfig
Before:
After, using Ray Serve:
Or, using static Dynamo replicas:
Do not pass Ray Serve’s deployment_config to DynamoVLLMModelConfig. Use num_replicas or the disaggregated role counts instead.
Troubleshooting
Model Does Not Become Healthy
- Increase
health_check_timeout_sfor large model downloads or slow actor-environment creation. - Check the model name returned by
/v1/models; local paths often need an explicitmodel_namealias. - Confirm the requested tensor-parallel groups fit the cluster topology.
- For Dynamo, inspect subprocess logs under the Ray session’s
nemo_curator_dynamo_<id>directory. - If the whole server (including other models) fails to start, check whether two models’
runtime_envpackage pins conflict — see Choosingruntime_envvs.subprocess_env.
Dynamo Cannot Start etcd or NATS
For setup options, refer to Prerequisites. When using
existing services, set etcd_endpoint and nats_url on DynamoServerConfig;
both endpoints must be reachable from all Ray nodes that may run Dynamo actors.
Local Multi-GPU Startup Hangs
On PCIe systems without peer-to-peer GPU access, restart Ray or the Python kernel and either set NCCL_P2P_DISABLE=1 before starting the cluster or reduce tensor_parallel_size to 1.
Runtime Environment Times Out
Confirm that cluster nodes can reach the package index and use compatible Python, CUDA, and Ray versions. Dynamo actor setup times out after 600 seconds. If runtime installation is too slow or unavailable, build a custom image with the required additional packages or prepare another environment with those dependencies.
General Debugging Tips for Dynamo Startup and Compatibility Issues
- Identify which environment failed. An error during actor creation, before any worker subprocess log appears, points to a
runtime_env/ virtualenv problem. An error inside the worker’s own logs, after the actor already exists, points to something in the already-built virtualenv orsubprocess_env. See Choosingruntime_envvs.subprocess_env. - Remember that installing packages via
runtime_envcan disturb a pin already cloned into the actor virtualenv. The actor virtualenv starts as a clone of the driver’s, but any packageruntime_envinstalls on top resolves its own dependencies — which can silently pull in a different version of something the driver already had pinned (its Ray version, an excluded incompatible build). Supply that pin or exclusion again through the model’sruntime_envrather than assuming the clone protects it. - Rule out GPU memory contention before suspecting a compatibility issue. A startup failure that looks like a version mismatch can also come from a competing process leaving too little free GPU memory for vLLM’s requested
gpu_memory_utilization. Check withnvidia-smifirst. - Reproduce with the smallest possible configuration: one model, one replica, and
enforce_eagerinengine_kwargsif graph capture is a suspect. Confirm a fix there before scaling back up to your full topology. - Smoke-test after any
runtime_env,subprocess_env, model, or engine-kwarg change, with one replica and one request, before trusting a full run. A model appearing at/v1/modelsconfirms registration, not that generation itself will succeed.
Example: A CUTLASS or QuACK Version-Mismatch Error
vLLM JIT-compiles some fused/quantized kernels using CUTLASS (NVIDIA’s GPU kernel template library, via its Python cutlass.cute interface) and, for some kernels, QuACK, which is built on that same interface. If the installed CUTLASS/QuACK build doesn’t match the rest of the stack, the failure usually doesn’t look like a version error — it shows up at kernel warm-up as a plain AttributeError naming a missing cutlass.cute.core symbol, for example:
This is easy to mistake for a model or vLLM bug, since nothing in the message mentions CUDA or CUTLASS versions. If you hit something like this:
- Check the CUDA tag on any newly added or newly resolved package against the
cu129baseline that NeMo Curator’s Dynamo stack targets (vllm==0.22.0+cu129). A straycu13-tagged wheel is one checkable cause. - If the tags all check out, the installed CUTLASS/QuACK build may simply be too old for your GPU architecture — try a newer
ai-dynamo[vllm]pin.