nemo_curator.core.serve
nemo_curator.core.serve
Subpackages
Submodules
nemo_curator.core.serve.basenemo_curator.core.serve.constantsnemo_curator.core.serve.placementnemo_curator.core.serve.servernemo_curator.core.serve.subprocess_mgr
Package Contents
Classes
Functions
API
Base public model config shared by inference backends.
Merge two runtime_env dicts while preserving package lists.
Base server-level config; subclasses declare which model config types they accept.
Per-role config for disaggregated Dynamo serving.
Frontend router config for Dynamo.
mode=None means “auto”: Curator picks "kv" if any model uses
mode="disagg", else leaves --router-mode unset so the Dynamo
frontend falls back to its own round_robin default. kv_events
only applies when mode == "kv": pass kv_events=True to opt into
exact ZMQ KV-cache event publishing; the default uses the router’s
approximate tree-based tracking. Anything else is forwarded to the
Dynamo frontend as CLI args via router_kwargs.
Bases: BaseServerConfig
Server-level Dynamo config.
Bases: BaseModelConfig
Dynamo vLLM model config.
Typed fields cover deployment/placement knobs Curator branches on; anything
else is forwarded to python -m dynamo.vllm via dynamo_kwargs.
kv_events_config and kv_transfer_config are Curator-managed
(init=False): events are derived from router state + port allocation,
transfer defaults to NixlConnector for disagg.
Serve one or more models behind a typed backend config.
OpenAI-compatible base URL for the served models.
Check every model is accepted by the backend and that all models share one concrete type.
Poll /v1/models until all expected models appear in the response.
Deploy all models and wait for them to become healthy.
Shut down the active inference backend and release resources.
Check whether any inference server is currently running in this process.