DynamoGraphDeployment (DGD) Reference
DynamoGraphDeployment (DGD) Reference
Field reference for the DynamoGraphDeployment custom resource — the canonical, live description of a Dynamo inference graph.
A DynamoGraphDeployment (DGD) is the canonical description of a running Dynamo inference graph. You list the components that make up the graph — frontend, workers, prefill and decode workers, planner, EPP — and the operator reconciles them into DynamoComponentDeployment resources, pods, and Services.
A DGD is the resource that persists and serves traffic. Author one directly when you have a known-good configuration or a tuned recipe; generate one from intent with a DynamoGraphDeploymentRequest (DGDR).
This page documents the nvidia.com/v1beta1 API — the served, current version.
LPX components are experimental and require the operator’s lpx.enabled setting. When disabled,
admission rejects new or changed LPX components but permits otherwise valid updates that leave them
unchanged, and deletion. Existing LPX DGDs report a disabled status without further workload
reconciliation; installed CRDs and workloads are preserved.
Each entry in spec.components is a DynamoComponentDeploymentSharedSpec — the same fields documented on the DCD Reference. This page covers the graph-level fields; per-component fields live there.
Minimal example
Spec reference
Components deployed as part of this graph. Each entry carries its own stable logical name (unique within the list, case-insensitively). Component types are repeatable except type: epp, which may appear at most once. Maximum 25 components. For the full per-component field set, see DCD Reference — Shared component spec.
Each new non-LPX component must supply a podTemplate.spec.containers entry named main with a non-empty image. For an image without a parseable semantic-version tag, set spec.components[*].runtimeVersionOverride to the image’s Dynamo runtime version. Also set it when the tag identifies a different version, such as the inference engine’s version. See Runtime Version Compatibility. DGD has no graph-level runtimeVersionOverride; the DGDR-level field supplies a default when generating components.
LPX components use roles[].podTemplate instead and do not require runtimeVersionOverride. See Shared component spec.
Customizes the root resource generated by the selected workload provider. Currently supported only for Grove, with apiVersion: grove.io/v1alpha1 and target PodCliqueSet. The value may set only spec.template.topologyConstraint. This field cannot select or change the provider. See ProviderOverride.
GPU backend framework for worker components. When omitted, the operator infers the framework from each worker’s command and arguments. When set, it must match the detected worker framework. LPX components omit this field.
Allowed values: sglang vllm trtllmEnvironment variables prepended to every component’s environment. A component-specific env entry with the same name takes precedence and may reference values from this list.
Annotations propagated to all child resources (scaling adapters, DCDs, Deployments, and pod templates). Component-level podTemplate values take precedence on conflict.
Labels propagated to all child resources. Same precedence rules as annotations.
For LPX, metadata synchronization keeps the PCS labels and clique templates consistent after queue edits. Moving an existing KAI PodGroup to another queue depends on the Grove backend; metadata synchronization does not recreate existing PodGroups.
Name of the PriorityClass to use for Grove PodCliqueSets. Requires the Grove pathway. See Multinode Orchestration.
Restart policy for the graph. Must be unset on creation; set or change restart.id on an existing DGD to trigger a restart.
Graphs containing LPX components require explicit restart.strategy.type: Parallel; Sequential and omitted strategies are rejected.
Arbitrary string; any change to it initiates a restart of the graph according to the configured strategy.
How components are restarted.
Whether components restart one at a time or all at once.
Allowed values: Sequential ParallelComplete ordered set of component names for a sequential restart. Omit to use the controller’s default order. Must not be set for parallel restarts.
Deployment-level topology constraint. Components without their own topologyConstraint inherit this value. See the SpecTopologyConstraint API.
Name of the ClusterTopology resource defining the topology hierarchy for this deployment.
Default topology domain to pack pods within. Optional at this level; omit when only components carry constraints.
Graph-level opt-in preview features whose API shape may change in breaking ways between v1beta1 releases. Component-level experimental features live under spec.components[*].experimental instead — see DCD Reference — ExperimentalSpec.
Topology-aware routing for KV-cache transfers between prefill and decode workers. Set exactly one of labelKey or clusterTopologyName. See the KvTransferPolicy API.
References a Grove ClusterTopology CR. The operator reads the CR’s topology levels and projects them through Dynamo-owned pod labels for worker topology metadata. Mutually exclusive with labelKey.
A Kubernetes node label key (for example topology.kubernetes.io/zone) whose value identifies the topology domain for each worker. The operator copies the node label onto worker pods so the runtime can publish it as worker metadata. Should correspond to the topology level named in domain. Mutually exclusive with clusterTopologyName.
Logical name for the topology level to enforce (for example zone, rack). The router uses this to match workers that share the same value for the label identified by labelKey. Free-form string matching ^[a-z0-9]([a-z0-9-]*[a-z0-9])?$.
How the selected prefill worker’s topology is applied to decode routing. required allows only decode workers in the same topology domain as the selected prefill worker; preferred keeps all decode workers eligible but biases selection toward the same domain.
Required and used only when enforcement is preferred. Higher values create a stronger same-domain routing preference but do not guarantee same-domain selection; the value is not a probability. 0 disables the topology preference; 1 is the strongest supported preference. Minimum 0, maximum 1.
Status
The operator maintains observed state under status.
High-level textual status of the graph deployment lifecycle.
Latest observed conditions, merged by type. Includes Available and DynamoComponentReady.
Per-component replica status, keyed by component name.
Most recent metadata.generation observed by the controller.
Status of a graph-level restart in progress.
The restart ID currently being processed. Matches spec.restart.id.
Phase of the restart.
Allowed values: Pending Restarting Completed Failed SupersededNames of the components currently being restarted.
Per-component checkpoint status, keyed by component name. See the ComponentCheckpointStatus API.
Progress of operator-managed rolling updates. Currently supported only for single-node, non-Grove deployments. See the RollingUpdateStatus API.
Current phase of the rolling update.
Allowed values: Pending InProgress Completed FailedWhen the rolling update began.
metav1.TimeWhen the rolling update completed, successfully or failed.
metav1.TimeComponents that have completed the rolling update.
Scheduler placement score and reporting state. The schema is available, but the controller does not yet populate this field.
Worst placement score across relevant scheduler placement units. Minimum 0, maximum 1; higher is better. Scores are comparable only between DGDs using the same scheduler scoring contract and version.
Reported means every scored placement unit has a score; Partial means only some do. Unsupported means the backend exposes no score. Unknown means scoring is supported but the current value is indeterminate, and score must be absent.
Additional Types
Nested struct types referenced by the fields above, broken out here to keep the field lists shallow. Types prefixed core/v1. or metav1. are standard Kubernetes types and link to their Go package documentation instead of being expanded.
ProviderOverride
A sparse provider-native fragment, supported only for Grove-backed DGD resources. The same shape is used at graph, component, and role scope; the location determines the allowed target. Standalone DCDs do not support provider overrides.
Provider schema version. Must be grove.io/v1alpha1.
Provider resource or embedded schema. Admission resolves and persists this value when omitted. Graph scope uses PodCliqueSet; component scope uses PodCliqueTemplateSpec or PodCliqueScalingGroupConfig according to the component shape; multinode leader and worker roles use PodCliqueTemplateSpec. Other targets are rejected.
Sparse fragment of the target schema. PodCliqueSet accepts only spec.template.topologyConstraint; PodCliqueTemplateSpec and PodCliqueScalingGroupConfig accept only topologyConstraint. Other fields are rejected.
ComponentReplicaStatus
Replica information for a single component, keyed by component name in status.components.
Underlying Kubernetes resource kind backing the component.
Allowed values: PodClique PodCliqueScalingGroup Deployment LeaderWorkerSetUnderlying Kubernetes resource names for this Dynamo component. During normal operation this contains a single name; during rolling updates it contains both the old and new resource names.
Effective Dynamo runtime namespace for the component. Workers may include a generation suffix; non-workers use the base namespace. During a rolling update, worker status retains the old active revision’s namespace until cutover completes.
GPUs assigned to one inference engine across all nodes of a component replica, excluding independent auxiliary GPU allocations. A reported 0 means the engine has no GPUs; consult gpusPerReplica for auxiliary allocations. Minimum 0.
Unique GPU allocation added by scaling the component up by one replica, across all nodes, application and initialization phases, and provider-owned pods. Scalar GPUs use the effective Kubernetes pod scheduling footprint; shared DRA claims count once. A reported 0 means a resolved non-GPU allocation; omission means no current shape is available. Minimum 0.
Total number of non-terminated replicas. Minimum 0.
Number of replicas at the current/desired revision. Minimum 0.
Number of ready replicas. Populated for PodClique, Deployment, and LeaderWorkerSet; not available for PodCliqueScalingGroup. Minimum 0.
Number of available replicas. Populated for Deployment and PodCliqueScalingGroup; not available for PodClique or LeaderWorkerSet. Minimum 0.
Replicas scheduled by the backend, in complete Dynamo component-replica units rather than raw pod counts. Helps distinguish scheduling shortfalls from runtime readiness. Omitted when the backend cannot report a reliable count; absence means not reported, not zero scheduled. Minimum 0.
ComponentCheckpointStatus
Checkpoint information for a single component, keyed by component name in status.checkpoints.
Name of the active PodSnapshot.
Deprecated legacy Dynamo artifact ID. Standalone snapshots leave this field empty.
Deprecated legacy checkpoint identity hash. Standalone snapshots leave this field empty.
Whether the checkpoint artifact is ready for future pods to restore.