Model Deployment

Deploy a model on a GPU cluster with Dynamo’s two deployment CRDs — DGD for manual control, DGDR for intent-driven auto deployment.

View as Markdown

Deploying a model on Dynamo means describing an inference graph — a Frontend plus one or more workers — as a Kubernetes Custom Resource, then letting the Dynamo operator reconcile it into running pods, services, and scheduling resources. There are two CRDs you can start from, and the whole section is organized around choosing between them.

Two ways to deploy: DGD and DGDR

DynamoGraphDeployment (DGD)DynamoGraphDeploymentRequest (DGDR)
PathManual — you author the specAutomatic — you describe intent
You provideThe full spec: components, parallelism, replicas, resource limitsModel, backend, workload, and optional SLA targets
What happensThe operator reconciles your spec into pods directlyThe profiler sizes the deployment and generates a DGD for you
Best forKnown-good configs, tuned recipes, full controlNew model/hardware combinations, SLA-driven sizing

DGD is the primary path

A DynamoGraphDeployment (DGD) is the canonical resource that serves traffic. You write the spec, kubectl apply it, and the operator runs it. Because you control every field, a DGD works for any model and backend, and it is the object that actually persists and serves. Even when you use DGDR, what it produces is a DGD. Start here: Deploy with DGD.

DGDR automates sizing — when your model is supported

A DynamoGraphDeploymentRequest (DGDR) is the deploy-by-intent path. Instead of hand-authoring parallelism and replica counts, you describe what you want to run and Dynamo’s profiler analyzes your GPUs, selects a configuration, and generates a DGD.

DGDR is convenient but does not yet cover every model, hardware, and backend combination. Its fast “rapid” profiling relies on the AIConfigurator support matrix; for combinations outside that matrix it falls back to a naive memory-fit configuration that may not be optimal. For this reason DGD remains the primary path, and even in the DGDR flow the generated DGD is something you can inspect and edit before it serves traffic. See Auto Deploy with DGDR.

Aggregated or disaggregated — supported by both

Independent of which CRD you author, you choose how prefill and decode are placed. Both DGD and DGDR support both options.

  • Aggregated — one worker runs both the prefill (prompt) and decode (generation) phases. The simplest deployment; a good default for smaller models and uniform traffic.
  • Disaggregated — separate prefill and decode worker pools, each sized and scaled independently, with the KV cache transferred between them over the network. Use it when prefill and decode have different bottlenecks — long prompts saturating prefill, or high concurrency saturating decode. It needs an RDMA-capable fabric for the KV transfer to perform well.

See Disaggregated Serving for the architecture and the disaggregation-specific configuration.

Where to go next

Run one model end to end with the Kubernetes Quickstart first if you are new here. Then pick your path:

GoalGuide
Author a deployment by hand (start here)Deploy with DGD
Let Dynamo size and generate the deploymentAuto Deploy with DGDR
Split prefill and decode for performanceDisaggregated Serving
Auto-pick TP/PP/DP for a latency targetSizing with AIConfigurator
Scale a model across multiple nodesMultinode Deployments
Understand the full resource model (DGD, DCD, DGDR, recipes)Deployment Overview
Load models faster across podsModel Caching

If you are still evaluating Dynamo locally, start with the Quickstart and Local Installation first.