Quickstart
Get a Dynamo OpenAI-compatible endpoint running locally on an NVIDIA GPU or Intel XPU.
Choose Your Path
You’re here. Container fast path.
Full walkthrough — PyPI, configuration.
Kubernetes-native production path.
For contributors against main.
Dynamo is backend-agnostic and Kubernetes-native without being Kubernetes-only. Use this container path to try the same frontend/router/worker stack locally; use the Kubernetes path when you want the operator, CRDs, Gateway API integration, autoscaling, scheduling, and cluster lifecycle management.
Run Dynamo Locally
Choose and install a build
Choose your local build
Unavailable combinations remain visible to show the current support boundary.
docker run --gpus all --network host --ipc host --rm -it nvcr.io/nvidia/ai-dynamo/sglang-runtime:1.3.0
Hugging Face token required for gated models. Llama, Kimi, Qwen-VL, and other gated models require HF_TOKEN in your environment and accepting the model card’s license on huggingface.co. Set export HF_TOKEN=hf_… before launching.
The remaining steps run inside the selected container. If you installed an NVIDIA wheel instead, run the same python3 -m dynamo.* commands in your Python environment. See Local Installation for host prerequisites and virtual environment setup.
For published NVIDIA container versions and tags, see Release Artifacts.
Start the frontend
Start the OpenAI-compatible frontend on port 8000:
--discovery-backend file avoids needing etcd. To run the frontend and worker in the same terminal, background each command with > logfile.log 2>&1 &.
Start a worker
In another terminal, select the hardware and backend you installed, then launch the worker:
Choose your worker command
Select the same hardware and backend you used for the installation.
python3 -m dynamo.sglang --model-path Qwen/Qwen3-0.6B --discovery-backend file
Verify the endpoint
Check that the endpoint is up:
If you see OK, send a chat completion:
Connection refused? The frontend takes a few seconds to start — retry. For production liveness and readiness probes, see Health Check Reference.
From the Digest
How Dynamo optimizes for agentic workloads at three layers: the frontend API, the router, and KV cache management.
How Dynamo’s concurrent global index evolved through six iterations to sustain over 100M ops/sec.
Dive Deeper
Pick a full install path from the four options above, or explore how Dynamo works under the hood: