Dynamo
Dynamo works with vLLM, SGLang, and TensorRT-LLM; supports NVIDIA and AMD GPUs and Intel XPUs; runs on Kubernetes, Slurm, or locally; and scales one component or the full stack.
Deployment walkthrough
See Dynamo in action
01 · Distributed inference
Scale the system around the model
Disaggregated serving, KV-aware routing, cache management, and autoscaling coordinate the full inference pipeline.
02 · Engine interoperability
Bring your inference engine
Use Dynamo with vLLM, SGLang, or TensorRT-LLM, then connect the network, storage, and KV cache layers you need.
03 · Infrastructure choice
Run where your workloads run
Deploy on Kubernetes, schedule with Slurm, or start locally across NVIDIA and AMD GPUs and Intel XPUs.
04 · Modular adoption
Adopt one component or the full stack
Start with the frontend, router, planner, or cache manager, and add the rest as your deployment grows.
Community Calendar
Open Google CalendarMeetups, talks, and community meetings
Community events
Previewed from the public Dynamo Google Calendar.Upcoming event