Dynamo

An open source inference framework for serving generative AI.

Dynamo works with vLLM, SGLang, and TensorRT-LLM; supports NVIDIA and AMD GPUs and Intel XPUs; runs on Kubernetes, Slurm, or locally; and scales one component or the full stack.

Get started

Deployment walkthrough

See Dynamo in action

DynamoQwen3-235B deploymentDemo runningShow playback controls

01 · Distributed inference

Scale the system around the model

Disaggregated serving, KV-aware routing, cache management, and autoscaling coordinate the full inference pipeline.

02 · Engine interoperability

Bring your inference engine

Use Dynamo with vLLM, SGLang, or TensorRT-LLM, then connect the network, storage, and KV cache layers you need.

03 · Infrastructure choice

Run where your workloads run

Deploy on Kubernetes, schedule with Slurm, or start locally across NVIDIA and AMD GPUs and Intel XPUs.

04 · Modular adoption

Adopt one component or the full stack

Start with the frontend, router, planner, or cache manager, and add the rest as your deployment grows.

Community Calendar

Open Google Calendar

Meetups, talks, and community meetings

Community events

Previewed from the public Dynamo Google Calendar.
Recent events