vLLM Sidecar
Experimental. The sidecars and their deployment examples are experimental. Manifests, flags, and behavior may change without notice.
dynamo-vllm-sidecar connects a Dynamo worker to vLLM’s native gRPC server
(vllm-rs). See Sidecar Backends
for the architecture.
For the best and latest support, use the upstream vLLM nightly image,
which carries the latest gRPC server updates: vllm/vllm-openai:nightly.
Support Matrix
Launch Locally
See lib/sidecar/vllm/launch/
for all topologies. For example, aggregated serving on one GPU:
In a second terminal:
Deploy on Kubernetes
See lib/sidecar/vllm/deploy/
for all manifests. For example, aggregated serving:
Before applying, replace <your-registry>/dynamo-sidecar in the manifest with a
sidecar image.
Topologies
The frontend reaches each sidecar over Dynamo’s request, discovery, and event planes; the sidecar reaches the engine over its native gRPC API. Dashed arrows carry KV events.
Single-Node TP
One engine on one node, with one sidecar.
Multi-Node TP
One engine spans two nodes. Only the leader node has a sidecar; the follower node holds the remaining TP ranks.
Multi-Node DP
Hybrid DP load balancing: each node runs vLLM for its local DP ranks plus a sidecar that serves requests. The frontend routes to either node.