vLLM Sidecar

Run Dynamo beside a stock vLLM engine through native gRPC.
以 Markdown 格式查看

Experimental. The sidecars and their deployment examples are experimental. Manifests, flags, and behavior may change without notice.

dynamo-vllm-sidecar connects a Dynamo worker to vLLM’s native gRPC server (vllm-rs). See Sidecar Backends for the architecture.

For the best and latest support, use the upstream vLLM nightly image, which carries the latest gRPC server updates: vllm/vllm-openai:nightly.

Support Matrix

FeatureSupported
AggregatedYes
DisaggregatedYes
KV routingYes

Launch Locally

See lib/sidecar/vllm/launch/ for all topologies. For example, aggregated serving on one GPU:

export DYN_DISCOVERY_BACKEND=file # single host: no etcd or NATS needed
./lib/sidecar/vllm/launch/agg.sh

In a second terminal:

curl -s localhost:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"Qwen/Qwen3-0.6B","messages":[{"role":"user","content":"Hello"}],"max_tokens":32}'

Deploy on Kubernetes

See lib/sidecar/vllm/deploy/ for all manifests. For example, aggregated serving:

Before applying, replace <your-registry>/dynamo-sidecar in the manifest with a sidecar image.

kubectl apply -f lib/sidecar/vllm/deploy/agg.yaml -n <namespace>
kubectl port-forward -n <namespace> svc/vllm-sidecar-agg-frontend 8000:8000

Topologies

The frontend reaches each sidecar over Dynamo’s request, discovery, and event planes; the sidecar reaches the engine over its native gRPC API. Dashed arrows carry KV events.

Single-Node TP

One engine on one node, with one sidecar.

On one node, a request reaches the vLLM tensor-parallel ranks through the Dynamo Sidecar. The Dynamo Frontend sends requests over the request plane to the sidecar.

Multi-Node TP

One engine spans two nodes. Only the leader node has a sidecar; the follower node holds the remaining TP ranks.

When one vLLM engine spans two nodes with tensor parallelism, only the leader node runs a Dynamo Sidecar. The Dynamo Frontend sends requests over the request plane to the sidecar on Node 0.

Multi-Node DP

Hybrid DP load balancing: each node runs vLLM for its local DP ranks plus a sidecar that serves requests. The frontend routes to either node.

vLLM hybrid data parallelism across two nodes. The Dynamo Frontend router picks a DP rank and sends requests over the request plane to the Dynamo Sidecar on the node that owns that rank: node 0 serves DP ranks 0-1 and node 1 serves DP ranks 2-3.