TensorRT-LLM Sidecar
TensorRT-LLM Sidecar
Run Dynamo beside a TensorRT-LLM engine through its OpenEngine gRPC API.
Experimental. The TensorRT-LLM sidecar, launcher, packaging, and feature coverage can change without notice.
dynamo-trtllm-sidecar is a CPU-only Dynamo worker that connects to
TensorRT-LLM’s OpenEngine gRPC API (openengine.v1), served by
trtllm-serve --grpc --grpc-protocol openengine. It preserves the upstream
engine process and argument surface while using Dynamo for request handling and
distributed serving. See the
Sidecar Backends page for the common
architecture.
Readiness
This table covers launch topology only. The
TensorRT-LLM feature matrix describes the
in-process backend; sidecar feature parity is still under evaluation.
Disaggregated prefill/decode is supported over the OpenEngine contract: a
prefill worker marks its request context_only and returns the PrefillReady
KV handoff that a decode worker replays. Running it needs an engine with a KV
cache transceiver configured on both legs. See the
TensorRT-LLM sidecar README
for other protocol limitations.
Launch Locally
Build or install Dynamo from a source checkout so dynamo-trtllm-sidecar is on
PATH. The engine must serve --grpc --grpc-protocol openengine, which means
nvcr.io/nvidia/tensorrt-llm/release:1.3.0rc27.dev202609170000 or newer; the
launchers install the pinned OpenEngine Python bindings, which those releases do
not ship.
Each launcher starts the Dynamo frontend, the engine, and the sidecar, and binds
the engine’s gRPC endpoint to loopback — it is unauthenticated and plaintext.
launch/disagg.sh is the same on two GPUs, with a NIXL cache transceiver on both
engines for the KV handoff. Either way the frontend serves the usual endpoint:
Deploy on Kubernetes
The source tree ships two
deployment manifests:
agg.yaml, and disagg.yaml, which runs prefill and decode as separate worker
pods. Both need two images you build yourself: dynamo-sidecar, which carries
all three engine-specific executables, and a TensorRT-LLM image with the pinned
OpenEngine bindings layered on — the release ships the servicer but not the
bindings that it and the manifests’ health probes import.
Read the disaggregated manifest’s header before applying it: it runs the engines
over TCP/CUDA-IPC and requests no rdma/ib, which you add on a fabric that
provides it. The
README has the full walkthrough.