TensorRT-LLM Sidecar
TensorRT-LLM Sidecar
Run Dynamo beside a stock TensorRT-LLM engine through native gRPC.
Experimental. The TensorRT-LLM sidecar, launcher, packaging, and feature coverage can change without notice.
dynamo-trtllm-sidecar is a CPU-only Dynamo worker that connects to
TensorRT-LLM’s native gRPC service. It preserves the upstream engine process
and argument surface while using Dynamo for request handling and distributed
serving. See the
Sidecar Backends page for the common
architecture.
Readiness
This table covers launch topology only. The TensorRT-LLM feature matrix describes the in-process backend; sidecar feature parity is still under evaluation. The current native gRPC contract does not provide the prefill/decode handoff needed for disaggregated serving. See the TensorRT-LLM sidecar README for other protocol limitations.
Launch Locally
From a Dynamo source checkout, build or install Dynamo so
dynamo-trtllm-sidecar is on PATH. Install a TensorRT-LLM release that
provides tensorrt_llm.commands.serve --grpc.
Start Dynamo’s local discovery services, then run the aggregated launcher:
The launcher starts the Dynamo frontend, TensorRT-LLM engine, and sidecar. It binds TensorRT-LLM’s native gRPC endpoint to loopback.
Verify the frontend:
Deploy on Kubernetes
No published sidecar image is available yet. Follow the
Kubernetes quick start
to build dynamo-sidecar, which contains all three engine-specific sidecar
executables. The TensorRT-LLM manifest runs dynamo-trtllm-sidecar as the
container command and pairs it with a stock upstream TensorRT-LLM image. The
source tree includes an
aggregated deployment manifest.