KV Cache Transfer
For general TensorRT-LLM features and configuration, see the Reference Guide.
In disaggregated serving architectures, KV cache must be transferred between prefill and decode workers. TensorRT-LLM supports three methods for this transfer:
- NIXL with UCX (default)
- NIXL with Libfabric
- using UCX directly
Using NIXL for KV Cache Transfer
Start the disaggregated service: See Disaggregated Serving to learn how to start the deployment.
Default Method: NIXL with UCX
By default, TensorRT-LLM uses NIXL (NVIDIA Inference Xfer Library) with UCX (Unified Communication X) as backend for KV cache transfer between prefill and decode workers. NIXL is NVIDIA’s high-performance communication library designed for efficient data transfer in distributed GPU environments.
Specify Backends for NIXL
TensorRT-LLM supports two NIXL communication backends: UCX and LIBFABRIC. By default, UCX is used if no backend is explicitly specified. Dynamo currently supports both backends. For AWS EFA deployments, UCX with SRD transport is the tested and recommended backend (see AWS EFA below).
Alternative Method: UCX
TensorRT-LLM can also leverage UCX (Unified Communication X) directly for KV cache transfer between prefill and decode workers. To enable UCX as the KV cache transfer backend, set cache_transceiver_config.backend: UCX in your engine configuration YAML file.
cache_transceiver_config.backend accepts the following values:
The precedence above matches TensorRT-LLM 1.3.0rc22; check CacheTransceiverConfig._resolve_default_backend in tensorrt_llm/llmapi/llm_args.py.
AWS EFA
On AWS, UCX uses the SRD (Scalable Reliable Datagram) transport over EFA devices. NIXL discovers EFA rdmap* devices automatically through UCX — no NIXL-level configuration changes are needed.
Image options:
- Pre-built EFA image: A dedicated EFA image with the EFA SDK baked in is available on NGC, for both AMD64 and ARM64:
On 1.2.1 the same image is tagged 1.2.1-efa-amd64. Despite the suffix that tag is a multi-arch manifest covering AMD64 and ARM64; the name was corrected to -efa in 1.3.0. Pull the tag exactly as written — 1.2.1-efa was never published.
See Release Artifacts for all available EFA images.
- Host-mount approach (ARM64 / GB200): Instead of the pre-built image, you can run the standard
tensorrtllm-runtimeimage and mount the EFA SDK from the host node, which keeps the SDK in step with the host driver. This is what we tested on GB200 NVL72:
EFA resource requests:
Required environment variables for EFA workers (set on both prefill and decode):
FI_EFA_ENABLE_SHM_TRANSFER must be 0. SHM transfers break NIXL GPU buffer registrations.Security context: AWS EFA currently requires privileged mode:
NIXL Plugin ABI Mismatch on Decode Multinode
When running multinode decode, the decode leader launches workers via mpirun -> mgmn_worker_node, which loads TRT-LLM’s bundled NIXL rather than the system nixl_cu13. The container’s default NIXL_PLUGIN_DIR points to system plugins that are ABI-incompatible with TRT-LLM’s bundled NIXL. Override this on the decode service only:
Do not set this on prefill workers — they use nixl_cu13 which is compatible with the system plugins.
ComputeDomain for GB200 NVL72
On GB200 NVL72 racks, NCCL requires a ComputeDomain CR for proper cuMem/NVLS initialization. Without it, workers fail with NCCL error 'unhandled system error' during model loading.
Both prefill and decode services must include ResourceClaims:
Required NCCL environment variables for GB200:
Verifying EFA is Active
After deployment, confirm NIXL is using SRD over EFA in the worker logs:
Expected output:
srd/rdmap*confirms SRD transport over EFA devices- Multiple
rdmapentries correspond to one EFA device per GPU