Feature Matrix

View as Markdown

This document provides a comprehensive compatibility matrix for key Dynamo features across the supported backends.

Updated for Dynamo v1.2.0

Legend:

  • ✅ : Supported
  • 🚧 : Work in Progress / Experimental / Limited

Quick Comparison

FeatureSGLangTensorRT-LLMvLLMSource
Disaggregated Serving✅✅✅Design Doc
KV-Aware Routing✅✅✅Router Doc
SLA-Based Planner✅✅✅Planner Doc
KV Block Manager🚧✅✅KVBM Doc
Multimodal (Image)✅✅✅Multimodal Doc
Multimodal (Video)✅✅Multimodal Doc
Multimodal (Audio)🚧Multimodal Doc
Request Migration✅🚧✅Migration Doc
Request Cancellation🚧✅✅Backend READMEs
LoRA✅K8s Guide
Tool Calling✅✅✅Tool Calling Doc
Speculative Decoding🚧✅✅Backend READMEs
Dynamo Snapshot✅✅Snapshot Docs

1. vLLM Backend

vLLM offers the broadest feature coverage in Dynamo, with full support for disaggregated serving, KV-aware routing, KV block management, LoRA adapters, and multimodal inference including video and audio.

Source: docs/backends/vllm/README.md

FeatureDisaggregated ServingKV-Aware RoutingSLA-Based PlannerKV Block ManagerMultimodalRequest MigrationRequest CancellationLoRATool CallingSpeculative Decoding
Disaggregated Serving—
KV-Aware Routing✅—
SLA-Based Planner✅✅—
KV Block Manager✅✅✅—
Multimodal✅✅1—✅—
Request Migration✅✅✅✅✅—
Request Cancellation✅✅✅✅✅✅—
LoRA✅✅2—✅—✅✅—
Tool Calling✅✅✅✅✅✅✅✅—
Speculative Decoding✅✅—✅—✅✅—✅—

Notes:

  1. Multimodal + KV-Aware Routing: Image-aware KV routing is supported in the documented vLLM paths. The default Rust frontend path supports model families handled by llm-multimodal; the Python chat-processor path delegates to vLLM’s multimodal processor. (Source)
  2. KV-Aware LoRA Routing: vLLM supports routing requests based on LoRA adapter affinity.
  3. Audio Support: vLLM supports audio models like Qwen2-Audio (experimental). (Source)
  4. Video Support: vLLM supports video input with frame sampling. (Source)
  5. Speculative Decoding: Eagle3 support documented. (Source)

2. SGLang Backend

SGLang is optimized for high-throughput serving with fast primitives, providing robust support for disaggregated serving, KV-aware routing, and request migration.

Source: docs/backends/sglang/README.md

FeatureDisaggregated ServingKV-Aware RoutingSLA-Based PlannerKV Block ManagerMultimodalRequest MigrationRequest CancellationLoRATool CallingSpeculative Decoding
Disaggregated Serving—
KV-Aware Routing✅—
SLA-Based Planner✅✅—
KV Block Manager🚧🚧🚧—
Multimodal✅21—🚧—
Request Migration✅✅✅🚧✅—
Request Cancellation🚧3✅✅🚧🚧✅—
LoRA🚧—
Tool Calling✅✅✅🚧✅✅✅—
Speculative Decoding🚧🚧—🚧—🚧—🚧—

Notes:

  1. Multimodal + KV-Aware Routing: Not supported. (Source)
  2. Multimodal Patterns: Supports simple Aggregated EPD, E/PD, and E/P/D patterns. Traditional Disagg EP/D is not supported. (Source)
  3. Request Cancellation: Cancellation during the remote prefill phase is not supported in disaggregated mode. (Source)
  4. Speculative Decoding: Code hooks exist (spec_decode_stats in publisher), but no examples or documentation yet.

3. TensorRT-LLM Backend

TensorRT-LLM delivers maximum inference performance and optimization, with full KVBM integration and robust disaggregated serving support.

Source: docs/backends/trtllm/README.md

FeatureDisaggregated ServingKV-Aware RoutingSLA-Based PlannerKV Block ManagerMultimodalRequest MigrationRequest CancellationLoRATool CallingSpeculative Decoding
Disaggregated Serving—
KV-Aware Routing✅—
SLA-Based Planner✅✅—
KV Block Manager✅✅✅—
Multimodal✅1✅2—✅—
Request Migration✅✅✅✅🚧—
Request Cancellation✅3✅3✅3✅3✅3✅3—
LoRA—
Tool Calling✅✅✅✅✅✅✅—
Speculative Decoding✅✅—✅—✅✅✅—

Notes:

  1. Multimodal Disaggregation: Supports EP/D (Traditional) and E/P/D (Full Disaggregation) image flows, including image URLs and pre-computed embeddings. (Source)
  2. Multimodal + KV-Aware Routing: Image-aware KV routing is supported through the dedicated TRT-LLM MM Router Worker. It requires KV event publishing on the TRT-LLM workers. (Source)
  3. Request Cancellation: Due to known issues, the TensorRT-LLM engine is temporarily not notified of request cancellations, meaning allocated resources for cancelled requests are not freed.