Encoder Disaggregation

Separate vision encoding into a dedicated worker for independent scaling
以 Markdown 格式查看

Overview

Encoder disaggregation separates the vision encoder from the prefill/decode pipeline into its own worker. Instead of running image encoding inline, a dedicated encode worker handles media processing and transfers the resulting embeddings to downstream workers via NIXL (RDMA).

This enables:

  • Independent scaling of encode workers based on vision workload
  • Reduced GPU memory pressure on prefill/decode workers
  • Better GPU utilization by matching worker counts to actual bottlenecks

When to Use

Use encoder disaggregation when:

  • Vision encoding is a bottleneck and you need to scale encoders independently of LLM workers
  • You want to run the vision encoder on different hardware (e.g., smaller GPUs for encoding, larger GPUs for LLM inference)
  • Your deployment handles high volumes of multimodal requests and encoding throughput is limiting

For simple deployments or development/testing, the aggregated (EPD) pattern is easier to set up.

Deployment Patterns

E/PD

Separate encoder, combined prefill and decode

Frontend → Processor → Encode Worker → PD Worker → Response

The encode worker runs the vision model and transfers embeddings over NIXL to the combined prefill and decode worker.

E/P/D

Separate encoder, prefill, and decode

Frontend → Processor → Encode Worker → Prefill Worker → Decode Worker → Response

The encode worker transfers embeddings over NIXL to the prefill worker. The prefill worker then transfers the KV cache to the decode worker.

Launching

cd $DYNAMO_HOME/examples/backends/vllm
# E/PD
bash launch/disagg_multimodal_e_pd.sh --model "Qwen/Qwen3-VL-30B-A3B-Instruct-FP8"
# E/P/D
bash launch/disagg_multimodal_epd.sh --model "Qwen/Qwen3-VL-30B-A3B-Instruct-FP8"

See the backend-specific documentation (vLLM, TRT-LLM, SGLang) for full configuration details and component flags.

Support Matrix

BackendE/PDE/P/DNotes
vLLMYesYesSeparate encode worker currently handles image_url inputs; video_url inputs stay on the prefill/PD path
TensorRT-LLM—YesSupports image URLs (via MultimodalEncoder) and pre-computed embeddings (via NIXL)
SGLangYesYesNIXL for embeddings; bootstrap mechanism for P/D KV transfer