Disaggregated Serving

以 Markdown 格式查看

Disaggregated serving separates the prefill and decode phases of inference into independent workers, each scalable on its own and connected by a KV cache transport such as NIXL or Mooncake (TokenSpeed only). This is the local equivalent of the disaggregated pattern in the Kubernetes Disaggregated Serving guide — the same architecture, driven by CLI flags instead of a DynamoGraphDeployment.

For the architecture and when to use it, see Disaggregated Serving.

How it works locally

A disaggregated deployment is still a frontend plus workers (see the Overview), except instead of one worker doing both phases, you run:

  • A prefill worker — computes the prompt’s KV cache.
  • A decode worker — receives the KV cache over the backend’s transfer protocol and generates tokens.

The vLLM, SGLang, and TensorRT-LLM launch/disagg.sh presets start the frontend and both workers, pinned to separate GPUs with the KV-transfer config wired up. Requires 2 GPUs.

Disaggregated serving

cd $DYNAMO_HOME/examples/backends/vllm
bash launch/disagg.sh

Each worker needs a unique VLLM_NIXL_SIDE_CHANNEL_PORT; the preset sets these for you.

TokenSpeed

For LongCat-Flash with two prefill replicas and KV-aware routing, use the TokenSpeed disaggregation example. This experimental backend uses Mooncake for KV transfer and has its own hardware and launch requirements.

Adding KV-aware routing

To scale each pool and route requests by cache overlap, use the disagg_router.sh presets (2 prefill + 2 decode workers, 4 GPUs). See KV-Aware Routing.

With a disaggregated deployment running, try adding another prefill worker — the frontend discovers and uses it automatically.

Troubleshooting

Workers fail to start with NIXL errors. Ensure NIXL is installed and side-channel ports don’t conflict. Each worker in a multi-worker setup needs a unique side-channel port.

See also