Disaggregated Serving
Disaggregated serving separates the prefill and decode phases of inference into independent workers, each scalable on its own and connected by a KV cache transport such as NIXL or Mooncake (TokenSpeed only). This is the local equivalent of the disaggregated pattern in the Kubernetes Disaggregated Serving guide — the same architecture, driven by CLI flags instead of a DynamoGraphDeployment.
For the architecture and when to use it, see Disaggregated Serving.
How it works locally
A disaggregated deployment is still a frontend plus workers (see the Overview), except instead of one worker doing both phases, you run:
- A prefill worker — computes the prompt’s KV cache.
- A decode worker — receives the KV cache over the backend’s transfer protocol and generates tokens.
The vLLM, SGLang, and TensorRT-LLM launch/disagg.sh presets start the frontend and both workers, pinned to separate GPUs with the KV-transfer config wired up. Requires 2 GPUs.
Disaggregated serving
vLLM
SGLang
TensorRT-LLM
Each worker needs a unique VLLM_NIXL_SIDE_CHANNEL_PORT; the preset sets these for you.
TokenSpeed
For LongCat-Flash with two prefill replicas and KV-aware routing, use the TokenSpeed disaggregation example. This experimental backend uses Mooncake for KV transfer and has its own hardware and launch requirements.
Adding KV-aware routing
To scale each pool and route requests by cache overlap, use the disagg_router.sh presets (2 prefill + 2 decode workers, 4 GPUs). See KV-Aware Routing.
With a disaggregated deployment running, try adding another prefill worker — the frontend discovers and uses it automatically.
Troubleshooting
Workers fail to start with NIXL errors. Ensure NIXL is installed and side-channel ports don’t conflict. Each worker in a multi-worker setup needs a unique side-channel port.
See also
- Disaggregated Serving (architecture) — design and rationale
- KV-Aware Routing — route across prefill/decode pools
- Tuning Disaggregated Performance — P/D tuning guide
- vLLM local deployment examples — copyable commands for each launch script