Standalone Router
Run KV-aware routing as its own Dynamo service.
The standalone router is a backend-agnostic KV-aware router service for Dynamo deployments. It is useful when worker selection should live outside the HTTP frontend, such as manual disaggregated serving, multi-tier architectures, or custom request pipelines that need to route to a specific worker endpoint.
For the routing cost model and worker-selection behavior, see Routing Concepts. For shared tuning guidance, see Configuration and Tuning.
When to use it
Use the frontend-embedded router for the normal HTTP serving path. It owns /v1/chat/completions, discovers backend workers, and routes requests inside the frontend process.
Use the standalone router when you need a dedicated routing component that targets a specific Dynamo endpoint. The standalone router does not provide the OpenAI-compatible HTTP frontend; clients call the router service through the Dynamo runtime.
Start the router
Run the router with the target worker endpoint. The endpoint uses the namespace.component.endpoint format.
Required option:
--endpoint: Full endpoint path for workers, such asdynamo.prefill.generate.
Standalone-specific options:
--router-block-size: KV block size used for routing decisions. Set this explicitly so it matches the backend worker block size.--no-router-track-active-blocks: Disables decode active-block tracking. This is commonly useful for prefill-only routing because decode load is not relevant.
Most KV tuning options use the --router-* prefix, but shared options such as --load-aware, --serve-indexer, --use-remote-indexer, and --shared-cache-* do not. Legacy names such as --block-size and --kv-events are still accepted but deprecated. Run python -m dynamo.router --help for the current command surface.
Runtime endpoints
The standalone router exposes three Dynamo runtime endpoints:
generate: Routes requests to the best worker and streams generation results.best_worker_id: Returns the best worker ID for the provided token IDs without forwarding the request.get_overlap_scores: Returns per-worker and per-DP-rank matched block counts for device, host-pinned, disk, and shared-cache tiers.
Clients can use generate when the router should forward the request, best_worker_id when an external scheduler wants to make the final call itself, or get_overlap_scores when the scheduler needs the raw tiered overlap signal.
Manual prefill routing example
This is an advanced alternative setup. For most disaggregated serving deployments, use the frontend’s automatic prefill routing, which activates when workers register with WorkerType.Prefill. See Disaggregated Serving.
Use a standalone router when you need explicit control over prefill routing configuration or want to manage prefill and decode routers separately.
For an integrated frontend disaggregated example, see examples/backends/vllm/launch/disagg_router.sh. For explicit multi-router composition, see the Global Router README.
For event-driven prefix-cache state, add the backend-specific KV event publishing flags to the workers that the router indexes. Use --no-router-kv-events on the router only when approximate cache-state prediction is acceptable.
Configuration guidance
Set --router-block-size on standalone routers. Unlike frontend-embedded routing, standalone routers do not infer block size from the ModelDeploymentCard during worker registration.
The block size must match across the standalone router and all target workers. For vLLM, that means --router-block-size on the router should match --block-size on the vLLM workers.
The endpoint must match where the target workers register:
- vLLM prefill workers:
dynamo.prefill.generate - vLLM decode workers:
dynamo.backend.generate - Custom workers:
<your_namespace>.<your_component>.<your_endpoint>
To integrate a backend with the standalone router:
- Register workers at the endpoint passed to
--endpoint. - Call
router.generateto let the router select and forward to a worker, or callrouter.best_worker_idand then send the request directly. - Let router state update as requests are routed; no separate free call is required.
See components/src/dynamo/vllm/handlers.py for a backend reference implementation.