LLM Request Router Load Balancing
Use this guide to deploy and validate Stargate load-balancer configuration in
self-managed NVCF. For the complete lb-config.json schema, algorithm
behavior, defaults, and tuning fields, see the
Stargate load balancer configuration.
This guide covers NVCF deployment ownership and trusted request metadata. It does not redefine the Stargate schema.
Configure the self-managed stack
Set the request-router configuration in the Helmfile environment:
The example uses these settings in the model-a configuration:
Replace model-a with the exact routed model name. For requests with an
affinity key, this example keeps global buckets closed for 200 ms. Stargate can
select an available affine candidate immediately. Set cache_affinity_wait_ms
in the model configuration.
After the wait, global buckets include all backends, including the affinity group. Global selection uses full prefill cost for every backend and retains the queue-admission, capacity, and retry-exclusion checks. Stargate still tries affinity selection first on each routing attempt.
When an engine reports a positive concurrency limit and has a free request slot, Stargate estimates zero queue delay. It counts prefill, decode, and pending assignments against that limit. The request’s own prefill time still contributes to estimated TTFT. Pylon uses the same rule for its local queue admission check. Engines with an unknown limit retain the existing queue-delay estimate.
The self-managed stack passes
addons.llm.requestRouter.loadBalancer to the request-router chart as
llmRequestRouter.loadBalancer.
The chart supports two configuration sources:
Inline config takes precedence when both values are set. When neither value
is set, Stargate uses its built-in power-of-two default and accepts a
routing-method override when it is in the allowlist of built-in algorithms.
Stargate reads and validates the file only during process startup. The request-router workload does not include a load-balancer ConfigMap checksum in its pod template. New installations default to a Deployment; existing installations can pin a StatefulSet. After a ConfigMap-only update, restart the selected workload so every replica loads the same configuration.
Existing StatefulSet installations must set
addons.llm.requestRouter.workload.kind=StatefulSet before upgrading. Changing
the workload kind is a controlled migration, not an in-place Kubernetes kind
mutation. A plain Helm upgrade across workload kinds can temporarily run both
the Deployment and StatefulSet; use the chart migration procedure to remove or
rename the old workload and verify that only the selected kind owns the router
Pods before scaling it.
Distinguish router algorithms from nvcf-cli routing methods
Algorithm availability is enforced at separate layers:
For example, when a configuration is set and power-of-two is the effective
algorithm, wait_and_widen requires a wait-and-widen entry in
request_algorithms. The legacy groq_multiregion value remains accepted
and resolves to the same algorithm.
Use wait-and-widen and pulsar-wait-and-widen in new function metadata,
lb-config.json files, request-algorithm maps, and deployment manifests.
Existing groq-multiregion and pulsar-multiregion values continue to work
through the Stargate and control-plane compatibility aliases.
Keep router headers trusted
The gateway can send the following headers to Stargate. Derive or validate their values from authenticated function metadata and the request.
An explicit x-max-wait-ms shorter than the configured affinity wait can end
routing with HTTP 503 before global buckets become eligible. Set the
affinity wait within the latency budget for the workload.
The NVCF LLM invocation HTTPRoute does not strip the other router-facing headers. Do not expose the route to untrusted callers until a managed ingress policy removes them before the request reaches the LLM API Gateway.
For a separately managed Gateway API route, add this filter to the rule that
forwards to llm-api-gateway:
The gateway-routes chart does not expose a value for this filter. Use an
equivalent policy at an external edge or maintain a route override. Preserve
x-multi-turn-session-id; clients can use it for session affinity. Chat
Completions and Responses request bodies can also supply prompt_cache_key.
See the
Gateway API header modifier guide
for filter semantics.
For affinity routing, the gateway can supply x-cache-affinity-key as a stable
session or prefix identifier. When using prompt_cache_key, place its
SHA-256-derived value in the header. The request body can retain the raw value
for the model backend. Enable require_cache_affinity_key only when the
gateway supplies a key for every endpoint served by the model.
Stargate returns HTTP 400 for a blank, unknown, or configured-but-unavailable
x-routing-method. It also returns HTTP 400 when a required router header is
missing or a numeric header is invalid.
Apply and roll out
Render the chart before applying it:
Confirm that the rendered request-router workload (a Deployment by default) has
the expected --lb-config-path argument and that inline JSON creates one
ConfigMap with the lb-config.json key.
If the LLM route accepts untrusted traffic, inspect the rendered or live HTTPRoute and confirm that the trusted-header filter is present:
Apply the self-managed environment from the unpacked stack directory:
For an inline configuration, inspect the live file and start argument:
Replace deployment with statefulset in these commands when
addons.llm.requestRouter.workload.kind is pinned to StatefulSet.
Restart after a ConfigMap-only change, then wait for all replicas:
Use statefulset/llm-request-router instead when the StatefulSet workload is
explicitly selected.
Confirm every listed pod was recreated after the ConfigMap update. For each
pod, check for the load balancer config loaded startup log and compare its
reported default and model-override count:
Do not continue if a pod predates the update, lacks a successful load log, or reports a different configuration summary. Stable affinity requires the same mounted configuration, seed, and candidate view across replicas.
Validate routing
Use a deployed function whose routingMethod is the effective configured
algorithm or is present in request_algorithms.
- Invoke without changing the function routing method and confirm success.
- Update the function to a method present in
request_algorithmsand confirm success. - Try a method accepted by
nvcf-clithat is neither the configured algorithm nor present inrequest_algorithms; confirm that Stargate returns HTTP400. - For an affinity-aware method, repeat a supported multi-turn request with the
same
prompt_cache_keyor the returnedx-multi-turn-session-id. - Exercise a failed or saturated backend and confirm selection and retry counters change.
Observe the request router
The chart exposes llm-request-router:9090/metrics when request-router metrics
are enabled. The current chart passes --metrics-port and uses Stargate’s
default stargate_ metric prefix.
See LLM Request Router Metrics for metric names, labels, and scrape configuration.
Troubleshoot
Use these logs together:
Use statefulset/llm-request-router for the first command when the StatefulSet
workload is explicitly selected.