> For clean Markdown content of this page, append .md to this URL. For the complete documentation index, see https://docs.nvidia.com/dynamo/llms.txt. For full content including API reference and SDK examples, see https://docs.nvidia.com/dynamo/llms-full.txt.

# vLLM Configuration (DynamoVllmConfig)

`DynamoVllmConfig` holds the Dynamo-specific configuration for the vLLM backend (`python -m dynamo.vllm`). Every field, type, default, and choice on this page comes from the [`DynamoVllmArgGroup` and `DynamoVllmConfig`](https://github.com/ai-dynamo/dynamo/blob/main/components/src/dynamo/vllm/backend_args.py) definitions. For features and operational details, see the [vLLM Reference Guide](/dynamo/dev/knowledge-base/modular-components/backends/v-llm/reference-guide).

These are **only** the Dynamo wrapper flags. The vLLM backend also accepts every native vLLM `EngineArgs` argument (`--model`, `--tensor-parallel-size`, `--max-model-len`, and so on) in the same command, plus the cross-cutting [Dynamo Runtime](/dynamo/dev/reference/components/runtime-configuration) flags (`--namespace`, `--endpoint`, and others). Except for the native KV event and KV transfer sections below, this page covers neither — only the vLLM-specific `DYN_VLLM_*` surface.

## How the config is loaded

Each field is both a CLI flag and an environment variable. The CLI flag takes precedence; the environment variable is the fallback. Boolean fields are negatable — `--headless` sets it on, `--no-headless` sets it off.

#### Kubernetes

Set flags in the worker container's `args` and environment variables in its `env`, under the vLLM worker service of a [DynamoGraphDeployment](/dynamo/dev/reference/api/kubernetes/dynamo-graph-deployment) (DGD). The operator passes both through to the process without validation.

```yaml
spec:
  services:
    VllmDecodeWorker:
      extraPodSpec:
        mainContainer:
          command:
            - python3
            - -m
            - dynamo.vllm
          args:
            - --model
            - meta-llama/Llama-3.1-8B-Instruct
            - --disaggregation-mode
            - decode
          env:
            - name: DYN_VLLM_EMBEDDING_TRANSFER_MODE
              value: nixl-read
```

#### Local

Pass the flags directly on the command line:

```bash
python -m dynamo.vllm \
    --model meta-llama/Llama-3.1-8B-Instruct \
    --disaggregation-mode decode \
    --embedding-transfer-mode nixl-read
```

## Native KV event configuration

`--kv-events-config` and `--enable-prefix-caching` are native vLLM engine arguments rather than `DynamoVllmConfig` fields, but they determine whether a worker publishes the cache events used by event-driven KV-aware routing.

Starting the frontend with `--router-mode kv` does not configure event publishing on vLLM workers. Enable publishing explicitly on every aggregated or prefill worker whose cache state the router should track.

```bash
python -m dynamo.vllm \
  --model Qwen/Qwen3-0.6B \
  --enable-prefix-caching \
  --kv-events-config '{"enable_kv_cache_events":true,"publisher":"zmq","topic":"kv-events","endpoint":"tcp://*:5557"}'
```

The `endpoint` value is the base ZeroMQ port. vLLM assigns each data-parallel rank the port at `base port + data-parallel rank`. When multiple workers share a host or network namespace, reserve one port per rank and choose base ports whose resulting ranges do not overlap.

If workers do not publish KV events, configure the frontend with `--no-router-kv-events` for prediction-based KV routing or `--load-aware` for load-only routing.

## Native KV transfer configuration

`--kv-transfer-config` is a native vLLM engine argument rather than a `DynamoVllmConfig` field, so it has no `DYN_VLLM_*` environment variable. It selects the KV connector that moves cache blocks between prefill and decode workers.

A worker started with `--disaggregation-mode prefill` must be passed `--kv-transfer-config` explicitly. Without it, the worker raises a `ValueError` during argument parsing and never starts. All non-prefill modes — `agg`, `pd`, `decode`, and `encode` — do not enforce this check.

The value is a JSON object. For NIXL-based prefill/decode disaggregation:

```bash
python -m dynamo.vllm \
  --model Qwen/Qwen3-0.6B \
  --disaggregation-mode prefill \
  --kv-transfer-config '{"kv_connector":"NixlConnector","kv_role":"kv_both"}'
```

Only the prefill worker is required to set it, but both halves of a NIXL pair must agree on a connector for transfers to succeed. Pass the same `--kv-transfer-config` value to the decode worker, as the [disaggregated vLLM launch script](https://github.com/ai-dynamo/dynamo/blob/main/examples/backends/vllm/launch/disagg.sh) does.

The earlier `--connector` flag is no longer accepted by the vLLM backend. Setting it — on the command line or through the `DYN_CONNECTOR` environment variable — raises a `ValueError` during argument parsing. The message depends on the value:

* An active connector, such as `--connector nixl` or `DYN_CONNECTOR=nixl`, reports the equivalent `--kv-transfer-config` JSON to use instead.
* `--connector none` or `--connector null` reports that the flag is no longer needed, because no connector is already the default. There is no equivalent value to migrate to, so none is shown.
* `DYN_CONNECTOR` set to an empty or whitespace-only value reports that the variable is no longer supported, without an equivalent value.

## Worker role and disaggregation

These flags control which role this worker plays in a disaggregated deployment. The default when no `--disaggregation-mode` is set is aggregated (`agg`).

**`--disaggregation-mode`** `string` — default: null

Worker disaggregation mode. `agg` (default when unset) runs a combined aggregated prefill+decode worker. `pd` is a legacy alias for `agg`. `prefill` and `decode` split the pipeline for prefill/decode disaggregation. `encode` starts a multimodal encode-only worker.

`prefill` additionally requires the native `--kv-transfer-config` argument — see [Native KV transfer configuration](#native-kv-transfer-configuration).

Allowed values:

pd

agg

prefill

decode

encode

Environment variable: `DYN_VLLM_DISAGGREGATION_MODE`

---

**`--headless`** `boolean` — default: false

Run in headless mode for multi-node tensor-parallel or pipeline-parallel deployments. Secondary nodes run vLLM workers only with no Dynamo endpoints. See the vLLM multi-node data parallel documentation for details.

Environment variable: `DYN_VLLM_HEADLESS`

---

## Tokenizer and multimodal

**`--use-vllm-tokenizer`** `boolean` — default: false

Use vLLM's tokenizer for pre- and post-processing. This bypasses Dynamo's preprocessor; only the `/v1/chat/completions` endpoint will be available through the Dynamo frontend.

Environment variable: `DYN_VLLM_USE_TOKENIZER`

---

**`--enable-multimodal`** `boolean` — default: false

Enable multimodal processing. Combine it with `--disaggregation-mode=encode`, `--disaggregation-mode=pd`, `--disaggregation-mode=prefill`, or `--disaggregation-mode=decode` to select a disaggregated multimodal role. Use the default `agg` mode for aggregated multimodal serving.

Environment variable: `DYN_VLLM_ENABLE_MULTIMODAL`

---

**`--route-to-encoder`** `boolean` — default: false

Enable routing to separate encoder workers for multimodal processing.

Environment variable: `DYN_VLLM_ROUTE_TO_ENCODER`

---

**`--mm-prompt-template`** `string` — default: USER: \<image>\n\<prompt> ASSISTANT:

Prompt template used to construct the final multimodal prompt sent to the model. The literal `<prompt>` token is replaced with the user's text at inference time; `<image>` marks where the image placeholder appears. Update this template to match your model's expected format.

Environment variable: `DYN_VLLM_MM_PROMPT_TEMPLATE`

---

**`--frontend-decoding`** `boolean` — default: false

Enable frontend decoding of multimodal images. Images are decoded in the Rust frontend and transferred to the backend via NIXL RDMA, bypassing in-engine HTTP fetch and decode.

Environment variable: `DYN_VLLM_FRONTEND_DECODING`

---

## Embedding

**`--embedding-transfer-mode`** `string` — default: nixl-write

Embedding transfer mode used between encode and decode workers. `local` keeps embeddings on the local file system. `nixl-write` and `nixl-read` transfer them over NIXL RDMA (writer-initiated or reader-initiated, respectively).

Allowed values:

local

nixl-write

nixl-read

Environment variable: `DYN_VLLM_EMBEDDING_TRANSFER_MODE`

---

**`--embedding-worker`** `boolean` — default: false

Run as a text-embedding worker. The vLLM engine must be started with `--runner pooling`. KV-event publishing, KV router registration, and `InstrumentedScheduler` injection are all skipped, as they do not apply to pooling models.

Environment variable: `DYN_VLLM_EMBEDDING_WORKER`

---

## Other

**`--enable-rl`** `boolean` — default: false

Enable reinforcement-learning training support. Selects RL-friendly vLLM defaults for token-in/token-out (TITO) workloads and per-token logprob parity. Mirrors `--enable-rl` on the SGLang backend.

Environment variable: `DYN_ENABLE_RL`

---

**`--gms-shadow-mode`** `boolean` — default: false

Enable GMS (GPU Memory Service) shadow/standby mode. Shadow engines skip KV cache allocation at startup, automatically pause after initialization, and resume on demand when the active engine dies. Requires `--load-format=gms`.

Environment variable: `DYN_VLLM_GMS_SHADOW_MODE`

---

## Benchmarking

These flags control the self-benchmark sweep that runs on startup before the worker begins accepting production requests.

**`--benchmark-mode`** `string` — default: null

Run a self-benchmark on startup before accepting requests. Sweeps prefill input sequence lengths (ISLs) and/or decode `(context_length × batch_size)` operating points, collecting `ForwardPassMetrics` at each point.

Allowed values:

prefill

decode

agg

Environment variable: `DYN_BENCHMARK_MODE`

---

**`--benchmark-prefill-granularity`** `integer` — default: 16

Number of ISL sample points for the prefill sweep.

Environment variable: `DYN_BENCHMARK_PREFILL_GRANULARITY`

---

**`--benchmark-decode-length-granularity`** `integer` — default: 6

Number of context length sample points for the decode sweep.

Environment variable: `DYN_BENCHMARK_DECODE_LENGTH_GRANULARITY`

---

**`--benchmark-decode-batch-granularity`** `integer` — default: 6

Number of batch size sample points per context length for the decode sweep.

Environment variable: `DYN_BENCHMARK_DECODE_BATCH_GRANULARITY`

---

**`--benchmark-warmup-iterations`** `integer` — default: 5

Number of warmup iterations to run before benchmark measurement begins.

Environment variable: `DYN_BENCHMARK_WARMUP_ITERATIONS`

---

**`--benchmark-output-path`** `string` — default: /tmp/benchmark\_results.json

File path where benchmark results are written in JSON format.

Environment variable: `DYN_BENCHMARK_OUTPUT_PATH`

---

**`--benchmark-timeout`** `integer` — default: 300

Maximum seconds to wait for the benchmark to complete. Worker startup fails if this limit is exceeded.

Environment variable: `DYN_BENCHMARK_TIMEOUT`

---

## Deprecated

These flags are retained for backward compatibility and will be removed in a future release. Each is mapped to its replacement at startup with a deprecation warning.

**`--model-express-url`** `string` — deprecated, default: null

**Deprecated** — accepted for compatibility with older ModelExpress manifests only. The vLLM ModelExpress plugin reads its own configuration.

Environment variable: `MODEL_EXPRESS_URL`

---

## Validation rules

* `--embedding-worker` is only valid with `--disaggregation-mode=agg` (or the default aggregated mode) and cannot be combined with `--enable-multimodal` or `--benchmark-mode`.

## Related pages

#### [vLLM Reference Guide](/dynamo/dev/knowledge-base/modular-components/backends/v-llm/reference-guide)

Features, worker roles, and operational details for the vLLM backend.

#### [Runtime Configuration](/dynamo/dev/reference/components/runtime-configuration)

Cross-cutting `DYN_*` flags shared by every backend and the frontend.

#### [SGLang Configuration](/dynamo/dev/reference/backends/sg-lang-configuration)

The equivalent Dynamo flag reference for the SGLang backend.

#### [TensorRT-LLM Configuration](/dynamo/dev/reference/backends/tensor-rt-llm-configuration)

The equivalent Dynamo flag reference for the TensorRT-LLM backend.