> For clean Markdown content of this page, append .md to this URL. For the complete documentation index, see https://docs.nvidia.com/dynamo/llms.txt. For full content including API reference and SDK examples, see https://docs.nvidia.com/dynamo/llms-full.txt.

# Mocker CLI Reference

`python3 -m dynamo.mocker` launches a simulated Dynamo worker that registers with the frontend,
publishes KV events, and exercises the router and planner paths without a GPU. This page is the flag
reference. For the deployment workflow and local/Kubernetes launch recipes, see
[Simulate a Kubernetes Deployment](/dynamo/knowledge-base/concepts/simulation/kubernetes-mocker) or
[Simulate a Local Deployment](/dynamo/knowledge-base/concepts/simulation/local-mocker); for the engine internals these flags
configure, see [Mocker Engine Architecture](/dynamo/knowledge-base/modular-components/backends/mocker/mocker-engine-architecture).

Run `python3 -m dynamo.mocker --help` for the complete option list supported by the installed
version.

AISimulate reuses the same engine core for offline virtual-clock prediction and recommendation, but
its public YAML is separate from these live-worker flags. See the
[DynoSim Replay CLI Reference](/dynamo/reference/components/dyno-sim-replay-cli-reference) for the offline contract. The
former public Replay online CLI remains unavailable, although the Python replay SDK retains online
mode.

## Core and model

**`--version`** `boolean` — default: false

Print the Dynamo Mocker version and exit.

---

**`--model-path`** `string` — default: null

Hugging Face model ID or local path for the tokenizer. Set this for normal frontend requests.

---

**`--model-name`** `string` — default: derived from --model-path

Model name used in API responses.

---

**`--endpoint`** `string` — default: auto-derived

Dynamo endpoint string. Defaults are namespace-dependent, and prefill workers use a different
default endpoint than aggregated or decode workers.

---

**`--engine-type`** `string` — default: vllm

Engine simulation type.

Allowed values:

vllm

sglang

trtllm

---

**`--extra-engine-args`** `path` — default: null

Path to a JSON file with mocker configuration. Overrides individual CLI arguments.

---

## KV cache

**`--num-gpu-blocks-override`** `integer` — default: auto

Usable KV cache blocks per data-parallel rank. Non-AIC timing defaults to 16,384 blocks; AIC
timing estimates capacity when this option is omitted.

---

**`--block-size`** `integer` — default: engine-specific

Tokens per KV cache block. Defaults are 64 for vLLM, 1 for SGLang, and 32 for TensorRT-LLM.

---

**`--max-model-len`** `integer` — default: null

Maximum sequence length, including prompt and generated tokens. Omitting the option leaves the
simulated engine without a model-length limit.

---

**`--enable-prefix-caching`** `boolean` — default: true

Enable prefix caching. Pass `--no-enable-prefix-caching` to disable.

---

**`--kv-cache-dtype`** `string` — default: auto

KV cache dtype for the bytes-per-token computation.

---

**`--kv-bytes-per-token`** `integer` — default: auto-computed

KV cache bytes per token. Overrides the auto-computation.

---

## Scheduling

**`--max-num-seqs`** `integer` — default: 256

Maximum concurrent sequences.

---

**`--max-num-batched-tokens`** `integer` — default: 8192

Maximum tokens per batch.

---

**`--enable-chunked-prefill`** `boolean` — default: true

Enable chunked prefill. Pass `--no-enable-chunked-prefill` to disable.

---

**`--preemption-mode`** `string` — default: lifo

Decode eviction policy under memory pressure. `lifo` is vLLM v1 style.

Allowed values:

lifo

fifo

---

## Timing

**`--speedup-ratio`** `float` — default: 1.0

Timing speedup factor applied to simulated prefill and decode durations.

---

**`--decode-speedup-ratio`** `float` — default: 1.0

Decode-only speedup multiplier, for example for Eagle speculation.

---

**`--startup-time`** `float` — default: null

Simulated startup delay, in seconds.

---

## Data parallelism and workers

**`--data-parallel-size`** `integer` — default: 1

Number of DP replicas.

---

**`--num-workers`** `integer` — default: 1

Workers per process. Prefer this over launching many separate mocker processes: all workers share
one tokio runtime and thread pool.

---

**`--stagger-delay`** `float` — default: -1 (auto)

Delay between worker launches, in seconds. `0` disables; `-1` enables auto mode.

---

## Performance modeling

**`--planner-profile-data`** `path` — default: null

Path to either a mocker-format `.npz` file or a profiler results directory.

---

**`--reasoning`** `JSON string` — default: null

JSON config for emitting reasoning token spans, with `start_thinking_token_id`,
`end_thinking_token_id`, and `thinking_ratio`.

---

**`--response-replay-trace-path`** `path` — default: null

Mooncake JSONL trace whose `output_token_ids` provide response replay annotations.

---

## AIC performance model

Opt-in flags for the AIConfigurator (AIC) latency model. Live Mocker workers do not use AIC by
default.
For the timing-model design, see [Mocker Engine Architecture](/dynamo/knowledge-base/modular-components/backends/mocker/mocker-engine-architecture#performance-model).

**`--aic-perf-model`** `boolean` — default: false

Use the AIC compatibility API in the AISimulate wheel for latency prediction instead of the
interpolated or polynomial models. Install `aisimulate` and use a supported
`system/backend/version` tuple.

---

**`--aic-system`** `string` — default: h200\_sxm when AIC is enabled

AIC system name used with `--aic-perf-model`. The raw CLI value is unset when omitted; runtime
resolution uses `h200_sxm` when AIC modeling is enabled.

---

**`--aic-backend`** `string` — default: --engine-type

Backend used for AIC performance-data lookup. Set it only to model timing for a different backend
than the simulated scheduler.

Allowed values:

vllm

sglang

trtllm

---

**`--aic-backend-version`** `string` — default: auto

AIC performance-database version. Use `current`, `previous`, or `next` when the slot is available
for the selected system and backend, or specify a version assigned to one of those slots.
When unset, uses the release database's `current` slot.

---

**`--aic-tp-size`** `integer` — default: 1

Tensor-parallel size for AIC latency prediction. Affects only AIC performance-model lookups, not
mocker scheduling.

---

**`--aic-moe-tp-size`** `integer` — default: null

Mixture-of-Experts tensor-parallel size for AIC latency prediction. Required by some MoE models.

---

**`--aic-moe-ep-size`** `integer` — default: null

Mixture-of-Experts expert-parallel size for AIC latency prediction. Required by some MoE models.

---

**`--aic-attention-dp-size`** `integer` — default: null

Attention data-parallel size for AIC latency prediction. Required by some MoE models.

---

**`--gpu-memory-utilization`** `float` — default: 0.9

vLLM GPU memory fraction used for AIC KV-capacity estimation.

---

**`--mem-fraction-static`** `float` — default: 0.88

SGLang static memory fraction used for AIC KV-capacity estimation.

---

**`--free-gpu-memory-fraction`** `float` — default: 0.9

TensorRT-LLM fraction of post-model-load free GPU memory used for AIC KV-capacity estimation.

---

**`--aic-nextn`** `integer` — default: null

Experimental Multi-Token Prediction (MTP) draft length, from 1 through 5.

---

**`--aic-nextn-accept-rates`** `string` — default: null

Experimental comma-separated conditional MTP acceptance rates.

---

**`--aic-mtp-seed`** `integer` — default: 42

Base random seed for Mocker MTP burst sampling.

---

## Disaggregation

**`--disaggregation-mode`** `string` — default: agg

Worker mode.

Allowed values:

agg

prefill

decode

---

**`--is-prefill-worker`** `boolean` — default: false

Deprecated alias for `--disaggregation-mode=prefill`. Do not combine it with
`--disaggregation-mode` or `--is-decode-worker`.

---

**`--is-decode-worker`** `boolean` — default: false

Deprecated alias for `--disaggregation-mode=decode`. Do not combine it with
`--disaggregation-mode` or `--is-prefill-worker`.

---

**`--bootstrap-ports`** `string` — default: null

Comma-separated rendezvous base ports, one per worker in disaggregated mode.

---

**`--kv-transfer-bandwidth`** `float` — default: 64.0

KV cache transfer bandwidth in GB/s. Set to `0` to disable.

---

**`--kv-transfer-timing-mode`** `string` — default: full\_prompt

Prompt footprint charged to a coordinated prefill/decode transfer.

Allowed values:

full\_prompt

destination\_missing

---

## Event and request transport

**`--event-plane`** `string` — default: auto

Event transport. When the environment does not set a value, Mocker uses ZMQ with file or memory
discovery and NATS with etcd or Kubernetes discovery.

Allowed values:

nats

zmq

---

**`--request-plane`** `string` — default: env-driven (tcp)

Request transport.

Allowed values:

nats

tcp

---

**`--discovery-backend`** `string` — default: env-driven (etcd)

Discovery backend.

Allowed values:

kubernetes

etcd

file

mem

---

**`--zmq-kv-events-ports`** `string` — default: null

Comma-separated ZMQ PUB base ports for KV event publishing, one per worker.

---

**`--zmq-replay-ports`** `string` — default: null

Comma-separated ZMQ ROUTER base ports for gap recovery, one per worker.

---

## SGLang-specific

Apply only when `--engine-type sglang`.

**`--sglang-schedule-policy`** `string` — default: fifo / fcfs

SGLang scheduling policy. `fifo`/`fcfs` is the default; `lpm` is longest prefix match.

Allowed values:

fifo

fcfs

lpm

---

**`--sglang-page-size`** `integer` — default: 1

SGLang radix-cache page size in tokens. Also becomes the effective block size when
`--engine-type sglang` and `--block-size` is omitted.

---

**`--sglang-max-prefill-tokens`** `integer` — default: 16384

SGLang max prefill-token budget per batch.

---

**`--sglang-chunked-prefill-size`** `integer` — default: 8192

SGLang chunked-prefill chunk size.

---

**`--sglang-clip-max-new-tokens`** `integer` — default: 4096

SGLang admission-budget cap for max new tokens.

---

**`--sglang-schedule-conservativeness`** `float` — default: 1.0

SGLang schedule conservativeness factor.

---

## TensorRT-LLM-specific

Apply only when `--engine-type trtllm`.

**`--trtllm-capacity-scheduler-policy`** `string` — default: guaranteed\_no\_evict

TensorRT-LLM capacity scheduler policy. The Mocker currently supports only
`guaranteed_no_evict`.

---

## Router advertisement

Mocker accepts the same `--router-*` worker-advertisement options as the real backend launchers.
Use them to override Router behavior for this worker set. See
[Router Configuration and Tuning](/dynamo/knowledge-base/modular-components/router/configuration-and-tuning)
for the shared fields.

**`--router-mode`** `string` — default: inherit frontend

Advertise a Router configuration for this worker set. When omitted, the worker advertises no
Router configuration and inherits the Frontend configuration.

Allowed values:

round-robin

random

power-of-two

kv

direct

least-loaded

device-aware-weighted

---

## Environment variables

| Variable                    | Default | Description                                                                            |
| --------------------------- | ------- | -------------------------------------------------------------------------------------- |
| `DYN_MOCKER_KV_CACHE_TRACE` | off     | Set to `1` or `true` to log structured native KV-cache allocation and eviction traces. |

## Related pages

#### [Simulate a Kubernetes Deployment](/dynamo/knowledge-base/concepts/simulation/kubernetes-mocker)

Deploy and run the Mocker worker on Kubernetes.

#### [Simulate a Local Deployment](/dynamo/knowledge-base/concepts/simulation/local-mocker)

Run a frontend and Mocker worker from the command line.

#### [Run a DynoSim Simulation](/dynamo/knowledge-base/concepts/simulation/simulation-runs)

Run one workload through an offline simulated configuration with `aisimulate predict --stack dynamo`.

#### [Mocker Engine Architecture](/dynamo/knowledge-base/modular-components/backends/mocker/mocker-engine-architecture)

Scheduler, KV block manager, eviction, and timing internals.

#### [Planner Configuration](/dynamo/reference/components/planner-configuration)

The PlannerConfig field reference for the Dynamo Planner service.