Mocker CLI Reference
Command-line flags for live Mocker workers
python3 -m dynamo.mocker launches a simulated Dynamo worker that registers with the frontend,
publishes KV events, and exercises the router and planner paths without a GPU. This page is the flag
reference. For the deployment workflow and local/Kubernetes launch recipes, see
Simulate a Kubernetes Deployment or
Simulate a Local Deployment; for the engine internals these flags
configure, see Mocker Engine Architecture.
Run python3 -m dynamo.mocker --help for the complete option list supported by the installed
version.
AISimulate reuses the same engine core for offline virtual-clock prediction and recommendation, but its public YAML is separate from these live-worker flags. See the DynoSim Replay CLI Reference for the offline contract. The former public Replay online CLI remains unavailable, although the Python replay SDK retains online mode.
Core and model
Print the Dynamo Mocker version and exit.
Hugging Face model ID or local path for the tokenizer. Set this for normal frontend requests.
Model name used in API responses.
Dynamo endpoint string. Defaults are namespace-dependent, and prefill workers use a different default endpoint than aggregated or decode workers.
Engine simulation type.
Allowed values: vllm sglang trtllmPath to a JSON file with mocker configuration. Overrides individual CLI arguments.
KV cache
Usable KV cache blocks per data-parallel rank. Non-AIS timing defaults to 16,384 blocks; AIS timing estimates capacity when this option is omitted.
Tokens per KV cache block. Defaults are 64 for vLLM, 1 for SGLang, and 32 for TensorRT-LLM.
Maximum sequence length, including prompt and generated tokens. Omitting the option leaves the simulated engine without a model-length limit.
Enable prefix caching. Pass --no-enable-prefix-caching to disable.
KV cache dtype for the bytes-per-token computation.
KV cache bytes per token. Overrides the auto-computation.
Scheduling
Maximum concurrent sequences.
Maximum tokens per batch.
Enable chunked prefill. Pass --no-enable-chunked-prefill to disable.
Decode eviction policy under memory pressure. lifo is vLLM v1 style.
Timing
Timing speedup factor applied to simulated prefill and decode durations.
Decode-only speedup multiplier, for example for Eagle speculation.
Simulated startup delay, in seconds.
Data parallelism and workers
Number of DP replicas.
Workers per process. Prefer this over launching many separate mocker processes: all workers share one tokio runtime and thread pool.
Delay between worker launches, in seconds. 0 disables staggering (default); -1 enables auto mode.
Performance modeling
Path to either a mocker-format .npz file or a profiler results directory.
JSON config for emitting reasoning token spans, with start_thinking_token_id,
end_thinking_token_id, and thinking_ratio.
Mooncake JSONL trace whose output_token_ids provide response replay annotations.
AIS performance model
Complete AISimulate ForwardPassPerfModelConfig, including immutable worker role,
data roots, selection policy and nested estimator controls. Defaults to
estimation_mode: auto and fallback_policy: deny. Do not combine this input
with flat AIS identity flags. Cold regression is rejected by Router and Mocker.
Mocker and Dynamo Replay require pp: 1; pipeline-parallel configurations are rejected.
Legacy --aic-* spellings remain input aliases. Use --ais-* in new configurations.
Opt-in flags for the AISimulate (AIS) latency model. Live Mocker workers do not use AIS by default. For the timing-model design, see Mocker Engine Architecture.
Use the canonical AISimulate performance-model API for latency prediction instead of the
interpolated or polynomial models. Install aisimulate and use a supported
system/backend/version tuple.
AIS system name used with --ais-perf-model. The raw CLI value is unset when omitted; runtime
resolution uses h200_sxm when AIS modeling is enabled.
Backend used for AIS performance-data lookup. Set it only to model timing for a different backend than the simulated scheduler.
Allowed values: vllm sglang trtllmAIS performance-database version. Use current, previous, or next when the slot is available
for the selected system and backend, or specify a version assigned to one of those slots.
When unset, uses the release database’s current slot.
Tensor-parallel size for AIS latency prediction. Affects only AIS performance-model lookups, not mocker scheduling.
Mixture-of-Experts tensor-parallel size for AIS latency prediction. Required by some MoE models.
Mixture-of-Experts expert-parallel size for AIS latency prediction. Required by some MoE models.
Attention data-parallel size for AIS latency prediction. Required by some MoE models.
vLLM GPU memory fraction used for AIS KV-capacity estimation.
SGLang static memory fraction used for AIS KV-capacity estimation.
TensorRT-LLM fraction of post-model-load free GPU memory used for AIS KV-capacity estimation.
Experimental Multi-Token Prediction (MTP) draft length, from 1 through 5.
Experimental comma-separated conditional MTP acceptance rates.
Base random seed for Mocker MTP burst sampling.
Disaggregation
Worker mode.
Allowed values: agg prefill decodeDeprecated alias for --disaggregation-mode=prefill. Do not combine it with
--disaggregation-mode or --is-decode-worker.
Deprecated alias for --disaggregation-mode=decode. Do not combine it with
--disaggregation-mode or --is-prefill-worker.
Comma-separated rendezvous base ports, one per worker in disaggregated mode.
KV cache transfer bandwidth in GB/s. Set to 0 to disable.
Prompt footprint charged to a coordinated prefill/decode transfer.
Allowed values: full_prompt destination_missingEvent and request transport
Event transport. When the environment does not set a value, Mocker uses ZMQ with file or memory discovery and NATS with etcd or Kubernetes discovery.
Allowed values: nats zmqRequest transport.
Allowed values: nats tcpDiscovery backend.
Allowed values: kubernetes etcd file memComma-separated ZMQ PUB base ports for KV event publishing, one per worker.
Comma-separated ZMQ ROUTER base ports for gap recovery, one per worker.
SGLang-specific
Apply only when --engine-type sglang.
SGLang scheduling policy. fifo/fcfs is the default; lpm is longest prefix match.
SGLang radix-cache page size in tokens. Also becomes the effective block size when
--engine-type sglang and --block-size is omitted.
SGLang max prefill-token budget per batch.
SGLang chunked-prefill chunk size.
SGLang admission-budget cap for max new tokens.
SGLang schedule conservativeness factor.
TensorRT-LLM-specific
Apply only when --engine-type trtllm.
TensorRT-LLM capacity scheduler policy. The Mocker currently supports only
guaranteed_no_evict.
Router advertisement
Mocker accepts the same --router-* worker-advertisement options as the real backend launchers.
Use them to override Router behavior for this worker set. See
Router Configuration and Tuning
for the shared fields.
Advertise a Router configuration for this worker set. When omitted, the worker advertises no Router configuration and inherits the Frontend configuration.
Allowed values: round-robin random power-of-two kv direct least-loaded device-aware-weighted