Mocker CLI Reference
Command-line flags for live Mocker workers
python3 -m dynamo.mocker launches a simulated Dynamo worker that registers with the frontend,
publishes KV events, and exercises the router and planner paths without a GPU. This page is the flag
reference. For the deployment workflow and local/Kubernetes launch recipes, see
Simulate a Kubernetes Deployment or
Simulate a Local Deployment; for the engine internals these flags
configure, see Mocker Engine Architecture.
Run python3 -m dynamo.mocker --help for the complete option list supported by the installed
version.
AISimulate reuses the same engine core for offline virtual-clock prediction and recommendation, but its public YAML is separate from these live-worker flags. See the DynoSim Replay CLI Reference for the offline contract. The former public Replay online CLI remains unavailable, although the Python replay SDK retains online mode.
Core and model
Print the Dynamo Mocker version and exit.
Hugging Face model ID or local path for the tokenizer. Set this for normal frontend requests.
Model name used in API responses.
Dynamo endpoint string. Defaults are namespace-dependent, and prefill workers use a different default endpoint than aggregated or decode workers.
Engine simulation type.
Allowed values: vllm sglang trtllmPath to a JSON file with mocker configuration. Overrides individual CLI arguments.
KV cache
Usable KV cache blocks per data-parallel rank. Non-AIC timing defaults to 16,384 blocks; AIC timing estimates capacity when this option is omitted.
Tokens per KV cache block. Defaults are 64 for vLLM, 1 for SGLang, and 32 for TensorRT-LLM.
Maximum sequence length, including prompt and generated tokens. Omitting the option leaves the simulated engine without a model-length limit.
Enable prefix caching. Pass --no-enable-prefix-caching to disable.
KV cache dtype for the bytes-per-token computation.
KV cache bytes per token. Overrides the auto-computation.
Scheduling
Maximum concurrent sequences.
Maximum tokens per batch.
Enable chunked prefill. Pass --no-enable-chunked-prefill to disable.
Decode eviction policy under memory pressure. lifo is vLLM v1 style.
Timing
Timing speedup factor applied to simulated prefill and decode durations.
Decode-only speedup multiplier, for example for Eagle speculation.
Simulated startup delay, in seconds.
Data parallelism and workers
Number of DP replicas.
Workers per process. Prefer this over launching many separate mocker processes: all workers share one tokio runtime and thread pool.
Delay between worker launches, in seconds. 0 disables; -1 enables auto mode.
Performance modeling
Path to either a mocker-format .npz file or a profiler results directory.
JSON config for emitting reasoning token spans, with start_thinking_token_id,
end_thinking_token_id, and thinking_ratio.
Mooncake JSONL trace whose output_token_ids provide response replay annotations.
AIC performance model
Opt-in flags for the AIConfigurator (AIC) latency model. Live Mocker workers do not use AIC by default. For the timing-model design, see Mocker Engine Architecture.
Use the AIC compatibility API in the AISimulate wheel for latency prediction instead of the
interpolated or polynomial models. Install aisimulate and use a supported
system/backend/version tuple.
AIC system name used with --aic-perf-model. The raw CLI value is unset when omitted; runtime
resolution uses h200_sxm when AIC modeling is enabled.
Backend used for AIC performance-data lookup. Set it only to model timing for a different backend than the simulated scheduler.
Allowed values: vllm sglang trtllmAIC backend engine version, for example 0.12.0 for vLLM. When unset, uses the default version for
the backend.
Tensor-parallel size for AIC latency prediction. Affects only AIC performance-model lookups, not mocker scheduling.
Mixture-of-Experts tensor-parallel size for AIC latency prediction. Required by some MoE models.
Mixture-of-Experts expert-parallel size for AIC latency prediction. Required by some MoE models.
Attention data-parallel size for AIC latency prediction. Required by some MoE models.
vLLM GPU memory fraction used for AIC KV-capacity estimation.
SGLang static memory fraction used for AIC KV-capacity estimation.
TensorRT-LLM fraction of post-model-load free GPU memory used for AIC KV-capacity estimation.
Experimental Multi-Token Prediction (MTP) draft length, from 1 through 5.
Experimental comma-separated conditional MTP acceptance rates.
Base random seed for Mocker MTP burst sampling.
Disaggregation
Worker mode.
Allowed values: agg prefill decodeDeprecated alias for --disaggregation-mode=prefill. Do not combine it with
--disaggregation-mode or --is-decode-worker.
Deprecated alias for --disaggregation-mode=decode. Do not combine it with
--disaggregation-mode or --is-prefill-worker.
Comma-separated rendezvous base ports, one per worker in disaggregated mode.
KV cache transfer bandwidth in GB/s. Set to 0 to disable.
Prompt footprint charged to a coordinated prefill/decode transfer.
Allowed values: full_prompt destination_missingEvent and request transport
Event transport. When the environment does not set a value, Mocker uses ZMQ with file or memory discovery and NATS with etcd or Kubernetes discovery.
Allowed values: nats zmqRequest transport.
Allowed values: nats tcpDiscovery backend.
Allowed values: kubernetes etcd file memComma-separated ZMQ PUB base ports for KV event publishing, one per worker.
Comma-separated ZMQ ROUTER base ports for gap recovery, one per worker.
SGLang-specific
Apply only when --engine-type sglang.
SGLang scheduling policy. fifo/fcfs is the default; lpm is longest prefix match.
SGLang radix-cache page size in tokens. Also becomes the effective block size when
--engine-type sglang and --block-size is omitted.
SGLang max prefill-token budget per batch.
SGLang chunked-prefill chunk size.
SGLang admission-budget cap for max new tokens.
SGLang schedule conservativeness factor.
TensorRT-LLM-specific
Apply only when --engine-type trtllm.
TensorRT-LLM capacity scheduler policy. The Mocker currently supports only
guaranteed_no_evict.
Router advertisement
Mocker accepts the same --router-* worker-advertisement options as the real backend launchers.
Use them to override Router behavior for this worker set. See
Router Configuration and Tuning
for the shared fields.
Advertise a Router configuration for this worker set. When omitted, the worker advertises no Router configuration and inherits the Frontend configuration.
Allowed values: round-robin random power-of-two kv direct least-loaded device-aware-weightedEnvironment variables
Related pages
Deploy and run the Mocker worker on Kubernetes.
Run a frontend and Mocker worker from the command line.
Run one workload through an offline simulated configuration with aisimulate predict --stack dynamo.
Scheduler, KV block manager, eviction, and timing internals.
The PlannerConfig field reference for the Dynamo Planner service.