DynoSim Replay CLI Reference
aisimulate predict --stack dynamo evaluates one concrete workload and deployment configuration
through the Dynamo simulation stack. AISimulate owns the traffic and engine schema. The ai-dynamo
package supplies the Dynamo runner and the optional router and planner configuration adapters.
For an end-to-end workflow, see Run a DynoSim Simulation. To search configuration domains, see Sweep DynoSim Configurations.
The former Replay online mode has no replacement in the unified AISimulate CLI yet. aisimulate predict and aisimulate recommend are offline-only. The separate python3 -m dynamo.mocker
command remains available for launching live workers but does not provide replay orchestration.
Command
YAML prediction configuration. Prediction files contain concrete values and reject search
domains, presets, optimization, and optimizer.
Execution stack. Set this option to dynamo to load the Dynamo runner and the Router and Planner
adapters registered by ai-dynamo.
Override a schema-valid configuration path after loading the YAML file. Repeat the option for multiple overrides. Values are parsed as YAML, and later assignments win. Sequence indexes are not supported.
Directory for prediction.json and optional per-request output.
Replace known AISimulate output files in an existing nonempty output directory. Unrelated files are preserved.
Standard-output format.
Allowed values: table jsonWrite one record per request to requests.jsonl.
Configuration Sections
The YAML document uses these top-level sections:
Unknown fields are rejected. predict accepts only concrete values. Use aisimulate recommend for
choices, range, or preset domains.
Dynamo Prediction Example
Traffic Rules
traffic contains source, load, and stop mappings.
- Omitting the entire section creates 100 independent synthetic requests at concurrency 10, with 1,024 input tokens and 128 output tokens per request.
source.type: syntheticcreates independent requests withinput_tokensandoutput_tokens.source.type: synthetic-sessioncreates ordered multi-turn sessions.source.type: tracereads one or more paths. Supported formats includemooncake,mooncake-delta,agentic_mooncake,weka,applied_compute_agentic, anddynamo.load.type: concurrencykeeps a fixed number of requests or sessions active.load.type: constant_rateschedules evenly spaced arrivals fromrequests_per_secondorsessions_per_second.load.type: poissonuses the matching rate with exponential inter-arrival times and an optionalseed, which defaults to 42.load.type: trace_timestampspreserves trace timing and accepts a positivespeedup.stop.requestsapplies to independent requests.stop.sessionsapplies to session sources.stop.requests_per_load_unitandstop.sessions_per_load_unitderive the count from the concrete concurrency or arrival rate.stop.max_virtual_time_secondsis a soft virtual-time cutoff. Requests admitted before the cutoff may finish afterward.
For Dynamo request traces, omit traffic.source.block_size; Dynamo derives it from the source and
rejects mixed block sizes across shards. For Weka corpora, AISimulate derives the block size from the
source. An explicitly configured Weka block size is only an assertion against the published source
metadata. Dynamo format accepts multiple trace shards. Weka accepts one file or directory; other
formats accept exactly one file.
For Weka traces, traffic.source.nested_timestamp_basis controls how nested timestamps
are interpreted:
auto(the default) scans the entire corpus and selects one basis for all nested requests. If any replayable child timestamp precedes its parent marker, it selectsrelative; otherwise, it selectsabsolute. A corpus without replayable nested requests resolves tonot_applicable. It does not choose a separate basis for each child or file.absoluteuses timestamps directly and rejects a child timestamp before its parent marker, apart from the importer’s timestamp-rounding tolerance.relativeadds the parent marker’s timestamp to each child timestamp.
Explicit absolute and relative overrides take precedence over inference. For example, a marker
at 1 second with a child timestamp of 2 seconds is interpreted as 3 seconds with relative, and
as 2 seconds with absolute. auto chooses 2 seconds for that corpus unless another child
provides evidence for relative timestamps.
AISimulate’s Weka importer also changes two behaviors from the former Dynamo importer:
- A source request with
out: 0retains zero output tokens and runs as a prefill-only request. The former importer usedmax(1, out). Output-token counts, throughput, and timing metrics can therefore change for the same corpus. - A missing or null
api_timeremains absent (None) inrecorded_api_time_ms, preserving the distinction from an explicit zero. Dependency classification uses a zero-duration interval at the request timestamp when that duration is unknown. These requests were previously rejected. The fallback can affect inferred dependency edges and replay timing; it does not assert that the measured API duration was zero.
mooncake-delta, agentic_mooncake, and weka require aggregated engine mode.
The two agentic formats also require trace_timestamps and reject the virtual-time cutoff.
applied_compute_agentic requires concurrency load.
With the Dynamo stack, omit planner or set planner.policy: disabled for mooncake-delta,
agentic_mooncake, and Dynamo traces that carry agent_context records.
Typed agentic replay through the Dynamo API
AISimulate owns public Weka ingestion, validation, and lowering into the canonical
ValidatedAgenticGraph. The lower-level
dynamo.replay.run_trace_replay(...) API keeps trace_format="weka" as a compatibility adapter,
but delegates graph construction directly to aisimulate-core before Dynamo composes its Router
and Mocker runtime behavior.
Use weka_nested_timestamp_basis="absolute" or "relative" to override nested timestamp
interpretation in run_trace_replay(...). Omitting it preserves the auto heuristic.
Set agentic_lanes to a positive integer to replay a fixed number of trajectories concurrently.
Plays are stable-sorted and stride-assigned to lanes. A lane starts its next play only after its
current play is quiescent, and lanes do not steal work. Agentic replay rejects
replay_concurrency and Planner scaling. A Weka corpus may retain multiple source-model labels as
provenance, but the current runtime projects every node onto the single configured
execution_model. Per-node heterogeneous timing models are not yet supported.
Agentic reports retain the sorted source-model labels in agentic_graph.source_models and record
the explicit projection policy and target in agentic_model_projection.
Weka reports also expose the resolved timestamp basis as the string
weka_nested_timestamp_basis: absolute, relative, or not_applicable. For offline
run_trace_replay(...), read report.summary["weka_nested_timestamp_basis"]; online replay
includes the same field in its summary dictionary. The Dynamo runner includes it in
report.metadata, including when raw-report capture is disabled. The requested basis remains
in the workload configuration.
Agentic Mooncake v2 begins with a required versioned header:
Each request row contains a globally unique request_id, nonempty play_id, session_id, model,
exact input and output metadata, hash_ids, and not_before_ms. Optional
source_play_ordinal and recorded_api_time_ms fields preserve source ordering and recorded timing
evidence. The optional dependencies array contains typed incoming edges. Each edge identifies its
predecessor, a dispatch or completion trigger, a nonnegative delay, and a sequence, spawn,
join, or replay_barrier relation.
Agentic Mooncake v2 remains an optional materialized interchange format. Dynamo intentionally does not ship a Weka parser, lowering implementation, or Weka-to-v2 converter; those source semantics and any future materialization tooling belong to AISimulate.
Engine and Adapter Rules
engine.mode: aggregatedrequiresengine.workers.aggregated.engine.mode: disaggregatedrequiresengine.workers.prefill,engine.workers.decode, andengine.kv_transfer.aisimulate predictandaisimulate recommendcurrently support TensorRT-LLM only in aggregated mode. This is an offline simulation limitation; Dynamo runtime deployments support TensorRT-LLM disaggregated serving.- Each
parallelismmapping is concrete and containsreplicas,tensor,pipeline,attention_data,moe_tensor, andmoe_expert. engine.context_lengthdefaults tomax, which AISimulate resolves from the model’s Hugging Face configuration. Default KV block sizes are 64 for vLLM, 1 for SGLang, and 32 for TensorRT-LLM.- Scheduler defaults are 8,192 batched tokens for every role and 256 sequences for aggregated and decode workers. Prefill workers default to one sequence.
timing.type: defaultuses the AIConfigurator forward-pass model shipped in theaisimulatewheel.fixedrequires bothprefill_msanddecode_ms;polynomialselects the built-in polynomial model.router.policy: kv_routerrequires more than one routable worker. Setrouter.prefill_load_model.typetononeoraic.planner.policyisdisabledorenabled. When enabled,planner.max_num_gpuslimits the Planner runtime budget; it is distinct from recommendation candidate constraints.evaluation.slaaccepts eithere2e_msalone orttft_msanditl_mstogether. The two forms are mutually exclusive.
Overrides
--set accepts dot-separated, schema-valid paths. The field does not need to be present in the
input YAML:
An override cannot create an unknown field. Override a complete mapping when changing a tagged
configuration shape, such as traffic.source.
Output
The output directory contains:
prediction.json preserves the Dynamo runner report, including summary metrics and available
Planner diagnostics. requests.jsonl is present only with --capture-per-request. --format
changes standard output but not durable files.
The command exits with 0 on success, 1 for execution failure, 2 for CLI or configuration
errors, and 130 when interrupted.