KVBench Commands and Examples
This page covers the KVBench command reference, model configuration schemas, and end-to-end LLM examples. For installation and build instructions, see Building KVBench. KVBench’s profile command invokes NIXLBench as a subprocess to run the actual transfer benchmarks.
Command Reference
plan
The plan command generates and displays recommended nixlbench command configurations based on your model architecture and parameters. It computes KV cache transfer sizes and produces the exact NIXLBench invocation without running the benchmark itself.
Use --format to control output format: text (default), json, or csv. The --model_configs flag accepts glob patterns to plan multiple configurations in a single invocation.
profile
The profile command runs NIXLBench with the planned configuration, collecting performance data across KV cache operations and access patterns. It computes the same parameters as plan and then executes nixlbench as a subprocess.
kvcache
The kvcache command analyzes and displays detailed information about the KV cache for a specified model configuration, including model type, sequence lengths, batch sizes, and I/O sizes.
Output:
ct-perftest
The ct-perftest command benchmarks the performance of a single custom traffic pattern. The pattern runs in multiple iterations and then metrics are reported. This is useful for optimizing specific traffic patterns.
Reports: Total latency (time elapsed between the first rank starting and the last rank finishing), average time per iteration, total size sent over the network, and average bandwidth by rank.
GPU memory is allocated with PyTorch on the GPU specified by the CUDA_VISIBLE_DEVICES environment variable. Make sure each process sets this variable to the correct device.
sequential-ct-perftest
The sequential-ct-perftest command benchmarks the performance of a series of traffic patterns executed one after the other. Before running each pattern, all ranks perform a barrier, optionally sleep for a configured duration, then run the pattern and measure execution time.
Reports: Total latency per matrix execution, along with isolated latency (latency when the pattern is run alone), which can be used to evaluate how well the network reacts to congestion.
Command Line Arguments
Common Arguments
These arguments are shared across KVBench commands (plan, kvcache, profile):
CLI Override Arguments
These arguments override values specified in model config files:
Plan Command Arguments
Specific to the plan command:
Shared Benchmark Arguments
These arguments are used by both plan and profile commands and are passed through to NIXLBench:
KVBench uses --etcd-endpoints (hyphens). NIXLBench uses --etcd_endpoints (underscores). Both forms are accepted by the CLI, but this documentation follows each tool’s convention.
CTP Command Arguments
Specific to CTP (Custom Traffic Performance) commands (ct-perftest and sequential-ct-perftest):
Model Configuration Guide
KVBench uses two YAML configuration files: a model architecture file describing the LLM structure, and a model config file specifying parallelism, runtime, and system settings. Both files are passed to KVBench commands via the --model and --model_config flags respectively.
Model Architecture YAML
The model architecture file defines the structural parameters of an LLM. Different attention mechanisms require different fields.
Common fields shared by all architectures:
MLA fields (Multi-Latent Attention, e.g., DeepSeek R1):
MHA/GQA fields (Multi-Head / Grouped-Query Attention, e.g., Llama 3.1):
DeepSeek R1 example (model_deepseek_r1.yaml):
Llama 3.1 70B example (model_llama_3_1_70b.yaml):
Model Config YAML
The model config file has three sections: strategy, runtime, and system.
Strategy fields:
Runtime fields:
System fields:
Block access example (block-tp1-pp8.yaml):
Block access groups KV cache entries into fixed-size pages. Layer access transfers KV cache one transformer layer at a time. Block access typically produces fewer, larger transfers; layer access produces more, smaller transfers.
LLM Examples
End-to-end examples showing model architecture YAML, model config YAML, and the plan and profile commands. These examples can be copy-pasted and run directly from the KVBench directory.
DeepSeek R1
Block Access (TP=1, PP=16)
Model architecture (model_deepseek_r1.yaml):
Model config (block-tp1-pp16.yaml):
Plan command:
Output:
Profile command:
Layer Access (TP=1, PP=16)
Uses the same model architecture YAML as above (model_deepseek_r1.yaml).
Model config (layer-tp1-pp16.yaml):
Plan command:
Output:
With layer access, the batch size increases and block size decreases compared to block access, reflecting the per-layer transfer granularity.
Llama 3.1 70B
Block Access (TP=1, PP=8)
Model architecture (model_llama_3_1_70b.yaml):
Model config (block-tp1-pp8.yaml):
Plan command:
The output follows the same format as the DeepSeek R1 example above, with values computed from the Llama 3.1 70B architecture.
Profile command: