KVBench Commands and Examples

View as Markdown

This page covers the KVBench command reference, model configuration schemas, and end-to-end LLM examples. For installation and build instructions, see Building KVBench. KVBench’s profile command invokes NIXLBench as a subprocess to run the actual transfer benchmarks.

Command Reference

plan

The plan command generates and displays recommended nixlbench command configurations based on your model architecture and parameters. It computes KV cache transfer sizes and produces the exact NIXLBench invocation without running the benchmark itself.

python main.py plan \
--model ./examples/model_deepseek_r1.yaml \
--model_config ./examples/block-tp1-pp8.yaml \
--backend GDS \
--source gpu

Use --format to control output format: text (default), json, or csv. The --model_configs flag accepts glob patterns to plan multiple configurations in a single invocation.

profile

The profile command runs NIXLBench with the planned configuration, collecting performance data across KV cache operations and access patterns. It computes the same parameters as plan and then executes nixlbench as a subprocess.

python main.py profile \
--model ./examples/model_deepseek_r1.yaml \
--model_config ./examples/block-tp1-pp8.yaml \
--backend GDS \
--source gpu \
--etcd-endpoints "http://localhost:2379"

kvcache

The kvcache command analyzes and displays detailed information about the KV cache for a specified model configuration, including model type, sequence lengths, batch sizes, and I/O sizes.

python main.py kvcache \
--model ./examples/model_deepseek_r1.yaml \
--model_config ./examples/block-tp1-pp8.yaml \
--isl 10000 \
--page_size 512

Output:

Model ISL Num Requests Batch Size IO Size TP PP Page Size Access
----------- ----- -------------- ------------ --------- ---- ---- ----------- --------
DEEPSEEK_R1 10000 10 1490 2.25 MB 1 8 512 block

ct-perftest

The ct-perftest command benchmarks the performance of a single custom traffic pattern. The pattern runs in multiple iterations and then metrics are reported. This is useful for optimizing specific traffic patterns.

python main.py ct-perftest ./config.yaml --verify-buffers

Reports: Total latency (time elapsed between the first rank starting and the last rank finishing), average time per iteration, total size sent over the network, and average bandwidth by rank.

GPU memory is allocated with PyTorch on the GPU specified by the CUDA_VISIBLE_DEVICES environment variable. Make sure each process sets this variable to the correct device.

sequential-ct-perftest

The sequential-ct-perftest command benchmarks the performance of a series of traffic patterns executed one after the other. Before running each pattern, all ranks perform a barrier, optionally sleep for a configured duration, then run the pattern and measure execution time.

python main.py sequential-ct-perftest ./config.yaml \
--verify-buffers \
--json-output-path ./results.json

Reports: Total latency per matrix execution, along with isolated latency (latency when the pattern is run alone), which can be used to evaluate how well the network reacts to congestion.

Transfer size (GB) Latency (ms) Isolated Latency (ms) Num Senders
-------------------- -------------- ----------------------- -------------
4.945 35.047 35.421 4
3.230 21.152 21.800 4
1.104 8.222 8.280 4
... ... ... ...
0.129 2.147 2.386 4

Command Line Arguments

Common Arguments

These arguments are shared across KVBench commands (plan, kvcache, profile):

ArgumentDescription
--modelPath to a model architecture config YAML file
--model_configPath to a single model config YAML file
--model_configsPath to multiple model config YAML files (supports glob patterns like configs/*.yaml)

CLI Override Arguments

These arguments override values specified in model config files:

ArgumentDescription
--ppPipeline parallelism size
--tpTensor parallelism size
--islInput sequence length
--oslOutput sequence length
--num_requestsNumber of requests
--page_sizePage size
--access_patternAccess pattern (block or layer)

Plan Command Arguments

Specific to the plan command:

ArgumentDescription
--formatOutput format of the nixlbench command: text, json, or csv (default: text)

Shared Benchmark Arguments

These arguments are used by both plan and profile commands and are passed through to NIXLBench:

ArgumentDescription
--sourceSource of the NIXL descriptors: file, memory, or gpu (default: file)
--destinationDestination of the NIXL descriptors: file, memory, or gpu (default: memory)
--backendCommunication backend: UCX, GDS, GDS_MT, POSIX, GPUNETIO, Mooncake, HF3FS, OBJ (default: UCX)
--worker_typeWorker to use to transfer data: nixl or nvshmem (default: nixl)
--initiator_seg_typeMemory segment type for initiator: DRAM, VRAM, FILE, or OBJ (default: DRAM)
--target_seg_typeMemory segment type for target: DRAM, VRAM, FILE, or OBJ (default: DRAM)
--schemeCommunication scheme: pairwise, manytoone, onetomany, or tp (default: pairwise)
--modeProcess mode: SG (single GPU per process) or MG (multi GPU per process) (default: SG)
--op_typeOperation type: READ or WRITE (default: WRITE)
--check_consistencyEnable consistency checking
--total_buffer_sizeTotal buffer size in bytes (default: 8 GiB)
--recreate_xferRecreate transfer handle for every iteration (default: false for all backends, true for GUSLI)
--start_block_sizeStarting block size in bytes (default: 4 KiB)
--max_block_sizeMaximum block size in bytes (default: 64 MiB)
--start_batch_sizeStarting batch size (default: 1)
--max_batch_sizeMaximum batch size (default: 1)
--num_iterNumber of iterations (default: 1000)
--warmup_iterNumber of warmup iterations (default: 100)
--num_threadsNumber of threads used by benchmark (default: 1)
--num_initiator_devNumber of devices in initiator processes (default: 1)
--num_target_devNumber of devices in target processes (default: 1)
--enable_ptEnable progress thread
--progress_threadsNumber of progress threads (default: 0)
--device_listComma-separated device names (default: all)
--runtime_typeType of runtime to use: ETCD (default: ETCD)
--etcd-endpointsetcd server URL for coordination (default: http://localhost:2379)
--storage_enable_directEnable direct I/O for storage operations
--filepathFile path for storage operations
--enable_vmmEnable VMM memory allocation when DRAM is requested

KVBench uses --etcd-endpoints (hyphens). NIXLBench uses --etcd_endpoints (underscores). Both forms are accepted by the CLI, but this documentation follows each tool’s convention.

CTP Command Arguments

Specific to CTP (Custom Traffic Performance) commands (ct-perftest and sequential-ct-perftest):

ArgumentDescription
config_filePath to YAML configuration file (required, positional argument)
--verify-buffers / --no-verify-buffersVerify buffer contents after transfer (default: false)
--print-recv-buffers / --no-print-recv-buffersPrint received buffer contents (default: false)
--json-output-pathPath to save JSON output (sequential-ct-perftest only)

Model Configuration Guide

KVBench uses two YAML configuration files: a model architecture file describing the LLM structure, and a model config file specifying parallelism, runtime, and system settings. Both files are passed to KVBench commands via the --model and --model_config flags respectively.

Model Architecture YAML

The model architecture file defines the structural parameters of an LLM. Different attention mechanisms require different fields.

Common fields shared by all architectures:

FieldDescription
model_nameModel identifier (e.g., DEEPSEEK_R1, LLAMA3.1_70B)
num_layersNumber of transformer layers
query_head_dimensionDimension of each query head
num_model_paramsTotal model parameter count

MLA fields (Multi-Latent Attention, e.g., DeepSeek R1):

FieldDescription
num_query_headsNumber of query attention heads
embedding_dimensionModel embedding dimension
rope_mla_dimensionRoPE dimension for MLA
mla_latent_vector_dimensionLatent vector dimension for MLA compression

MHA/GQA fields (Multi-Head / Grouped-Query Attention, e.g., Llama 3.1):

FieldDescription
num_query_heads_with_mhaNumber of query heads (MHA variant)
gqa_num_queries_in_groupNumber of queries per KV head group (GQA)

DeepSeek R1 example (model_deepseek_r1.yaml):

model_name: 'DEEPSEEK_R1' # Model identifier
num_layers: 61 # 61 transformer layers
num_query_heads: 128 # 128 query attention heads
query_head_dimension: 128 # 128-dim per query head
embedding_dimension: 7168 # Model embedding dimension
rope_mla_dimension: 64 # RoPE dimension for MLA
mla_latent_vector_dimension: 512 # Latent vector dimension for MLA compression
num_model_params: 671000000000 # 671B parameters

Llama 3.1 70B example (model_llama_3_1_70b.yaml):

model_name: 'LLAMA3.1_70B' # Model identifier
num_layers: 80 # 80 transformer layers
num_query_heads_with_mha: 64 # 64 query heads (MHA)
query_head_dimension: 128 # 128-dim per query head
gqa_num_queries_in_group: 8 # 8 queries per KV head group (GQA)
num_model_params: 70000000000 # 70B parameters

Model Config YAML

The model config file has three sections: strategy, runtime, and system.

Strategy fields:

FieldDescription
tp_sizeTensor parallelism size — number of GPUs for tensor-parallel execution (default: 1)
pp_sizePipeline parallelism size — number of GPUs for pipeline-parallel execution (default: 1)
model_quant_modeModel weight quantization mode, e.g., fp8, fp16, int8 (default: "fp8")
kvcache_quant_modeKV cache quantization mode, e.g., fp8, fp16, int8 (default: "fp8")

Runtime fields:

FieldDescription
islInput sequence length in tokens (default: 1)
oslOutput sequence length in tokens (default: 1)
num_requestsNumber of inference requests (default: 1)

System fields:

FieldDescription
hardwareHardware platform (e.g., "H100", "A100")
backendInference backend engine (e.g., "SGLANG")
access_patternKV cache access pattern: "block" or "layer"
page_sizePage size for access pattern (default: 1)
sourceSource descriptor type
destinationDestination descriptor type

Block access example (block-tp1-pp8.yaml):

strategy:
tp_size: 1 # Tensor parallelism -- 1 GPU for tensor-parallel
pp_size: 8 # Pipeline parallelism -- 8 GPUs for pipeline-parallel
model_quant_mode: "fp8" # Model weight quantization
kvcache_quant_mode: "fp8" # KV cache quantization
runtime:
isl: 1000 # Input sequence length (tokens)
osl: 100 # Output sequence length (tokens)
num_requests: 10 # Number of inference requests
system:
hardware: "H100" # Hardware platform
backend: "SGLANG" # Inference backend engine
access_pattern: "block" # KV cache access pattern
page_size: 16 # Page size for block access

Block access groups KV cache entries into fixed-size pages. Layer access transfers KV cache one transformer layer at a time. Block access typically produces fewer, larger transfers; layer access produces more, smaller transfers.

LLM Examples

End-to-end examples showing model architecture YAML, model config YAML, and the plan and profile commands. These examples can be copy-pasted and run directly from the KVBench directory.

DeepSeek R1

Block Access (TP=1, PP=16)

Model architecture (model_deepseek_r1.yaml):

model_name: 'DEEPSEEK_R1'
num_layers: 61
num_query_heads: 128
query_head_dimension: 128
embedding_dimension: 7168
rope_mla_dimension: 64
mla_latent_vector_dimension: 512
num_model_params: 671000000000

Model config (block-tp1-pp16.yaml):

strategy:
tp_size: 1
pp_size: 16
model_quant_mode: "fp8"
kvcache_quant_mode: "fp8"
runtime:
isl: 1000
osl: 100
num_requests: 10
system:
hardware: "H100"
backend: "SGLANG"
access_pattern: "block"
page_size: 16

Plan command:

python main.py plan \
--model ./examples/model_deepseek_r1.yaml \
--model_config ./examples/block-tp1-pp16.yaml \
--backend GDS \
--source gpu \
--etcd-endpoints "http://localhost:2379"

Output:

================================================================================
Model Config: ./examples/block-tp1-pp16.yaml
ISL: 10000 tokens
Page Size: 256
Requests: 10
TP: 1
PP: 16
================================================================================
nixlbench \
--backend GDS \
--max_batch_size 5958 \
--max_block_size 589824 \
--start_batch_size 5958 \
--start_block_size 589824 \
--target_seg_type VRAM

Profile command:

python main.py profile \
--model ./examples/model_deepseek_r1.yaml \
--model_config ./examples/block-tp1-pp16.yaml \
--backend GDS \
--source gpu \
--etcd-endpoints "http://localhost:2379"

Layer Access (TP=1, PP=16)

Uses the same model architecture YAML as above (model_deepseek_r1.yaml).

Model config (layer-tp1-pp16.yaml):

strategy:
tp_size: 1
pp_size: 16
model_quant_mode: "fp8"
kvcache_quant_mode: "fp8"
runtime:
isl: 1000
osl: 100
num_requests: 10
system:
hardware: "H100"
backend: "SGLANG"
access_pattern: "layer"
page_size: 16

Plan command:

python main.py plan \
--model ./examples/model_deepseek_r1.yaml \
--model_config ./examples/layer-tp1-pp16.yaml \
--backend GDS \
--source gpu \
--etcd-endpoints "http://localhost:2379"

Output:

================================================================================
Model Config: ./examples/layer-tp1-pp16.yaml
ISL: 10000 tokens
Page Size: 256
Requests: 10
TP: 1
PP: 16
================================================================================
nixlbench \
--backend GDS \
--max_batch_size 23829 \
--max_block_size 147456 \
--start_batch_size 23829 \
--start_block_size 147456 \
--target_seg_type VRAM

With layer access, the batch size increases and block size decreases compared to block access, reflecting the per-layer transfer granularity.

Llama 3.1 70B

Block Access (TP=1, PP=8)

Model architecture (model_llama_3_1_70b.yaml):

model_name: 'LLAMA3.1_70B'
num_layers: 80
num_query_heads_with_mha: 64
query_head_dimension: 128
gqa_num_queries_in_group: 8
num_model_params: 70000000000

Model config (block-tp1-pp8.yaml):

strategy:
tp_size: 1
pp_size: 8
model_quant_mode: "fp8"
kvcache_quant_mode: "fp8"
runtime:
isl: 1000
osl: 100
num_requests: 10
system:
hardware: "H100"
backend: "SGLANG"
access_pattern: "block"
page_size: 16

Plan command:

python main.py plan \
--model ./examples/model_llama_3_1_70b.yaml \
--model_config ./examples/block-tp1-pp8.yaml \
--backend GDS \
--source gpu \
--etcd-endpoints "http://localhost:2379"

The output follows the same format as the DeepSeek R1 example above, with values computed from the Llama 3.1 70B architecture.

Profile command:

python main.py profile \
--model ./examples/model_llama_3_1_70b.yaml \
--model_config ./examples/block-tp1-pp8.yaml \
--backend GDS \
--source gpu \
--etcd-endpoints "http://localhost:2379"