NIXLBench Usage and Troubleshooting
This page covers running NIXLBench benchmarks end-to-end, including worker coordination, the four communication patterns, storage backend examples, and essential CLI options. For installation prerequisites, see Building NIXLBench.
Worker Coordination
For the common two-worker case, we recommend ASIO, which connects the workers directly over TCP without requiring etcd:
Both workers must use the listener host’s IP address and the same port. For two processes on one host, the default address and port are 127.0.0.1:12345.
Use etcd when coordinating more than two workers, such as many-to-one, one-to-many, or larger pairwise and TP runs. Start an etcd server with Docker:
When using etcd, all workers in a benchmark group must connect within 60 seconds of the first worker joining. Workers that miss this window cause the barrier to fail and the benchmark to abort.
Single-instance storage benchmarks run without etcd or ASIO.
Communication Patterns
NIXLBench supports four communication patterns selected with the --scheme flag. All examples below use the UCX backend with VRAM transfers.
Pairwise
Pairwise is the default pattern. Data transfers between matched pairs of initiators and targets, making it ideal for point-to-point throughput measurement.
On the host that owns 192.0.2.10, start the listener:
--mode SG is the default and assigns one GPU to each process. Then run the same command on the peer host. It connects to the listener automatically.
To have the two processes each drive multiple GPUs, use --mode MG and set the device counts. For example, to use four GPUs in each process:
Run the same MG command on the peer host.
Many-to-One
Multiple initiators send data to a single target. This pattern measures how a target handles concurrent incoming transfers from several sources.
Launch one target worker and multiple initiator workers, all pointing to the same etcd server.
One-to-Many
A single initiator sends data to multiple targets. This pattern is useful for measuring fan-out performance such as broadcast or scatter workloads.
Launch one initiator worker and multiple target workers, all pointing to the same etcd server.
TP (Tensor Parallel)
All-to-all exchange where every worker communicates with every other worker. This pattern simulates tensor-parallel distributed training workloads.
Launch two or more workers, all pointing to the same etcd server. Each worker acts as both initiator and target.
Storage Backend Examples
Storage backends benchmark file and object I/O operations. They can run without etcd when launched as a single instance.
GPUDirect Storage (GDS)
Run a single-instance GDS benchmark with direct I/O. For backend-specific flags, see the GPUDirect Storage page.
OBJ (S3)
Run an S3 object storage benchmark using CLI flags for credentials. For backend-specific flags, see the OBJ page.
INFINIA
Run a single-instance INFINIA benchmark using a backend configuration file. For the file format and backend requirements, see INFINIA.
For backend-specific options not listed on this page, see the corresponding backend page in the User Guide.
CLI Options
Core Configuration
Memory and Transfer Configuration
NIXLBench accepts a TOML configuration file via --config_file where CLI parameter names are used as keys in the global scope. For example, backend="UCX" in the config file is equivalent to --backend UCX on the command line. When a parameter appears in both the config file and on the command line, the command-line value takes precedence.
The --worker_type nvshmem option selects the NVSHMEM worker for GPU-only VRAM-to-VRAM transfers. NVSHMEM workers require --initiator_seg_type VRAM and --target_seg_type VRAM.
Reading Benchmark Output
After each run, NIXLBench prints a results table with one row per block-size and batch-size combination. Block sizes sweep from --start_block_size to --max_block_size, doubling each step. For each block size, batch sizes sweep from --start_batch_size to --max_batch_size in the same way.
When running pairwise benchmarks with more than two workers, NIXLBench adds the Aggregate B/W and Network Util columns. These columns do not appear for other communication patterns or two-worker pairwise runs.
The latency columns break down each transfer into three phases. Prep measures buffer registration and transfer handle setup overhead. Tx measures the actual data movement time. Post measures completion checking and cleanup. Together, Avg Lat. reflects the end-to-end latency across all three phases.
Troubleshooting
etcd Connection Failures
Symptoms: Workers fail to join the benchmark group, barrier timeout errors, or “connection refused” messages.
Resolution:
Verify etcd is running and reachable:
If a previous NIXLBench run failed or was interrupted, stale keys may prevent new runs. Clean them up:
Confirm that all workers use the same --etcd_endpoints value and that the etcd server is accessible from every host in the benchmark.
Build Failures
Symptoms: Compilation errors during the native build, missing header files, or linker errors.
Resolution:
For UCX builds missing RDMA libraries:
For etcd-cpp-api build errors related to cpprestsdk or protobuf:
For Docker build failures, clear the cache and rebuild:
For full build instructions, see Building NIXLBench.
CUDA / GPU Not Found
Symptoms: CUDA-related errors at launch, nvcc not found, or GPU detection failures.
Resolution:
Ensure CUDA is in your environment:
Verify the installation:
Backend Library Missing
Symptoms: “library not found” errors at runtime when launching NIXLBench.
Resolution:
Update the shared library cache:
Check that all required libraries are resolved:
If libraries are installed in non-standard paths, add them to LD_LIBRARY_PATH before running NIXLBench.