NIXLBench
NIXLBench is a benchmarking tool for the NVIDIA Inference Xfer Library (NIXL) that measures data transfer performance across distributed computing environments. It enables developers to evaluate throughput and latency for point-to-point communication, multi-node transfers, and storage I/O by exercising the full range of NIXL backends. Single-node benchmarks require no coordination service. Two-node benchmarks can use either ASIO or etcd, while benchmarks spanning more than two nodes require etcd.
Features
- Network backends — UCX, Libfabric, Mooncake, and DOCA GPUNetIO for high-speed network communication
- Storage backends — GPUDirect Storage, GPUDirect Storage MT, POSIX, HF3FS, OBJ, Azure Blob, INFINIA, and GUSLI for storage operations
- Communication patterns — Pairwise, many-to-one, one-to-many, and TP (tensor parallel)
- Memory types — CPU (DRAM) and GPU (VRAM) transfers
- Worker types — NIXL worker with full backend support, and NVSHMEM worker for GPU-focused VRAM-only transfers
- Coordination — Direct ASIO coordination for two workers, or etcd coordination for larger groups
- Performance metrics — Multi-threading support, VMM memory allocation, latency percentiles, and data consistency validation
Next Steps
- Building NIXLBench — Docker and native build instructions
- Usage and Troubleshooting — Running benchmarks and resolving common issues