Overview
What is NIXL?
NIXL (NVIDIA Inference Xfer Library) accelerates point-to-point data transfers in AI inference frameworks such as NVIDIA Dynamo. It abstracts heterogeneous memory types — VRAM, DRAM, file, block, and object storage — through a modular plug-in architecture.
Distributed inference demands high-bandwidth networking across heterogeneous data paths and dynamic scaling. NIXL addresses these requirements with a unified transfer API.
NIXL delivers high-bandwidth, low-latency point-to-point data transfers across VRAM (HBM), DRAM, local and remote SSDs, and distributed storage. Multiple backend plug-ins — UCX, Libfabric, GPUDirect Storage, S3 over RDMA, and others — handle the underlying transports. NIXL abstracts connection management, addressing schemes, and memory characteristics so inference frameworks integrate through a single API.
Key Capabilities
- High-performance point-to-point data transfers with low latency and high bandwidth
- Unified abstraction across heterogeneous memory types — VRAM (HBM), DRAM, block, file, and object storage
- Multiple backend plug-ins — UCX, Libfabric, GPUDirect Storage, S3 over RDMA, and others
- Dynamic scaling — agents can be added or removed at runtime without disrupting existing transfers
- Asynchronous transfer model with non-blocking status checking for overlapping computation with data movement
- Modular plug-in architecture for extensibility, allowing new backends to be added without modifying the core library.
NIXL automatically selects the optimal backend based on the source and destination memory types and the backends available on both agents. You do not need to manually specify which transport to use. See Backend Selection for the full support matrix.