Backend Selection

View as Markdown

How NIXL Selects Backends

NIXL automatically selects the optimal backend based on the source and destination memory types and the backends available on both the local and remote agents. When multiple backends support a given transfer, NIXL chooses the most efficient one.

To override automatic selection, specify the desired backend when creating a transfer request. In most cases, automatic selection is the recommended approach.

In Python, backends listed in nixl_agent_config.backends (default: ['UCX']) are auto-initialized when the agent is created. In C++ and Rust, you must call createBackend() / create_backend() explicitly for each backend you want to use.

For a detailed explanation of how backends interact with the Transfer Agent, Memory Sections, and Metadata Handler, see Architecture and Concepts.

Memory Types

NIXL provides a unified interface for registering and transferring data across five memory and storage types. Each backend supports a specific subset of these types.

Memory TypeAPI EnumPython StringDescription
VRAM (HBM)VRAM_SEG"VRAM" or "cuda"GPU high-bandwidth memory
DRAMDRAM_SEG"DRAM" or "cpu"Standard host memory
FileFILE_SEG"FILE"Local and remote file systems
Object StorageOBJ_SEG"OBJ"Distributed object stores (S3, Azure Blob, DDN INFINIA)
Block StorageBLK_SEG"BLOCK"Block-level storage devices

For more on how memory types fit into the NIXL architecture, see the Overview.

Backend Support Matrix

NIXL loads each backend as a plug-in, and each plug-in supports specific memory types and transport protocols.

BackendTransfer TypesProtocol / Technology
UCXVRAM ↔ VRAM
VRAM ↔ DRAM
DRAM ↔ DRAM
RoCE, InfiniBand, TCP
LibfabricOFI (verbs, EFA, TCP providers)
MooncakeMooncake Transfer Engine
UCCL-P2PPeer-to-peer transport
DOCA GPUNetIODOCA GPUDirect Async
GDSVRAM ↔ File
DRAM ↔ File
GPUDirect Storage (cuFile API)
GDS_MTGPUDirect Storage (cuFile API, multi-threaded)
POSIXDRAM ↔ Fileio_uring / linux_aio / posix_aio
HF3FS3FS fuse filesystem
OBJDRAM ↔ ObjectS3, S3_CRT, S3/RDMA
Azure BlobAzure Blob REST API
INFINIADRAM ↔ Object
VRAM ↔ Object
DDN INFINIA Client and Async libraries
GUSLIDRAM ↔ BlockBlock I/O

Common Scenarios

The table below maps common scenarios to recommended backends. NIXL selects the right backend automatically if it is initialized on both agents.

ScenarioSourceDestinationRecommended BackendWhy
GPU-to-GPU (same or remote node)VRAMVRAMUCXRDMA-capable, supports RoCE, InfiniBand, and TCP
CPU-to-CPU (remote node)DRAMDRAMUCXStandard high-performance network transport
GPU-to-file (local NVMe)VRAMFileGDS / GDS_MTGPUDirect Storage bypasses CPU bounce buffer
CPU-to-fileDRAMFilePOSIXStandard filesystem I/O
CPU-to-S3DRAMObject StorageOBJS3, S3_CRT, S3/RDMA protocol support
GPU-to-GPU (Mooncake network)VRAMVRAMMooncakeMooncake Transfer Engine for specialized deployments
GPU-to-CPU (DOCA/GPUDirect Async)VRAMDRAMDOCA GPUNetIODOCA GPUDirect Async path on SmartNIC-equipped systems
Block storage I/OBlock StorageDRAMGUSLIBlock-level I/O operations
CPU-to-Azure BlobDRAMObject StorageAzure BlobAzure Blob REST API
CPU/GPU to DDN INFINIADRAM / VRAMObject StorageINFINIAAsynchronous object I/O from registered host or GPU memory
File I/O (3FS)FileDRAMHF3FS3FS fuse-based filesystem

When multiple backends can handle a transfer (e.g., both UCX and Libfabric support DRAM-to-VRAM), NIXL automatically selects the most efficient one based on the available backends on both agents. You only need to ensure the desired backends are installed and initialized.

Backend plug-ins use the Southbound API exposed by the NIXL core.