Quick Start

View as Markdown

Install

Install via PyPI:

pip install nixl[cu12]

For CUDA 13:

pip install nixl[cu13]

Bare pip install nixl defaults to nixl[cu12] for backwards compatibility, so existing workflows continue to work without changes.

NIXL is supported on Linux only. It is tested on Ubuntu (22.04/24.04) and Fedora. macOS and Windows are not currently supported.

To verify the installation:

python3 -c "import nixl; agent = nixl.nixl_agent('agent1')"

Expected output:

NIXL INFO _api.py:363 Backend UCX was instantiated
NIXL INFO _api.py:253 Initialized NIXL agent: agent1

The sections below follow the Transfer Agent lifecycle: initialization, backend creation, memory registration, metadata exchange, transfer, and teardown.

If you installed NIXL via pip above, you’re ready to go. For building from source, see Building NIXL from Source.

The NIXL workflow follows a strict order:

  1. Create an agent
  2. Create backends (C++ and Rust only — Python auto-initializes)
  3. Register memory
  4. Exchange metadata with remote agents
  5. Create and execute transfers
  6. Check transfer status
  7. Clean up resources

Memory must be registered before metadata exchange. Metadata exchanged before memory registration will not include the memory segment information needed for transfers.

Agent Initialization

A Transfer Agent represents one endpoint in a data transfer. Create one per process with a unique name.

from nixl._api import nixl_agent, nixl_agent_config
config = nixl_agent_config(
enable_prog_thread=True,
enable_listen_thread=True,
listen_port=5555
)
agent = nixl_agent("my_agent", config)
Python Auto-Initialization

The Python API auto-initializes backends listed in nixl_agent_config.backends (default: ['UCX']). In C++ and Rust, backends must be created explicitly as shown in the next section.

Backend Creation

Backends handle data transfer over specific transports. Python auto-initializes backends from the agent config. C++ and Rust require explicit creation.

# Backends are auto-created from nixl_agent_config.backends (default: ['UCX']).
# To add additional backends after agent creation:
agent.create_backend("GDS")

You can create multiple backends on the same agent. NIXL will automatically select the best one for each transfer based on source and destination memory types. See NIXL Backends for guidance on which backends to enable.

Memory Registration

Register memory segments before they can participate in transfers. Registration creates the internal structures backends need for tracking and remote access metadata.

import torch
torch.set_default_device("cuda:0")
tensor = torch.zeros((10, 16), dtype=torch.float32)
# Register the tensor -- NIXL detects GPU memory automatically
reg_descs = agent.register_memory(tensor)
Python Memory Type Detection

Python automatically detects the memory type from the tensor’s device. A CUDA tensor registers as VRAM, a CPU tensor as DRAM. You can also pass raw memory tuples (address, size, device_id, tag) for manual control.

Metadata Exchange

Transfer Agents must exchange metadata before transfers. NIXL supports three modes:

The simplest approach for getting started. Agents exchange metadata directly over a TCP connection. One agent listens for incoming metadata, while the other fetches and sends.

# On the initiator side:
# Fetch the target's metadata (target is listening on ip:port)
agent.fetch_remote_metadata("target", target_ip, target_port)
# Send this agent's metadata to the target
agent.send_local_metadata(target_ip, target_port)
# Wait for metadata to be available
while not agent.check_remote_metadata("target"):
pass

The Python side-channel API (send_local_metadata/fetch_remote_metadata) uses built-in TCP communication. The C++ and Rust examples above show the programmatic approach for same-process usage. For cross-process C++/Rust with side-channel, use the etcd mode or implement your own transport for the metadata blobs.

Creating and Executing Transfers

Create a transfer request, post it (non-blocking), and poll for completion.

# Build transfer descriptors from tensors
local_rows = [tensor[i, :] for i in range(tensor.shape[0])]
local_descs = agent.get_xfer_descs(local_rows)
# Initialize transfer (READ from target into local memory)
xfer_handle = agent.initialize_xfer(
"READ", local_descs, target_descs, "target", b"notification_msg"
)
# Post transfer (non-blocking)
state = agent.transfer(xfer_handle)

Checking Transfer Status

Poll transfer status after posting. The post call returns immediately.

# Poll for completion
while True:
state = agent.check_xfer_state(xfer_handle)
if state == "DONE":
break
elif state == "ERR":
raise RuntimeError("Transfer failed")
# state == "PROC" means still in progress

Transfer status returns differ across languages. Python returns strings ("DONE", "PROC", "ERR"). C++ returns nixl_status_t enum values (NIXL_SUCCESS, NIXL_IN_PROG, negative error codes). Rust returns Result<XferStatus, NixlError> with variants Success and InProgress. Always use the language-appropriate status check.

Teardown

Release resources in order: transfer handles, then memory, then remote metadata.

# Release transfer handle
agent.release_xfer_handle(xfer_handle)
# Deregister memory
agent.deregister_memory(reg_descs)
# Remove remote agent metadata
agent.remove_remote_agent("target")

In Rust, resources are automatically cleaned up when they go out of scope via the Drop trait. Explicit cleanup calls are optional but can be useful for controlling the order of resource release.