Communicator Class

class nccl.core.Communicator(ptr: int | None = None)

Bases: object

NCCL communicator for collective and point-to-point operations.

A communicator represents a group of participants that perform NCCL operations. Each participant is assigned an integer rank in [0, nranks).

Most users should create communicators with init() or init_all(). The constructor is a low-level interoperability entry point for wrapping an existing NCCL communicator pointer or creating a null communicator for later initialization.

A communicator instance provides collective and point-to-point operations, lifecycle and resource management, and properties describing its rank, device, topology, and capabilities.

__init__(ptr: int | None = None) → None

Wraps an existing NCCL communicator pointer.

This is a low-level interoperability entry point for an NCCL communicator pointer obtained from another library or framework. Most users should create communicators with init() or init_all() instead.

Omitting ptr or passing 0 creates a null communicator. A null communicator can later be initialized with initialize(), or used with grow() to join an existing communicator.

Parameters:

ptr – Address of an existing NCCL communicator, represented as a Python integer. None and 0 create a null communicator. Defaults to None.

Properties

Communicator.properties returns an NCCLCommProperties holding the properties NCCL reports for the communicator, including fields that have no dedicated accessor. The groups below provide per-field accessors for the values needed most often.

Communicator.properties

All properties NCCL reports for this communicator.

Use this to read several properties at once, or to reach the fields that have no dedicated accessor.

Returns:

An NCCLCommProperties holding every field NCCL reports for this communicator.

Raises:

NcclInvalid – If the communicator is not initialized.

Identity

Communicator.ptr

Integer value of the underlying ncclComm_t (0 if destroyed or null).

Communicator.is_valid

Whether the communicator is valid (not destroyed or null).

Communicator.nranks

Total number of ranks in the communicator.

Raises:

NcclInvalid – If the communicator is not initialized.

Communicator.device

CUDA device associated with this communicator.

Returns a cuda.core.Device providing additional functionality such as to_system_device for obtaining the NVML device, device properties, and synchronization.

Raises:

NcclInvalid – If the communicator is not initialized.

Communicator.rank

This caller’s rank within the communicator (0 to nranks - 1).

Raises:

NcclInvalid – If the communicator is not initialized.

Device-API capability

These properties reflect the underlying NCCL ncclCommProperties_t structure.

Communicator.cuda_dev

CUDA device ID associated with this communicator.

Raises:

NcclInvalid – If the communicator is not initialized.

Communicator.nvml_dev

NVML device ID for the GPU associated with this communicator.

Uses the NVML indexing space, which may differ from CUDA indexing.

Raises:

NcclInvalid – If the communicator is not initialized.

Communicator.device_api_support

Whether device-side NCCL operations are supported on this platform.

If False, a device communicator cannot be created.

Raises:

NcclInvalid – If the communicator is not initialized.

Communicator.multimem_support

Whether ranks in the same LSA team can communicate using multimem.

If False, a device communicator cannot be created with multimem resources.

Raises:

NcclInvalid – If the communicator is not initialized.

Communicator.gin_type

GPU-Initiated Networking (GIN) type reaching every rank.

If equal to NcclGinType.NONE, a device communicator cannot be created with GIN connection type NcclGinConnectionType.FULL. A rail-restricted transport may still be available; see railed_gin_type.

Raises:

NcclInvalid – If the communicator is not initialized.

Communicator.n_lsa_teams

Number of Load/Store Accessible (LSA) teams for this communicator.

Raises:

NcclInvalid – If the communicator is not initialized.

Communicator.host_rma_support

Whether host RMA is supported on this communicator.

Raises:

NcclInvalid – If the communicator is not initialized.

Communicator.railed_gin_type

Railed GIN type supported by this communicator.

If equal to NcclGinType.NONE, a device communicator cannot be created with GIN connection type NcclGinConnectionType.RAIL.

Raises:

NcclInvalid – If the communicator is not initialized.

NCCLCommProperties

Covers the accessors in both groups above, plus the fields that have none.

class nccl.core.NCCLCommProperties(*, rank: int, n_ranks: int, cuda_dev: int, nvml_dev: int, device_api_support: bool, multimem_support: bool, gin_type: NcclGinType, n_lsa_teams: int, host_rma_support: bool, railed_gin_type: NcclGinType, comm_hash: int | None = None, gin_min_stride: int | None = None, gin_connection_type: NcclGinConnectionType | None = None, available_gin_types: frozenset[NcclGinType] | None = None, dev_comm_runtime_version_size: int | None = None, cft_support: bool | None = None, cft_multicast_support: bool | None = None, cft_counted_support: bool | None = None)

Bases: object

The properties NCCL reports for a communicator.

Returned by Communicator.properties. These values are fixed for the lifetime of the communicator. Version-marked fields are None when nccl4py was built against an older NCCL.

See also

ncclCommProperties for the description of each field.

rank: int

This caller’s rank within the communicator.

n_ranks: int

Number of ranks in the communicator.

cuda_dev: int

CUDA device ID associated with the communicator.

nvml_dev: int

NVML device ID for the GPU. Uses the NVML indexing space, which may differ from CUDA indexing.

device_api_support: bool

Whether device-side NCCL operations are supported.

multimem_support: bool

Whether ranks in the same LSA team can communicate using multimem.

gin_type: NcclGinType

GIN transport reaching every rank. NONE unless gin_connection_type is FULL, even when a rail-restricted transport is available.

n_lsa_teams: int

Number of LSA teams.

host_rma_support: bool

Whether host RMA is supported.

railed_gin_type: NcclGinType

GIN transport reaching ranks within a rail. NONE only when no GIN transport is available at all.

comm_hash: int | None = None

Hash identifying the communicator, shared by all its ranks (NCCL 2.31+).

gin_min_stride: int | None = None

Granularity of the GIN rank stride this communicator supports. A stride passed as NCCLDevCommRequirements.gin_custom_stride must be a multiple of this value, and no larger than the rail team’s stride. It is 1 when gin_connection_type is FULL (NCCL 2.31+).

gin_connection_type: NcclGinConnectionType | None = None

NONE, RAIL or FULL. A device communicator may request this topology or a narrower one via NCCLDevCommRequirements.gin_connection_type (NCCL 2.31+).

Type:

Widest GIN connection topology this communicator supports

available_gin_types: frozenset[NcclGinType] | None = None

The GIN transports this communicator can use, e.g. NcclGinType.GDAKI in props.available_gin_types. Empty when GIN is unavailable (NCCL 2.31+).

dev_comm_runtime_version_size: int | None = None

Size, in bytes, of the device communicator structure in the running NCCL library (NCCL 2.31+).

cft_support: bool | None = None

Whether every rank in the communicator supports CFT unicast logical endpoints, which requires CUDA and driver 13.3+ on each. NCCL reduces this across ranks, so False does not mean the local GPU lacks support (NCCL 2.32+).

cft_multicast_support: bool | None = None

Whether every rank in the communicator supports multicast CFT logical endpoints. Independent of cft_support; a GPU may support multicast endpoints without unicast ones (NCCL 2.32+).

cft_counted_support: bool | None = None

Whether every rank in the communicator supports counted CFT logical-endpoint operations (NCCL 2.32+).

Teams

A NCCLTeam names a strided subset of the communicator’s ranks. The members below return the predefined teams; pass one to TeamRequirement or to the rank converters. See Teams for team semantics.

Communicator.team_world

The world team for this communicator.

Raises:

NcclInvalid – If the communicator is not initialized.

Communicator.team_lsa

The LSA team for this communicator.

Raises:

NcclInvalid – If the communicator is not initialized.

Communicator.team_rail

The rail team for this communicator.

Raises:

NcclInvalid – If the communicator is not initialized.

Communicator.team_cft_multimem

The CFT multimem team for this communicator.

Raises:

NcclInvalid – If the communicator is not initialized.

Communicator.team_cft(mode: NcclCftTeamMode = NcclCftTeamMode.FLAT) → NCCLTeam

The CFT team for this communicator, in the requested layout.

Parameters:

mode – Team layout. Defaults to NcclCftTeamMode.FLAT, matching the C default.

Raises:

NcclInvalid – If the communicator is not initialized.

See also

ncclTeamCft()

Communicator.team_rank_to_world(team: NCCLTeam, team_rank: int) → int

Maps a rank within team to its rank in this communicator.

team is anchored at this rank, so team.rank maps back to rank and neighbours are offset by team.stride.

Parameters:
Returns:

The corresponding rank in this communicator.

Raises:

NcclInvalid – If the communicator is not initialized.

Communicator.team_rank_to_lsa(team: NCCLTeam, team_rank: int) → int

Maps a rank within team to its rank in the LSA team.

The LSA-relative counterpart of team_rank_to_world(): team.rank maps back to this rank’s index in team_lsa. Only meaningful when team_rank names a peer that shares this rank’s LSA team.

Parameters:
Returns:

The corresponding rank in the LSA team, or -1 if the device resource state could not be initialized.

Raises:

NcclInvalid – If the communicator is not initialized.

NcclCftTeamMode

class nccl.core.NcclCftTeamMode(value, names=<not given>, *values, module=None, qualname=None, type=None, start=1, boundary=None)

Bases: IntEnum

CFT team layout, mirroring ncclCftTeamMode_t.

Selects which ranks in the CFT unicast group the CFT team Communicator.team_cft() includes.

FLAT = 0

Every rank in the CFT unicast group.

HIER_MULTIMEM = 1

The ranks of the CFT unicast group sharing the same index across multicast CFT groups.

HIER_LSA = 2

The ranks of the CFT unicast group sharing the same index across LSA groups.