Configuration

Configuration objects passed to communicator creation methods and to individual collectives, plus the flag enums they consume.

NCCLConfig

Used by Communicator.init(), Communicator.initialize(), Communicator.split(), Communicator.shrink(), and Communicator.grow(). Fields left unset (None) remain at NCCL’s internal default; values are validated by the C library when the config is consumed.

class nccl.core.NCCLConfig(*, blocking: bool | None = None, cga_cluster_size: int | None = None, min_ctas: int | None = None, max_ctas: int | None = None, net_name: str | None = None, split_share: bool | None = None, traffic_class: int | None = None, comm_name: str | None = None, collnet_enable: bool | None = None, cta_policy: CTAPolicy | None = None, shrink_share: bool | None = None, nvls_ctas: int | None = None, n_channels_per_net_peer: int | None = None, nvlink_centric_sched: bool | None = None, graph_usage_mode: int | None = None, num_rma_ctx: int | None = None, max_p2p_peers: int | None = None, graph_stream_ordering: int | None = None, launch_order_implicit: bool | None = None, num_rma_sig: int | None = None, rma_eager_init: bool | None = None, host_cft_mode: NcclHostCftMode | None = None, nvls_host_mode: NcclNvlsHostMode | None = None)

Bases: LowppSpec

NCCL configuration for communicator initialization.

Provides configuration options for NCCL communicators, allowing fine-tuning of performance and behavior characteristics. Fields not set in the constructor remain at NCCL’s internal default; values are validated by the C library when the config is consumed.

See also

ncclConfig_t for the description of each field.

blocking: bool | None = None

Blocking (True) or non-blocking (False) communicator behavior. If unset, NCCL uses True.

Available since NCCL 2.14.0.

cga_cluster_size: int | None = None

Cooperative Group Array (CGA) size for kernels (0-8). If unset, NCCL uses 4 for sm90+, 0 otherwise.

Available since NCCL 2.17.0.

min_ctas: int | None = None

Minimum number of CTAs per kernel; positive integer up to 32. If unset, NCCL uses 1.

Available since NCCL 2.17.0.

max_ctas: int | None = None

Maximum number of CTAs per kernel; positive integer up to 32. If unset, NCCL uses 32.

Available since NCCL 2.17.0.

net_name: str | None = None

Network module name (e.g. ‘IB’, ‘Socket’). Case-insensitive. If unset, NCCL auto-selects.

Available since NCCL 2.17.0.

split_share: bool | None = None

Share resources with the child communicator during split. If unset, NCCL uses False.

Available since NCCL 2.18.0.

traffic_class: int | None = None

Traffic class (TC) for network operations (>= 0). Network-specific meaning.

Available since NCCL 2.26.0.

comm_name: str | None = None

User-defined communicator name for logging and profiling.

Available since NCCL 2.27.0.

collnet_enable: bool | None = None

Enable (True) or disable (False) IB SHARP. If unset, NCCL uses False.

Available since NCCL 2.27.0.

cta_policy: CTAPolicy | None = None

CTA scheduling policy. If unset, NCCL uses CTAPolicy.DEFAULT.

Available since NCCL 2.27.0.

shrink_share: bool | None = None

Share resources with the child communicator during shrink. If unset, NCCL uses False.

Available since NCCL 2.27.0.

nvls_ctas: int | None = None

Total number of CTAs for NVLS kernels (positive integer). If unset, NCCL auto-determines.

Available since NCCL 2.27.1.

n_channels_per_net_peer: int | None = None

Number of network channels for pairwise communication. Positive integer, rounded up to power of 2. If unset, NCCL uses an AlltoAll-optimized value.

Available since NCCL 2.28.0.

Enable NVLink-centric scheduling. If unset, NCCL uses False.

Available since NCCL 2.28.2.

graph_usage_mode: int | None = None

Graph usage mode. Supported values are 0 (no graphs), 1 (one graph), and 2 (multiple graphs or a mix of graph and non-graph). If unset, NCCL uses 2.

Available since NCCL 2.29.0.

num_rma_ctx: int | None = None

Number of RMA contexts. Positive integer. If unset, NCCL uses 1.

Available since NCCL 2.29.0.

max_p2p_peers: int | None = None

Maximum number of peers any rank will concurrently communicate with using P2P. Positive integer. If unset, NCCL uses the communicator size.

Available since NCCL 2.30.0.

graph_stream_ordering: int | None = None

Whether NCCL preserves stream-ordering semantics for collectives captured into CUDA graphs. Supported values are 0 (disabled) or 1 (enabled). The value 0 cannot be combined with graph_usage_mode=2. Also controllable via the NCCL_GRAPH_STREAM_ORDERING environment variable. If unset, NCCL uses 1.

Available since NCCL 2.30.5.

launch_order_implicit: bool | None = None

Whether this communicator takes part in implicit launch ordering. Within one CUDA context, operations on communicators that enable it must not overlap with operations on communicators that do not. Also controllable via the NCCL_LAUNCH_ORDER_IMPLICIT environment variable, which takes precedence. If unset, NCCL uses False.

Available since NCCL 2.31.0.

num_rma_sig: int | None = None

Number of one-sided RMA signal indexes available per context. Non-negative integer; bounds the signal_index accepted by the signal and wait-signal operations. If unset, NCCL uses 1.

Available since NCCL 2.31.0.

rma_eager_init: bool | None = None

Whether the collective one-sided RMA signal setup is initialized at communicator creation rather than at the first window registration. True is required if the communicator issues signal or wait-signal operations without first registering a symmetric window. Also controllable via the NCCL_RMA_EAGER_INIT environment variable, which takes precedence. If unset, NCCL uses False.

Available since NCCL 2.31.0.

host_cft_mode: NcclHostCftMode | None = None

Host-side Compute Fabric Transport mode. Controls whether the communicator creates the CUDA fabric logical endpoints backing the host-side CFT queries. If unset, NCCL uses NcclHostCftMode.DEFAULT.

Available since NCCL 2.31.1.

nvls_host_mode: NcclNvlsHostMode | None = None

Host-side NVLS mode. Selects which host NVLS components the communicator uses. If unset, NCCL uses its library-defined default, which is currently equivalent to NcclNvlsHostMode.ENABLE.

Available since NCCL 2.32.0.

NcclHostCftMode

Value of NCCLConfig.host_cft_mode.

class nccl.core.NcclHostCftMode(value, names=<not given>, *values, module=None, qualname=None, type=None, start=1, boundary=None)

Bases: IntEnum

Host-side Compute Fabric Transport (CFT) mode, mirroring ncclHostCftMode_t.

Set on NCCLConfig.host_cft_mode to control whether the communicator creates the CUDA fabric logical endpoints that back the host-side CFT queries.

DEFAULT = -2147483648

Use the version-specific default.

ENABLE = 1

Enable host-side CFT support, creating the communicator’s unicast and multicast logical endpoints during the first window registration.

DISABLE = 2

Disable host-side CFT support.

FALLBACK = 3

Try to create the logical endpoints; on error, disable host-side CFT instead of failing.

NcclNvlsHostMode

Value of NCCLConfig.nvls_host_mode.

class nccl.core.NcclNvlsHostMode(value, names=<not given>, *values, module=None, qualname=None, type=None, start=1, boundary=None)

Bases: IntFlag

Host-side NVLS mode, mirroring ncclNvlsHostMode_t.

Set on NCCLConfig.nvls_host_mode to select which host NVLS components the communicator uses. The two DISABLE_ members combine; leave the field unset for the library default.

ENABLE = 0

Enable all host NVLS components.

DISABLE_TRANSPORT = 1

Disable the NVLS transport and the registered-buffer optimization.

DISABLE_SYMMETRIC_MULTIMEM = 2

Disable multimem in NCCL’s symmetric kernels and copy-engine paths.

DISABLE = 2147483647

Disable all present and future host NVLS components.

NCCLCollConfig

Accepted as the config argument of every collective on Communicator. See the individual field documentation for unset behavior and usage requirements.

class nccl.core.NCCLCollConfig(*, min_ctas: int | None = None, max_ctas: int | None = None, nvls_ctas: int | None = None, cga_cluster_size: int | None = None, alg_selection: str | None = None, force_alg_selection: bool | None = None, cta_policy: CTAPolicy | None = None, user_profiler_tag: int | None = None, launch_completion_event: Event | int | None = None, vendor_options: tuple[VendorOption, ...] = ())

Bases: LowppSpec

Per-call configuration for a single collective.

Accepted as the config argument of every collective on Communicator, tuning that one call. The same configuration must be set on every rank; NCCL validates it only locally, when the call is issued.

See also

ncclCollConfig_t

min_ctas: int | None = None

Lower bound on channels/CTAs for this call. Also set by NCCL_MIN_CTAS, which takes precedence. If unset, inherits NCCLConfig.min_ctas.

Available since NCCL 2.31.0.

max_ctas: int | None = None

Upper bound on channels/CTAs for this call, clamped to the communicator’s max_ctas. Also set by NCCL_MAX_CTAS, which takes precedence. If unset, inherits NCCLConfig.max_ctas.

Available since NCCL 2.31.0.

nvls_ctas: int | None = None

NVLS-pool-specific channel cap for this call. Also set by NCCL_NVLS_NCHANNELS, which takes precedence. If unset, inherits NCCLConfig.nvls_ctas.

Available since NCCL 2.31.0.

cga_cluster_size: int | None = None

CUDA thread-block-cluster size (0-8, Hopper+). Inconsistent values within one group are undefined behavior. Also set by NCCL_CGA_CLUSTER_SIZE, which takes precedence. If unset, inherits NCCLConfig.cga_cluster_size.

Available since NCCL 2.31.0.

alg_selection: str | None = None

Selection string filtering which algorithms this call may use, e.g. "ring", "tree,ring", "^symk". If unset or empty, NCCL selects automatically.

Available since NCCL 2.31.0.

force_alg_selection: bool | None = None

Whether an unsatisfiable alg_selection is an error rather than a fallback to automatic selection. If unset, NCCL uses True.

Available since NCCL 2.31.0.

cta_policy: CTAPolicy | None = None

CTA scheduling policy for this call. Also set by NCCL_CTA_POLICY, which takes precedence. If unset, inherits NCCLConfig.cta_policy.

Available since NCCL 2.31.0.

user_profiler_tag: int | None = None

Opaque value delivered verbatim to profiler plugins with this call’s profiler events; does not affect execution. Values with the most-significant bit set are reserved by NCCL. If unset, NCCL uses 0.

Available since NCCL 2.31.0.

launch_completion_event: NcclEventSpec | None = None

Caller-owned, rank-local CUDA event recorded at collective kernel launch completion. With CUDA versions earlier than 12.3, NCCL records the event before the kernel launch instead. A cuda.core.Event created by Device().create_event() has the required timing-disabled configuration; events passed as integer handles must likewise have timing disabled. Interprocess and interop events are unsupported. Either every rank passes an event or none does, and at most one per communicator in a group. Keep it alive through all queued waits and captured-graph executions. If unset, NCCL records no event.

Available since NCCL 2.32.0.

vendor_options: tuple[VendorOption, ...] = ()

Vendor-specific options; (vendor_id, option_id) keys must be unique.

Available since NCCL 2.31.0.

VendorOption

class nccl.core.VendorOption(vendor_id: int, option_id: int, int_value: int | None = None, str_value: str | None = None, raw_value: int | None = None)

Bases: object

A single vendor-specific option attached to an NCCLCollConfig.

Mirrors one ncclConfigExt_t node. Options are identified by the (vendor_id, option_id) pair; the official NCCL library ignores every extension, so an option only has an effect on a vendor library that recognizes its vendor_id. Vendors pick a non-zero vendor_id less than 2**24 that is unlikely to collide.

Exactly one of the three value fields must be set.

See also

ncclConfigExt_t

vendor_id: int

Vendor-chosen identifier, unique across vendor libraries.

option_id: int

Vendor-defined identifier distinguishing options within a vendor.

int_value: int | None = None

Integer value (val.i).

str_value: str | None = None

String value (val.s), encoded to UTF-8.

raw_value: int | None = None

Value of any other type (val.raw), as an integer. If it is an address, the referent must stay valid for the duration of the call.

CTAPolicy

class nccl.core.CTAPolicy(value, names=<not given>, *values, module=None, qualname=None, type=None, start=1, boundary=None)

Bases: IntFlag

NCCL performance policy for CTA scheduling, used by NCCLConfig.cta_policy and NCCLCollConfig.cta_policy.

DEFAULT = 0

Default CTA policy.

EFFICIENCY = 1

Optimize for efficiency.

ZERO = 2

Zero-CTA optimization.

NCCLDevCommRequirements

Used by Communicator.create_dev_comm(). Fields left unset (None) remain at NCCL’s internal default.

class nccl.core.NCCLDevCommRequirements(*, lsa_multimem: bool | None = None, barrier_count: int | None = None, lsa_barrier_count: int | None = None, rail_gin_barrier_count: int | None = None, lsa_ll_a2a_block_count: int | None = None, lsa_ll_a2a_slot_count: int | None = None, gin_force_enable: bool | None = None, gin_context_count: int | None = None, gin_signal_count: int | None = None, gin_counter_count: int | None = None, gin_connection_type: NcclGinConnectionType | None = None, gin_exclusive_contexts: bool | None = None, gin_queue_depth: int | None = None, gin_traffic_class: int | None = None, world_gin_barrier_count: int | None = None, gin_strong_signals_required: bool | None = None, gin_va_signals_required: bool | None = None, gin_custom_stride: int | None = None, gin_type: NcclGinType | None = None, cft_caps: NcclCftCap | None = None, cft_barrier_count: int | None = None, teams: tuple[TeamRequirement, ...] = (), resources: tuple[LsaBarrierRequirement | GinBarrierRequirement | LLA2ARequirement, ...] = ())

Bases: LowppSpec

NCCL device communicator requirements configuration.

This is a reusable high-level Python request consumed by Communicator.create_dev_comm(). Per-team requirements are declared through the teams tuple. Each call snapshots the request into independent low-level ncclDevCommRequirements_t and linked ncclTeamRequirements_t storage, including separate multimem output handles. NCCL copies the requirements and linked-list nodes before the call returns; the resulting DevCommResource retains the storage referenced by each outMultimemHandle. This object may therefore be changed between calls without affecting device communicators that were already created. Do not mutate it concurrently with Communicator.create_dev_comm().

See also

ncclDevCommRequirements for the description of each field.

lsa_multimem: bool | None = None

Enable multimem on the LSA team. If unset, NCCL uses False.

barrier_count: int | None = None

Number of barriers required. If unset, NCCL uses 0.

lsa_barrier_count: int | None = None

Number of LSA barriers. If unset, NCCL uses 0.

rail_gin_barrier_count: int | None = None

Number of railed GIN barriers. If unset, NCCL uses 0.

lsa_ll_a2a_block_count: int | None = None

LSA low-latency all-to-all block count. If unset, NCCL uses 0.

lsa_ll_a2a_slot_count: int | None = None

LSA low-latency all-to-all slot count. If unset, NCCL uses 0.

gin_force_enable: bool | None = None

Force-enable GPU-Initiated Networking (GIN). If unset, NCCL uses False.

gin_context_count: int | None = None

Number of GIN contexts (hint; actual count may differ). If unset, NCCL uses 4.

gin_signal_count: int | None = None

Number of GIN signals (guaranteed to start at id=0). If unset, NCCL uses 0.

gin_counter_count: int | None = None

Number of GIN counters (guaranteed to start at id=0). If unset, NCCL uses 0.

gin_connection_type: NcclGinConnectionType | None = None

GIN connection type. If unset, NCCL uses NcclGinConnectionType.NONE.

gin_exclusive_contexts: bool | None = None

Use exclusive GIN contexts. If unset, NCCL uses False.

gin_queue_depth: int | None = None

GIN queue depth. If unset, NCCL uses 0.

gin_traffic_class: int | None = None

GIN traffic class. If unset, NCCL uses its internal default.

world_gin_barrier_count: int | None = None

Number of world GIN barriers. If unset, NCCL uses 0.

gin_strong_signals_required: bool | None = None

Whether GIN strong signals are required by kernels using this devComm. When False, using GIN strong signals results in undefined behavior. If unset, NCCL uses True.

gin_va_signals_required: bool | None = None

Whether GIN VA signals are required by kernels using this devComm. When False, using GIN VA signals results in undefined behavior. If unset, NCCL uses True.

gin_custom_stride: int | None = None

Stride of ranks to connect for GIN. Only consulted when gin_connection_type is NcclGinConnectionType.CUSTOM_STRIDE, and must be a multiple of NCCLCommProperties.gin_min_stride. If unset, NCCL uses 1.

gin_type: NcclGinType | None = None

GIN transport to require. If unset, NCCL uses NcclGinType.NONE, accepting any available transport.

cft_caps: NcclCftCap | None = None

Compute Fabric Transport capabilities to request, as a bitmask of NcclCftCap values. Creation fails if CFT resources are requested on a communicator where not all ranks support CFT. If unset, NCCL uses NcclCftCap.NONE.

cft_barrier_count: int | None = None

Number of CFT barriers to allocate, one per independently addressed barrier slot the kernel uses (commonly one per CTA). If unset, NCCL uses 0.

teams: tuple[TeamRequirement, ...] = ()

Per-team requirements. Entries for the same team (by value) are merged, keeping first-appearance order; multimem is requested for a team if any of its entries sets it. A team requested with multimem=True yields a multimem handle retrievable via multimem_handle().

resources: tuple[LsaBarrierRequirement | GinBarrierRequirement | LLA2ARequirement, ...] = ()

Device resource requirements (LSA/GIN barriers, low-latency all-to-all). Each entry yields, in order, a handle in resource_handles. Entries are kept as-is (not merged): each is a distinct resource.

NcclCftCap

Bitmask value of NCCLDevCommRequirements.cft_caps.

class nccl.core.NcclCftCap(value, names=<not given>, *values, module=None, qualname=None, type=None, start=1, boundary=None)

Bases: IntFlag

Compute Fabric Transport capabilities, mirroring ncclCftCap_t.

Combined as a bitmask on NCCLDevCommRequirements.cft_caps.

NONE = 0

No CFT capability requested.

CFT = 1

Request unicast CFT logical endpoints.

MULTIMEM = 2

Request multicast CFT operations and multimem CFT barriers.

Requirement entries

The element types of NCCLDevCommRequirements.teams and NCCLDevCommRequirements.resources.

TeamRequirement

class nccl.core.TeamRequirement(team: NCCLTeam, multimem: bool = False)

Bases: object

A per-team requirement for device communicator creation.

Pass a tuple of these as NCCLDevCommRequirements.teams. When multimem is True, NCCL allocates a multicast handle for the team, retrievable afterwards via multimem_handle().

team: NCCLTeam
multimem: bool = False

LsaBarrierRequirement

class nccl.core.LsaBarrierRequirement(team: NCCLTeam, n_barriers: int)

Bases: object

Requests an LSA barrier resource on team with n_barriers barriers.

Add to NCCLDevCommRequirements.resources; the finalized LsaBarrierHandle is returned in resource_handles.

team: NCCLTeam
n_barriers: int

GinBarrierRequirement

class nccl.core.GinBarrierRequirement(team: NCCLTeam, n_barriers: int)

Bases: object

Requests a GIN barrier resource on team with n_barriers barriers.

Add to NCCLDevCommRequirements.resources; the finalized GinBarrierHandle is returned in resource_handles.

team: NCCLTeam
n_barriers: int

LLA2ARequirement

class nccl.core.LLA2ARequirement(n_blocks: int, max_elements: int, max_element_size: int)

Bases: object

Requests a low-latency all-to-all resource with n_blocks blocks, sized to hold up to max_elements elements of at most max_element_size bytes each.

Add to NCCLDevCommRequirements.resources; the finalized LLA2AHandle is returned in resource_handles.

n_blocks: int
max_elements: int
max_element_size: int