Configuration
Configuration objects passed to communicator creation methods and to individual collectives, plus the flag enums they consume.
NCCLConfig
Used by Communicator.init(), Communicator.initialize(),
Communicator.split(), Communicator.shrink(), and
Communicator.grow(). Fields left unset (None) remain at NCCL’s
internal default; values are validated by the C library when the config is
consumed.
- class nccl.core.NCCLConfig(*, blocking: bool | None = None, cga_cluster_size: int | None = None, min_ctas: int | None = None, max_ctas: int | None = None, net_name: str | None = None, split_share: bool | None = None, traffic_class: int | None = None, comm_name: str | None = None, collnet_enable: bool | None = None, cta_policy: CTAPolicy | None = None, shrink_share: bool | None = None, nvls_ctas: int | None = None, n_channels_per_net_peer: int | None = None, nvlink_centric_sched: bool | None = None, graph_usage_mode: int | None = None, num_rma_ctx: int | None = None, max_p2p_peers: int | None = None, graph_stream_ordering: int | None = None, launch_order_implicit: bool | None = None, num_rma_sig: int | None = None, rma_eager_init: bool | None = None, host_cft_mode: NcclHostCftMode | None = None, nvls_host_mode: NcclNvlsHostMode | None = None)
Bases:
LowppSpecNCCL configuration for communicator initialization.
Provides configuration options for NCCL communicators, allowing fine-tuning of performance and behavior characteristics. Fields not set in the constructor remain at NCCL’s internal default; values are validated by the C library when the config is consumed.
See also
ncclConfig_tfor the description of each field.- blocking: bool | None = None
Blocking (True) or non-blocking (False) communicator behavior. If unset, NCCL uses True.
Available since NCCL 2.14.0.
- cga_cluster_size: int | None = None
Cooperative Group Array (CGA) size for kernels (0-8). If unset, NCCL uses 4 for sm90+, 0 otherwise.
Available since NCCL 2.17.0.
- min_ctas: int | None = None
Minimum number of CTAs per kernel; positive integer up to 32. If unset, NCCL uses 1.
Available since NCCL 2.17.0.
- max_ctas: int | None = None
Maximum number of CTAs per kernel; positive integer up to 32. If unset, NCCL uses 32.
Available since NCCL 2.17.0.
- net_name: str | None = None
Network module name (e.g. ‘IB’, ‘Socket’). Case-insensitive. If unset, NCCL auto-selects.
Available since NCCL 2.17.0.
Share resources with the child communicator during split. If unset, NCCL uses False.
Available since NCCL 2.18.0.
- traffic_class: int | None = None
Traffic class (TC) for network operations (>= 0). Network-specific meaning.
Available since NCCL 2.26.0.
- comm_name: str | None = None
User-defined communicator name for logging and profiling.
Available since NCCL 2.27.0.
- collnet_enable: bool | None = None
Enable (True) or disable (False) IB SHARP. If unset, NCCL uses False.
Available since NCCL 2.27.0.
- cta_policy: CTAPolicy | None = None
CTA scheduling policy. If unset, NCCL uses CTAPolicy.DEFAULT.
Available since NCCL 2.27.0.
Share resources with the child communicator during shrink. If unset, NCCL uses False.
Available since NCCL 2.27.0.
- nvls_ctas: int | None = None
Total number of CTAs for NVLS kernels (positive integer). If unset, NCCL auto-determines.
Available since NCCL 2.27.1.
- n_channels_per_net_peer: int | None = None
Number of network channels for pairwise communication. Positive integer, rounded up to power of 2. If unset, NCCL uses an AlltoAll-optimized value.
Available since NCCL 2.28.0.
- nvlink_centric_sched: bool | None = None
Enable NVLink-centric scheduling. If unset, NCCL uses False.
Available since NCCL 2.28.2.
- graph_usage_mode: int | None = None
Graph usage mode. Supported values are 0 (no graphs), 1 (one graph), and 2 (multiple graphs or a mix of graph and non-graph). If unset, NCCL uses 2.
Available since NCCL 2.29.0.
- num_rma_ctx: int | None = None
Number of RMA contexts. Positive integer. If unset, NCCL uses 1.
Available since NCCL 2.29.0.
- max_p2p_peers: int | None = None
Maximum number of peers any rank will concurrently communicate with using P2P. Positive integer. If unset, NCCL uses the communicator size.
Available since NCCL 2.30.0.
- graph_stream_ordering: int | None = None
Whether NCCL preserves stream-ordering semantics for collectives captured into CUDA graphs. Supported values are 0 (disabled) or 1 (enabled). The value 0 cannot be combined with
graph_usage_mode=2. Also controllable via theNCCL_GRAPH_STREAM_ORDERINGenvironment variable. If unset, NCCL uses 1.Available since NCCL 2.30.5.
- launch_order_implicit: bool | None = None
Whether this communicator takes part in implicit launch ordering. Within one CUDA context, operations on communicators that enable it must not overlap with operations on communicators that do not. Also controllable via the
NCCL_LAUNCH_ORDER_IMPLICITenvironment variable, which takes precedence. If unset, NCCL uses False.Available since NCCL 2.31.0.
- num_rma_sig: int | None = None
Number of one-sided RMA signal indexes available per context. Non-negative integer; bounds the
signal_indexaccepted by the signal and wait-signal operations. If unset, NCCL uses 1.Available since NCCL 2.31.0.
- rma_eager_init: bool | None = None
Whether the collective one-sided RMA signal setup is initialized at communicator creation rather than at the first window registration. True is required if the communicator issues signal or wait-signal operations without first registering a symmetric window. Also controllable via the
NCCL_RMA_EAGER_INITenvironment variable, which takes precedence. If unset, NCCL uses False.Available since NCCL 2.31.0.
- host_cft_mode: NcclHostCftMode | None = None
Host-side Compute Fabric Transport mode. Controls whether the communicator creates the CUDA fabric logical endpoints backing the host-side CFT queries. If unset, NCCL uses
NcclHostCftMode.DEFAULT.Available since NCCL 2.31.1.
- nvls_host_mode: NcclNvlsHostMode | None = None
Host-side NVLS mode. Selects which host NVLS components the communicator uses. If unset, NCCL uses its library-defined default, which is currently equivalent to
NcclNvlsHostMode.ENABLE.Available since NCCL 2.32.0.
NcclHostCftMode
Value of NCCLConfig.host_cft_mode.
- class nccl.core.NcclHostCftMode(value, names=<not given>, *values, module=None, qualname=None, type=None, start=1, boundary=None)
Bases:
IntEnumHost-side Compute Fabric Transport (CFT) mode, mirroring
ncclHostCftMode_t.Set on
NCCLConfig.host_cft_modeto control whether the communicator creates the CUDA fabric logical endpoints that back the host-side CFT queries.- DEFAULT = -2147483648
Use the version-specific default.
- ENABLE = 1
Enable host-side CFT support, creating the communicator’s unicast and multicast logical endpoints during the first window registration.
- DISABLE = 2
Disable host-side CFT support.
- FALLBACK = 3
Try to create the logical endpoints; on error, disable host-side CFT instead of failing.
NcclNvlsHostMode
Value of NCCLConfig.nvls_host_mode.
- class nccl.core.NcclNvlsHostMode(value, names=<not given>, *values, module=None, qualname=None, type=None, start=1, boundary=None)
Bases:
IntFlagHost-side NVLS mode, mirroring
ncclNvlsHostMode_t.Set on
NCCLConfig.nvls_host_modeto select which host NVLS components the communicator uses. The twoDISABLE_members combine; leave the field unset for the library default.- ENABLE = 0
Enable all host NVLS components.
- DISABLE_TRANSPORT = 1
Disable the NVLS transport and the registered-buffer optimization.
- DISABLE_SYMMETRIC_MULTIMEM = 2
Disable multimem in NCCL’s symmetric kernels and copy-engine paths.
- DISABLE = 2147483647
Disable all present and future host NVLS components.
NCCLCollConfig
Accepted as the config argument of every collective on
Communicator. See the individual field documentation for unset
behavior and usage requirements.
- class nccl.core.NCCLCollConfig(*, min_ctas: int | None = None, max_ctas: int | None = None, nvls_ctas: int | None = None, cga_cluster_size: int | None = None, alg_selection: str | None = None, force_alg_selection: bool | None = None, cta_policy: CTAPolicy | None = None, user_profiler_tag: int | None = None, launch_completion_event: Event | int | None = None, vendor_options: tuple[VendorOption, ...] = ())
Bases:
LowppSpecPer-call configuration for a single collective.
Accepted as the
configargument of every collective onCommunicator, tuning that one call. The same configuration must be set on every rank; NCCL validates it only locally, when the call is issued.See also
- min_ctas: int | None = None
Lower bound on channels/CTAs for this call. Also set by
NCCL_MIN_CTAS, which takes precedence. If unset, inheritsNCCLConfig.min_ctas.Available since NCCL 2.31.0.
- max_ctas: int | None = None
Upper bound on channels/CTAs for this call, clamped to the communicator’s
max_ctas. Also set byNCCL_MAX_CTAS, which takes precedence. If unset, inheritsNCCLConfig.max_ctas.Available since NCCL 2.31.0.
- nvls_ctas: int | None = None
NVLS-pool-specific channel cap for this call. Also set by
NCCL_NVLS_NCHANNELS, which takes precedence. If unset, inheritsNCCLConfig.nvls_ctas.Available since NCCL 2.31.0.
- cga_cluster_size: int | None = None
CUDA thread-block-cluster size (0-8, Hopper+). Inconsistent values within one group are undefined behavior. Also set by
NCCL_CGA_CLUSTER_SIZE, which takes precedence. If unset, inheritsNCCLConfig.cga_cluster_size.Available since NCCL 2.31.0.
- alg_selection: str | None = None
Selection string filtering which algorithms this call may use, e.g.
"ring","tree,ring","^symk". If unset or empty, NCCL selects automatically.Available since NCCL 2.31.0.
- force_alg_selection: bool | None = None
Whether an unsatisfiable
alg_selectionis an error rather than a fallback to automatic selection. If unset, NCCL uses True.Available since NCCL 2.31.0.
- cta_policy: CTAPolicy | None = None
CTA scheduling policy for this call. Also set by
NCCL_CTA_POLICY, which takes precedence. If unset, inheritsNCCLConfig.cta_policy.Available since NCCL 2.31.0.
- user_profiler_tag: int | None = None
Opaque value delivered verbatim to profiler plugins with this call’s profiler events; does not affect execution. Values with the most-significant bit set are reserved by NCCL. If unset, NCCL uses 0.
Available since NCCL 2.31.0.
- launch_completion_event: NcclEventSpec | None = None
Caller-owned, rank-local CUDA event recorded at collective kernel launch completion. With CUDA versions earlier than 12.3, NCCL records the event before the kernel launch instead. A
cuda.core.Eventcreated byDevice().create_event()has the required timing-disabled configuration; events passed as integer handles must likewise have timing disabled. Interprocess and interop events are unsupported. Either every rank passes an event or none does, and at most one per communicator in a group. Keep it alive through all queued waits and captured-graph executions. If unset, NCCL records no event.Available since NCCL 2.32.0.
- vendor_options: tuple[VendorOption, ...] = ()
Vendor-specific options;
(vendor_id, option_id)keys must be unique.Available since NCCL 2.31.0.
VendorOption
- class nccl.core.VendorOption(vendor_id: int, option_id: int, int_value: int | None = None, str_value: str | None = None, raw_value: int | None = None)
Bases:
objectA single vendor-specific option attached to an
NCCLCollConfig.Mirrors one
ncclConfigExt_tnode. Options are identified by the(vendor_id, option_id)pair; the official NCCL library ignores every extension, so an option only has an effect on a vendor library that recognizes itsvendor_id. Vendors pick a non-zerovendor_idless than 2**24 that is unlikely to collide.Exactly one of the three value fields must be set.
See also
- vendor_id: int
Vendor-chosen identifier, unique across vendor libraries.
- option_id: int
Vendor-defined identifier distinguishing options within a vendor.
- int_value: int | None = None
Integer value (
val.i).
- str_value: str | None = None
String value (
val.s), encoded to UTF-8.
- raw_value: int | None = None
Value of any other type (
val.raw), as an integer. If it is an address, the referent must stay valid for the duration of the call.
CTAPolicy
- class nccl.core.CTAPolicy(value, names=<not given>, *values, module=None, qualname=None, type=None, start=1, boundary=None)
Bases:
IntFlagNCCL performance policy for CTA scheduling, used by
NCCLConfig.cta_policyandNCCLCollConfig.cta_policy.- DEFAULT = 0
Default CTA policy.
- EFFICIENCY = 1
Optimize for efficiency.
- ZERO = 2
Zero-CTA optimization.
NCCLDevCommRequirements
Used by Communicator.create_dev_comm(). Fields left unset
(None) remain at NCCL’s internal default.
- class nccl.core.NCCLDevCommRequirements(*, lsa_multimem: bool | None = None, barrier_count: int | None = None, lsa_barrier_count: int | None = None, rail_gin_barrier_count: int | None = None, lsa_ll_a2a_block_count: int | None = None, lsa_ll_a2a_slot_count: int | None = None, gin_force_enable: bool | None = None, gin_context_count: int | None = None, gin_signal_count: int | None = None, gin_counter_count: int | None = None, gin_connection_type: NcclGinConnectionType | None = None, gin_exclusive_contexts: bool | None = None, gin_queue_depth: int | None = None, gin_traffic_class: int | None = None, world_gin_barrier_count: int | None = None, gin_strong_signals_required: bool | None = None, gin_va_signals_required: bool | None = None, gin_custom_stride: int | None = None, gin_type: NcclGinType | None = None, cft_caps: NcclCftCap | None = None, cft_barrier_count: int | None = None, teams: tuple[TeamRequirement, ...] = (), resources: tuple[LsaBarrierRequirement | GinBarrierRequirement | LLA2ARequirement, ...] = ())
Bases:
LowppSpecNCCL device communicator requirements configuration.
This is a reusable high-level Python request consumed by
Communicator.create_dev_comm(). Per-team requirements are declared through theteamstuple. Each call snapshots the request into independent low-levelncclDevCommRequirements_tand linkedncclTeamRequirements_tstorage, including separate multimem output handles. NCCL copies the requirements and linked-list nodes before the call returns; the resultingDevCommResourceretains the storage referenced by eachoutMultimemHandle. This object may therefore be changed between calls without affecting device communicators that were already created. Do not mutate it concurrently withCommunicator.create_dev_comm().See also
ncclDevCommRequirementsfor the description of each field.- lsa_multimem: bool | None = None
Enable multimem on the LSA team. If unset, NCCL uses False.
- barrier_count: int | None = None
Number of barriers required. If unset, NCCL uses 0.
- lsa_barrier_count: int | None = None
Number of LSA barriers. If unset, NCCL uses 0.
- rail_gin_barrier_count: int | None = None
Number of railed GIN barriers. If unset, NCCL uses 0.
- lsa_ll_a2a_block_count: int | None = None
LSA low-latency all-to-all block count. If unset, NCCL uses 0.
- lsa_ll_a2a_slot_count: int | None = None
LSA low-latency all-to-all slot count. If unset, NCCL uses 0.
- gin_force_enable: bool | None = None
Force-enable GPU-Initiated Networking (GIN). If unset, NCCL uses False.
- gin_context_count: int | None = None
Number of GIN contexts (hint; actual count may differ). If unset, NCCL uses 4.
- gin_signal_count: int | None = None
Number of GIN signals (guaranteed to start at id=0). If unset, NCCL uses 0.
- gin_counter_count: int | None = None
Number of GIN counters (guaranteed to start at id=0). If unset, NCCL uses 0.
- gin_connection_type: NcclGinConnectionType | None = None
GIN connection type. If unset, NCCL uses NcclGinConnectionType.NONE.
- gin_exclusive_contexts: bool | None = None
Use exclusive GIN contexts. If unset, NCCL uses False.
- gin_queue_depth: int | None = None
GIN queue depth. If unset, NCCL uses 0.
- gin_traffic_class: int | None = None
GIN traffic class. If unset, NCCL uses its internal default.
- world_gin_barrier_count: int | None = None
Number of world GIN barriers. If unset, NCCL uses 0.
- gin_strong_signals_required: bool | None = None
Whether GIN strong signals are required by kernels using this devComm. When False, using GIN strong signals results in undefined behavior. If unset, NCCL uses True.
- gin_va_signals_required: bool | None = None
Whether GIN VA signals are required by kernels using this devComm. When False, using GIN VA signals results in undefined behavior. If unset, NCCL uses True.
- gin_custom_stride: int | None = None
Stride of ranks to connect for GIN. Only consulted when
gin_connection_typeisNcclGinConnectionType.CUSTOM_STRIDE, and must be a multiple ofNCCLCommProperties.gin_min_stride. If unset, NCCL uses 1.
- gin_type: NcclGinType | None = None
GIN transport to require. If unset, NCCL uses
NcclGinType.NONE, accepting any available transport.
- cft_caps: NcclCftCap | None = None
Compute Fabric Transport capabilities to request, as a bitmask of
NcclCftCapvalues. Creation fails if CFT resources are requested on a communicator where not all ranks support CFT. If unset, NCCL usesNcclCftCap.NONE.
- cft_barrier_count: int | None = None
Number of CFT barriers to allocate, one per independently addressed barrier slot the kernel uses (commonly one per CTA). If unset, NCCL uses 0.
- teams: tuple[TeamRequirement, ...] = ()
Per-team requirements. Entries for the same team (by value) are merged, keeping first-appearance order; multimem is requested for a team if any of its entries sets it. A team requested with
multimem=Trueyields a multimem handle retrievable viamultimem_handle().
- resources: tuple[LsaBarrierRequirement | GinBarrierRequirement | LLA2ARequirement, ...] = ()
Device resource requirements (LSA/GIN barriers, low-latency all-to-all). Each entry yields, in order, a handle in
resource_handles. Entries are kept as-is (not merged): each is a distinct resource.
NcclCftCap
Bitmask value of NCCLDevCommRequirements.cft_caps.
- class nccl.core.NcclCftCap(value, names=<not given>, *values, module=None, qualname=None, type=None, start=1, boundary=None)
Bases:
IntFlagCompute Fabric Transport capabilities, mirroring
ncclCftCap_t.Combined as a bitmask on
NCCLDevCommRequirements.cft_caps.- NONE = 0
No CFT capability requested.
- CFT = 1
Request unicast CFT logical endpoints.
- MULTIMEM = 2
Request multicast CFT operations and multimem CFT barriers.
Requirement entries
The element types of NCCLDevCommRequirements.teams and
NCCLDevCommRequirements.resources.
TeamRequirement
- class nccl.core.TeamRequirement(team: NCCLTeam, multimem: bool = False)
Bases:
objectA per-team requirement for device communicator creation.
Pass a tuple of these as
NCCLDevCommRequirements.teams. Whenmultimemis True, NCCL allocates a multicast handle for the team, retrievable afterwards viamultimem_handle().- multimem: bool = False
LsaBarrierRequirement
- class nccl.core.LsaBarrierRequirement(team: NCCLTeam, n_barriers: int)
Bases:
objectRequests an LSA barrier resource on
teamwithn_barriersbarriers.Add to
NCCLDevCommRequirements.resources; the finalizedLsaBarrierHandleis returned inresource_handles.- n_barriers: int
GinBarrierRequirement
- class nccl.core.GinBarrierRequirement(team: NCCLTeam, n_barriers: int)
Bases:
objectRequests a GIN barrier resource on
teamwithn_barriersbarriers.Add to
NCCLDevCommRequirements.resources; the finalizedGinBarrierHandleis returned inresource_handles.- n_barriers: int
LLA2ARequirement
- class nccl.core.LLA2ARequirement(n_blocks: int, max_elements: int, max_element_size: int)
Bases:
objectRequests a low-latency all-to-all resource with
n_blocksblocks, sized to hold up tomax_elementselements of at mostmax_element_sizebytes each.Add to
NCCLDevCommRequirements.resources; the finalizedLLA2AHandleis returned inresource_handles.- n_blocks: int
- max_elements: int
- max_element_size: int