NCCL Release 2.31.2

This is the NCCL 2.31.2 release notes. For previous NCCL release notes, refer to the NCCL Archives.

Compatibility

NCCL 2.31.2 has been tested with the following:

Key Features and Enhancements

This NCCL release includes the following key features and enhancements.

  • Compute Fabric Transport (CFT): Added CFT host and device APIs for registering window memory with CUDA logical endpoints and issuing device-side Put, Get, Red and NVLS operations. CFT is supported on Blackwell GPUs with CUDA toolkit 13.3 or later.

  • Per-Collective Configuration and Tuning: Added new ncclCollConfig_t and nccl*Config APIs for all collectives, with a vendor-defined field to make it easier for custom forks to integrate. Supports per-collective algorithm selection, CTA/CGA size and CTA policy overrides. Leverages collConfig to add userTag in profiler API, motivated by community RFE https://github.com/NVIDIA/nccl/issues/1916.

  • GIN Enhancements: AWS-EFA team contributed the EFA GDA backend for GIN Put, Signal, and Flush operations (Github PR #2273). Added per-DevComm GIN backend selection, allowing users to create multiple DevComms with different backends. Added support for connecting GIN with custom strides. Added device-side timeouts to blocking GIN APIs such as Flush, Wait, WaitSignal, and WaitCounter operations. Reduced QP usage when using railed GIN. Reduced file-descriptor consumption in GDAKI when using many QPs. Optimized error reporting in GDAKI via event-based CQ error reporting. Added support for out-of-order delivery (DDP) in GDAKI on SPCX, improving performance without impacting GIN correctness guarantees. Added GRH support and automatic path-MTU discovery to the GIN GDAKI backend.

  • Parallel Aggregated Tree (PAT) Enhancements: Enhanced the PAT algorithm for ReduceScatter/AllGather by adding hierarchical kernels that use NVLS within the node and PAT across nodes. Enables better small/medium message performance. Disabled by default. Use NCCL_ALGO=PAT, or the collConfig API to enable.

  • One-Sided RMA and Copy Engine Collectives: Added support for multiple contexts and multiple signals to one-sided RMA operations, allowing one-sided RMA traffic from a single rank to use multiple NICs. Leverage multiple contexts and signals for hierarchical 0-SM AllGather and AllToAll. Optimizes small message latency by using NVLink multicast for AllGather.

  • Tuning and Cost Model: Introduced a new unified cost model interface for querying the cost estimate for both legacy and device-API-based kernels. On Blackwell, TMA kernels are now enabled by default and integrated in the cost model for symmetric registered memory.

  • Diagnostics and Profiling: Added NCCL RAS Diagnostics checks through NCCL_RUN_RAS_DIAGNOSTICS=1 or RAS client, including GPU inventory, CUDA driver versions, ECC errors, NVLink state, and NCCL_* environment consistency checks. Added NCCL Diagnostics through NCCL_RUN_DIAGNOSTICS=1, including active P2P connectivity check and actionable P2P remediation guidance. Moved Kernel-channel profiling to a dedicated per-communicator thread, extending kernel timing to proxy-less NVLink/SHM and graph-captured collectives. Profiler V7 exposes per-kernel initial_sync, compute and final_sync phase events and symmetric-kernel variant metadata. Added per-QP CPU WQE post-to-poll latency monitoring for the IB transport.

  • Added backward compatibility support for applications that JIT-compile NCCL Device API code.

  • Added support for LTO IR for NCCL device API.

  • Reduced communicator host-memory use and topology initialization time on large systems by allocating topology paths according to their actual lengths.

  • Added multiple GIN proxy progress threads with per-thread endpoint assignment through GIN_PROXY_NTHREADS (Github PR #2279).

  • Improved GIN host-proxy throughput by processing multiple GIN operations per progress iteration (Github PR #2232) and adding hints to the RMA plugin to aggregate requests (Github PR #2254).

  • Added CuTeDSL bindings for GIN Get, Flush, Signal, ReadSignal, and PutValue operations (Github PR #2266).

  • Added an internal CUDA 13.3 DMA-BUF mmap backend for GDRCopy.

  • Added event-based load balancing for the net_ib transport.

  • Refreshed local GIDs after IBV_EVENT_GID_CHANGE events during IB port recovery.

  • Updated the RMA plugin interface to v15 (Github PR #2254).

  • Improved AMD EPYC topology modeling (Github PR #2036).

  • Improved algorithm selection on newer Intel CPUs.

  • Made nccl_device.h compatible with C99 host translation units.

  • Contrib updates: Added PACE, a parallelism-aware collective engine that fuses layout conversion, data-type conversion, and scatter/gather operations into collective kernels (Github PR #2319). Added experimental nccl4rust host and Device API bindings using LTO IR. Added NIIN, a header-only NVSHMEM-compatible interface implemented with NCCL Device API primitives. Added community-contributed architecture learning guides for communicator initialization, topology, tuning, transports, and collective execution (Github PR #2081). Added a community-contributed analysis of DeepEPv2.

Fixed Issues

The following issues have been resolved in NCCL 2.31.2:

  • Fixed a hang when PXN connection initialization races with communicator teardown.

  • Fixed data corruption when multiple symmetric windows are carved from the same backing memory allocation (Github Issue #2198).

  • Fixed a hang during the first Copy Engine collective on a two-rank communicator (Github Issue #2241).

  • Fixed NIC/GPU assignment perf regressions by limiting consistent start-NIC selection to Blackwell systems with ConnectX-8.

  • Fixed missing GIN signal and counter requirements when Device API communicators are created asynchronously (Github PR #2208).

  • Fixed a memory leak when Device API communicator creation fails (Github PR #2225).

  • Fixed an incorrect GIN context count reported to applications when the allocated context count is rounded up (Github PR #2301).

  • Fixed a resource leak when GIN connection setup fails (Github PR #2206).

  • Fixed the CUDA thread capture mode remaining changed when GIN setup fails (Github PR #2229).

  • Fixed inconsistent NVLS enablement when multiple ranks share a GPU (Github PR #2257).

  • Fixed valid communicator configurations being rejected when only one of minCTAs or maxCTAs is set (Github PR #2256).

  • Fixed communicator initialization failures on multi-system and MNNVL topologies when NCCL_P2P_PXN_LEVEL=1 (Github PR #2258).

  • Fixed a Device API loadConst performance regression caused by __ldg.

  • Fixed CuTeDSL GIN barrier initialization failures caused by under-aligned by-value arguments (Github PR #2243, Issue #2242).

Known Issues

  • PAT + NVLS has performance regressions on H100 platforms. This will be fixed in the next release.

  • B40 symmetric TMA kernels: On B40 GPUs, TMA-based symmetric kernels can exceed available shared memory and fail with an illegal memory access. Set NCCL_SYM_TMA_ENABLE=0 to disable these kernels.

  • B100 PCIe MLoPart: Multi-rank-per-GPU workloads using MPS MLoPart on B100 PCIe systems can fail during memory allocation with CUDA error 101 (invalid device ordinal). Set NCCL_CUMEM_ENABLE=0 as a workaround.

Updating the GPG Repository Key

To best ensure the security and reliability of our RPM and Debian package repositories, NVIDIA is updating and rotating the signing keys used by apt, dnf/yum, and zypper package managers beginning on April 27, 2022. Failure to update your repository signing keys will result in package management errors when attempting to access or install NCCL packages. To ensure continued access to the latest NCCL release, please follow the updated NCCL installation guide.