Using CFT Handles with NVSHMEM#
CUDA Compute Fabric Transport (CFT) handles let NVSHMEM address symmetric memory through CUDA logical endpoints instead of requiring every peer heap to be directly mapped into the caller’s virtual address space. This is useful on Multi-Node NVLink (MNNVL) systems where a GPU can be fabric-reachable even when the regular pointer path is unavailable or intentionally limited.
CFT handle support does not add new put, get, atomic, signal, collective, or tile APIs. Applications continue to pass NVSHMEM symmetric pointers. When CFT support is built, enabled at run time, and eligible for a particular operation, NVSHMEM translates the symmetric pointer into a logical-endpoint ID plus a heap-relative offset and issues CUDA fabric operations. If the handle path is not available for an operation, NVSHMEM uses the existing pointer, NVLS, or transport path.
Build-Time Setup#
CFT handle support is compiled out by default. Build NVSHMEM with:
cmake -DNVSHMEM_CFT_HANDLES_SUPPORT=ON <other-options> ..
To make eligible handle paths take priority even when a direct peer pointer is available, also build with:
cmake -DNVSHMEM_CFT_HANDLES_SUPPORT=ON \
-DNVSHMEM_PRIORITIZE_LOGICAL_ENDPOINT=ON \
<other-options> ..
NVSHMEM_PRIORITIZE_LOGICAL_ENDPOINT is useful for validating the CFT path.
Without this build option, ordinary unicast operations prefer a reachable peer
pointer and use a CFT handle only when the pointer path is unavailable and the
handle path is eligible.
Application Setup#
CFT uses TMA for shared-memory staging and fabric completion tracking, so the
TMA setup requirements also apply to CFT. See Using TMA with NVSHMEM for the detailed
nvshmemx_ask_smem(), nvshmemx_give_smem(), dynamic shared-memory launch,
and nvshmemx_release_smem() usage.
In short, each CTA that may execute a CFT-backed device operation must donate CTA shared memory to NVSHMEM before issuing that operation and release the registration before the CTA exits. Point-to-point CFT PUT sources and GET destinations may be in global memory or in application shared memory. A global-memory operand is staged through the donated shared memory, while an eligible shared-memory PUT source or GET destination can be used directly by the CUDA fabric instruction.
Run-Time Setup#
Enable logical endpoints and TMA before NVSHMEM initialization:
export NVSHMEM_ENABLE_LOGICAL_ENDPOINT=1
export NVSHMEM_TMA_POLICY=ENABLE
NVSHMEM_ENABLE_LOGICAL_ENDPOINT=1Enables CUDA logical-endpoint discovery, endpoint creation, endpoint exchange, and symmetric-heap binding. The default is
0.NVSHMEM_TMA_POLICY=ENABLEEnables the TMA support required by the CFT path. Use
ENABLEfor typical runs.FORCEis also accepted for validation on supported systems. Initialization fails if logical endpoints are enabled while TMA is disabled.
For bring-up, NVSHMEM_DEBUG=INFO can help confirm whether logical endpoint
support was enabled and which paths were selected.
Platform Requirements#
CFT handles require:
NVSHMEM built with
NVSHMEM_CFT_HANDLES_SUPPORT=ON.CUDA Runtime 13.3 or newer for bulk CFT put, get, signal, tile, and multicast operations.
CUDA Runtime 13.4 or newer for CFT atomics and CUDA fabric-clique discovery.
Device code compiled for SM100 or newer.
NVSHMEM_TMA_POLICY=ENABLEorFORCEand registered CTA shared memory.
Limitations#
CFT handle support currently has the following limitations:
CFT handles do not provide fabric error reporting or recovery.
Eligible cooperative warp- and block-scoped put/get operations, supported atomics, SET signal operations, unicast tile put/get, and NVLS tile allgather/fcollect, broadcast, and reduce/allreduce operations may use CFT handles. For non-tile collectives, dedicated CFT paths are limited to barrier/sync signaling, NVLS simple fcollect, NVLS one-shot and two-shot SUM reductions, and NVLS reduce-scatter for supported operation/type pairs. All-to-all, broadcast, and non-NVLS collective algorithms do not have dedicated multicast CFT paths, although eligible non-thread put steps may inherit unicast CFT routing.
Standard single-thread bulk put/get operations and scalar
gdo not use CFT handles. Scalarpoperations may use CFT handles when eligible.Bulk CFT operations can use the handle path only when the target peer or team is reachable through CFT, the CTA has registered shared memory, and the remote symmetric address and local PUT source or GET destination address are both 16-byte aligned.
GET and some collective CFT paths require a 16-byte-multiple payload size. PUT and PUT NBI can handle payload sizes that are not multiples of 16 bytes.
Handle-backed NBI PUT defers final fabric completion to
nvshmem_quiet. Handle-backed NBI GET waits before returning because the local destination must already be valid when the call completes.For handle-backed NBI PUTs issued by a CTA, use one elected CTA thread to call device
nvshmem_quiet()after the issuing operation returns, then synchronize the CTA before reusing or releasing the donated shared memory. Concurrentnvshmem_quiet()calls from multiple CTA threads are not supported for draining shared handle-barrier state.CFT atomics require CUDA 13.4 or newer, registered CTA shared memory, a target peer that is reachable through CFT, a supported atomic operation/type mapping, and a 16-byte-aligned destination block. Unsupported or resource-ineligible atomics retain the existing pointer or transport implementation.
EGM and static or non-VMM memory are not supported for logical endpoints. CFT handle memory must use supported CUDA VMM device allocations.
A CTA should serialize CFT bulk operations and CFT atomics because their donated shared-memory staging regions may overlap. This is not required when bulk operations use user-controlled shared memory, because the internal staging buffer is not used.