Using Locality Domains with NVSHMEM#
A locality domain groups memory and execution resources within a physical GPU. Applications can place work and memory in the same domain to improve locality. NVSHMEM supports CUDA locality domains through NVLink utilization- centric scheduling, application-managed programmatic localization, and MPS MLOPart.
Requirements#
Locality-domain support requires:
CUDA Toolkit 13.4 or later for NVSHMEM and the application.
An NVIDIA driver compatible with that toolkit.
A GPU and system configuration that report locality-domain support.
Applications must query the available locality domains instead of assuming a fixed count. Use the CUDA Programming Guide locality-domain documentation for CUDA discovery, allocation, green-context, and MPS setup.
Choosing a Model#
Model |
PE mapping |
Placement |
NVSHMEM integration |
|---|---|---|---|
NVLink utilization-centric scheduling (NUCS) |
One PE per physical GPU |
CUDA distributes the full grid’s CTAs across locality domains. Memory remains ordinary and nonlocalized. |
Pass the NUCS attribute to NVSHMEMX_COLLECTIVE_LAUNCH_ATTR. |
Programmatic localization |
One PE per physical GPU |
The application creates localized VMM allocations and locality-bound green-context streams. |
Register each allocation with
|
MPS MLOPart |
One PE per exposed MLOPart CUDA device |
MPS localizes execution and default allocations to the selected MLOPart. |
Configure MPS before starting the NVSHMEM application. NVSHMEM uses its existing allocation and RMA APIs. |
NVLink Utilization-Centric Scheduling#
NUCS is a scheduling hint. CUDA distributes the CTAs of one full kernel grid across the GPU’s locality domains. It does not create a grid or PE for each domain, and it does not localize memory.
Set cudaLaunchAttributeNvlinkUtilCentricScheduling in a CUDA launch
attribute, place that attribute in nvshmemx_collective_launch_attr_t, and
launch the kernel with nvshmemx_collective_launch_attr:
cudaLaunchAttribute cuda_attr{};
cuda_attr.id = cudaLaunchAttributeNvlinkUtilCentricScheduling;
cuda_attr.val.nvlinkUtilCentricScheduling = 1;
nvshmemx_collective_launch_attr_t launch_attr{};
launch_attr.cuda_config.gridDim = grid;
launch_attr.cuda_config.blockDim = block;
launch_attr.cuda_config.dynamicSmemBytes = shared_mem;
launch_attr.cuda_config.stream = stream;
launch_attr.cuda_config.attrs = &cuda_attr;
launch_attr.cuda_config.numAttrs = 1;
void *args[] = {&destination, &source, &count};
int status = nvshmemx_collective_launch_attr(
&launch_attr, (const void *)kernel, args);
This replaces nvshmemx_collective_launch for that launch. The call affects
only the GPU associated with the calling PE and does not coordinate with other
PEs. If the kernel invokes an NVSHMEM synchronization or collective operation,
every required PE must launch and participate according to that operation’s
normal rules. The cooperative grid-size requirements still apply. NUCS does
not change NVSHMEM memory placement or transport selection.
Programmatic Localization#
Programmatic localization gives the application explicit control over memory and execution placement while retaining one NVSHMEM PE per physical GPU. For each selected locality domain, the application:
Uses CUDA VMM to create and map a locality-domain allocation.
Creates a green context and stream using SM resources from the same domain.
Calls
nvshmemx_buffer_register_symmetric(local_mapping, size, 0)collectively and in the same order on every PE.Uses the returned symmetric alias with NVSHMEM RMA operations. The original CUDA mapping is not a symmetric pointer.
Calls
nvshmemx_buffer_unregister_symmetric(symmetric_alias, size)collectively before releasing the CUDA allocation.
This registration requires the dynamic VMM NVSHMEM heap. Allocation sizes and registration order must be symmetric across PEs. The CUDA allocation’s handle type and granularity must also be compatible with the NVSHMEM heap.
NVSHMEM does not create or manage locality domains, localized allocations, green contexts, or locality-bound streams. When an application runs concurrent NVSHMEM collectives on different localized streams, each concurrent collective must use a distinct NVSHMEM team.
MPS MLOPart#
MPS MLOPart exposes each locality domain as a CUDA device. Start and configure the MPS daemon in MLOPart mode before launching the application, then launch one NVSHMEM PE per exposed device. Execution and default allocations are localized automatically. See the CUDA Programming Guide locality-domain documentation for MPS MLOPart configuration details.
A VMM allocation created with
CUmemAllocationProp::allocFlags.gpuDirectRDMACapable opts out of MLOPart
localization and can be exported for RDMA. That allocation does not receive
locality-domain placement benefits. cudaMalloc has no equivalent opt-out.
NVSHMEM detects when multiple PEs occupy one physical GPU and disables NVLink SHARP (NVLS). NVLS multicast cannot span MLOPart devices on the same physical GPU.
Communication Restrictions#
Warning
Localized buffers support only direct NVLink/P2P communication. They cannot be used with NVSHMEM remote transports.
Network communication that begins or ends in localized memory requires application-managed staging. Copy between the localized buffer and a separate nonlocalized, RDMA-capable symmetric buffer as needed, then issue the remote NVSHMEM operation using that staging buffer. NVSHMEM does not stage localized buffers implicitly.
Programmatically localized buffers do not support NVLS. NVSHMEM also disables NVLS for MPS MLOPart and other multiple-PE-per-GPU configurations. NUCS uses ordinary nonlocalized memory, so these localized-buffer restrictions do not apply to NUCS launches.