Using Locality Domains with NVSHMEM#

A locality domain groups memory and execution resources within a physical GPU. Applications can place work and memory in the same domain to improve locality. NVSHMEM supports CUDA locality domains through NVLink utilization- centric scheduling, application-managed programmatic localization, and MPS MLOPart.

Requirements#

Locality-domain support requires:

  • CUDA Toolkit 13.4 or later for NVSHMEM and the application.

  • An NVIDIA driver compatible with that toolkit.

  • A GPU and system configuration that report locality-domain support.

Applications must query the available locality domains instead of assuming a fixed count. Use the CUDA Programming Guide locality-domain documentation for CUDA discovery, allocation, green-context, and MPS setup.

Choosing a Model#

Model

PE mapping

Placement

NVSHMEM integration

NVLink utilization-centric scheduling (NUCS)

One PE per physical GPU

CUDA distributes the full grid’s CTAs across locality domains. Memory remains ordinary and nonlocalized.

Pass the NUCS attribute to NVSHMEMX_COLLECTIVE_LAUNCH_ATTR.

Programmatic localization

One PE per physical GPU

The application creates localized VMM allocations and locality-bound green-context streams.

Register each allocation with nvshmemx_buffer_register_symmetric and use the returned symmetric alias.

MPS MLOPart

One PE per exposed MLOPart CUDA device

MPS localizes execution and default allocations to the selected MLOPart.

Configure MPS before starting the NVSHMEM application. NVSHMEM uses its existing allocation and RMA APIs.

Programmatic Localization#

Programmatic localization gives the application explicit control over memory and execution placement while retaining one NVSHMEM PE per physical GPU. For each selected locality domain, the application:

  1. Uses CUDA VMM to create and map a locality-domain allocation.

  2. Creates a green context and stream using SM resources from the same domain.

  3. Calls nvshmemx_buffer_register_symmetric(local_mapping, size, 0) collectively and in the same order on every PE.

  4. Uses the returned symmetric alias with NVSHMEM RMA operations. The original CUDA mapping is not a symmetric pointer.

  5. Calls nvshmemx_buffer_unregister_symmetric(symmetric_alias, size) collectively before releasing the CUDA allocation.

This registration requires the dynamic VMM NVSHMEM heap. Allocation sizes and registration order must be symmetric across PEs. The CUDA allocation’s handle type and granularity must also be compatible with the NVSHMEM heap.

NVSHMEM does not create or manage locality domains, localized allocations, green contexts, or locality-bound streams. When an application runs concurrent NVSHMEM collectives on different localized streams, each concurrent collective must use a distinct NVSHMEM team.

MPS MLOPart#

MPS MLOPart exposes each locality domain as a CUDA device. Start and configure the MPS daemon in MLOPart mode before launching the application, then launch one NVSHMEM PE per exposed device. Execution and default allocations are localized automatically. See the CUDA Programming Guide locality-domain documentation for MPS MLOPart configuration details.

A VMM allocation created with CUmemAllocationProp::allocFlags.gpuDirectRDMACapable opts out of MLOPart localization and can be exported for RDMA. That allocation does not receive locality-domain placement benefits. cudaMalloc has no equivalent opt-out.

NVSHMEM detects when multiple PEs occupy one physical GPU and disables NVLink SHARP (NVLS). NVLS multicast cannot span MLOPart devices on the same physical GPU.

Communication Restrictions#

Warning

Localized buffers support only direct NVLink/P2P communication. They cannot be used with NVSHMEM remote transports.

Network communication that begins or ends in localized memory requires application-managed staging. Copy between the localized buffer and a separate nonlocalized, RDMA-capable symmetric buffer as needed, then issue the remote NVSHMEM operation using that staging buffer. NVSHMEM does not stage localized buffers implicitly.

Programmatically localized buffers do not support NVLS. NVSHMEM also disables NVLS for MPS MLOPart and other multiple-PE-per-GPU configurations. NUCS uses ordinary nonlocalized memory, so these localized-buffer restrictions do not apply to NUCS launches.