NVIDIA OpenSHMEM Library (NVSHMEM) Documentation#
NVSHMEM implements the OpenSHMEM parallel programming model for clusters of NVIDIA ® GPUs. The NVSHMEM Partitioned Global Address Space (PGAS) spans the memory across GPUs and includes an API for fine-grained GPU-GPU data movement from within a CUDA kernel, on CUDA streams, and from the CPU.
Contents:
- Introduction
- Installing NVSHMEM
- Using NVSHMEM
- NVSHMEM and the CUDA Model
- Using Locality Domains with NVSHMEM
- Using TMA with NVSHMEM
- Using CFT Handles with NVSHMEM
- Using Counted Writes with NVSHMEM
- Memory Model
- Execution Model
- Library Constants
- Library Handles
- Environment Variables
- NVSHMEM APIs
- Overview of the APIs
- Library Setup, Exit, and Query
- Thread Support
- Kernel Launch Routines
- Memory Management
- Queue Pair (QP) Specific APIs
- Team Management
- Remote Memory Access
- Atomic Memory Operations
- Signaling Operations
- Collective Communication
- Point-To-Point Synchronization
- Memory Ordering
- Regions
- Language Bindings
- Examples
- Language Bindings Examples
- Attribute-Based Initialization Example
- Collective Launch Example
- On-Stream Example
- Threadgroup Example
- Put on Block Example
- Regions Example
- TMA Put Example
- Threadgroup Fence and Signal Example
- Ring Broadcast Example
- Ring Allreduce Example
- User Buffer Registration Example
- GEMM + AllReduce Fused Kernel Example
- Troubleshooting and FAQs
- License
- Acknowledgements