Multi-vGPU and P2P#
Multi vGPU#
Multi-vGPU technology allows a single VM to simultaneously use multiple vGPUs, enhancing its computational capabilities. Unlike standard vGPU configurations that virtualize a single physical GPU for sharing across multiple VMs, Multi-vGPU presents resources from several vGPU devices into a single VM. These vGPUs can be time-sliced or MIG-backed. These vGPU devices are not required to reside on the same physical GPU and can be distributed across separate physical GPUs, pooling their collective power to meet the demands of high-performance workloads.
This technology is advantageous for AI training and inference workloads that require extensive computational power. It optimizes resource allocation by enabling applications within a VM to access dedicated GPU resources. For instance, a VM configured with two NVIDIA A100 GPUs using Multi-vGPU can run large-scale AI models more efficiently than with a single GPU. This dedicated assignment eliminates resource contention between different AI processes within the same VM, ensuring optimal and predictable performance for critical tasks. The ability to aggregate computational power from multiple vGPUs makes Multi-vGPU a solution for scaling complex AI model development and deployment.
vGPU Support for Multi-vGPU#
You can assign multiple vGPUs with differing amounts of frame buffer to a single VM, provided the board type and the series of all the vGPUs are the same. For example, you can assign an A40-48C vGPU and an A40-16C timesliced vGPUs to the same VM. You can also assign an A100-4-20C vGPU and one A100-2-10C vGPU to a VM, both on MIG instances from an A100 board. However, you cannot assign an A30-8C vGPU and an A16-8C vGPU to the same VM.
Board |
vGPU [1] |
|---|---|
NVIDIA HGX B200 180GB |
Generic Linux with KVM hypervisors [4], Red Hat Enterprise Linux KVM, and Ubuntu: - All NVIDIA vGPU for Compute |
NVIDIA RTX PRO 6000 Blackwell Server Edition 96GB |
|
Board |
vGPU [1] |
|---|---|
NVIDIA H800 PCIe 94GB (H800 NVL) |
All NVIDIA vGPU for Compute |
NVIDIA H800 PCIe 80GB |
All NVIDIA vGPU for Compute |
NVIDIA H800 SXM5 80GB |
NVIDIA vGPU for Compute [3] |
NVIDIA H200 PCIe 141GB (H200 NVL) |
All NVIDIA vGPU for Compute |
NVIDIA H200 SXM5 141GB |
NVIDIA vGPU for Compute [3] |
NVIDIA H100 PCIe 94GB (H100 NVL) |
All NVIDIA vGPU for Compute |
NVIDIA H100 SXM5 94GB |
NVIDIA vGPU for Compute [3] |
NVIDIA H100 PCIe 80GB |
All NVIDIA vGPU for Compute |
NVIDIA H100 SXM5 80GB |
NVIDIA vGPU for Compute [3] |
NVIDIA H100 SXM5 64GB |
NVIDIA vGPU for Compute [3] |
NVIDIA H20 SXM5 141GB |
NVIDIA vGPU for Compute [3] |
NVIDIA H20 SXM5 96GB |
NVIDIA vGPU for Compute [3] |
Board |
vGPU |
|---|---|
NVIDIA L40 |
|
NVIDIA L40S |
|
NVIDIA L20 |
|
NVIDIA L4 |
|
NVIDIA L2 |
|
NVIDIA RTX 6000 Ada |
|
NVIDIA RTX 5880 Ada |
|
NVIDIA RTX 5000 Ada |
|
Board |
vGPU [1] |
|---|---|
|
|
NVIDIA A800 PCIe 40GB active-cooled |
|
NVIDIA A800 HGX 80GB |
|
|
|
NVIDIA A100 HGX 80GB |
|
NVIDIA A100 PCIe 40GB |
|
NVIDIA A100 HGX 40GB |
|
NVIDIA A40 |
|
|
|
NVIDIA A16 |
|
NVIDIA A10 |
|
NVIDIA RTX A6000 |
|
NVIDIA RTX A5500 |
|
NVIDIA RTX A5000 |
|
Board |
vGPU |
|---|---|
Tesla T4 |
|
Quadro RTX 6000 passive |
|
Quadro RTX 8000 passive |
|
Board |
vGPU |
|---|---|
Tesla V100 SXM2 |
|
Tesla V100 SXM2 32GB |
|
Tesla V100 PCIe |
|
Tesla V100 PCIe 32GB |
|
Tesla V100S PCIe 32GB |
|
Tesla V100 FHHL |
|
Peer-To-Peer (P2P) CUDA Transfers#
Peer-to-Peer (P2P) CUDA transfers enable device memory between vGPUs on different GPUs that are assigned to the same VM to be accessed from within CUDA kernels. NVLink is a high-bandwidth interconnect that enables fast communication between such vGPUs.
P2P CUDA transfers over NVLink are supported only on a subset of vGPUs, hypervisor releases, and guest OS releases.
Peer-to-Peer CUDA Transfers Known Issues and Limitations#
Only time-sliced vGPUs are supported. MIG-backed vGPUs are not supported.
P2P transfers over PCIe are not supported.
Warning
On NVIDIA A100 (Ampere) GPUs, P2P over NVLink and NVSwitch is disabled
when Unified Virtual Memory (UVM) is enabled on a multi-vGPU VM in
SR-IOV Heavy mode. Inside the guest, nvidia-smi topo -m reports
PIX (PCIe) instead of NV12 (NVLink). To restore NVLink P2P
between vGPUs, disable UVM on the vGPU devices (for example,
pciPassthru<n>.cfg.enable_uvm = 0 on VMware vSphere). This
restriction is specific to the NVIDIA Ampere A100 architecture and
does not apply to NVIDIA Hopper (H100/H200) or newer architectures.
vGPU Support for P2P#
Only NVIDIA vGPU for Compute time-sliced vGPUs allocated all of the physical GPU framebuffer on physical GPUs supporting NVLink are supported.
Board |
vGPU |
|---|---|
NVIDIA HGX B200 180GB |
NVIDIA B200X-180C |
Board |
vGPU |
|---|---|
NVIDIA H800 PCIe 94GB (H800 NVL) |
H800L-94C |
NVIDIA H800 PCIe 80GB |
H800-80C |
NVIDIA H200 PCIe 141GB (H200 NVL) |
H200-141C |
NVIDIA H200 SXM5 141GB |
H200X-141C |
NVIDIA H100 PCIe 94GB (H100 NVL) |
H100L-94C |
NVIDIA H100 SXM5 94GB |
H100XL-94C |
NVIDIA H100 PCIe 80GB |
H100-80C |
NVIDIA H100 SXM5 80GB |
H100XM-80C |
NVIDIA H100 SXM5 64GB |
H100XS-64C |
NVIDIA H20 SXM5 141GB |
H20X-141C |
NVIDIA H20 SXM5 96GB |
H20-96C |
Board |
vGPU |
|---|---|
|
A800D-80C |
NVIDIA A800 PCIe 40GB active-cooled |
A800-40C |
NVIDIA A800 HGX 80GB |
A800DX-80C [2] |
|
A100D-80C |
NVIDIA A100 HGX 80GB |
A100DX-80C [2] |
NVIDIA A100 PCIe 40GB |
A100-40C |
NVIDIA A100 HGX 40GB |
A100X-40C [2] |
NVIDIA A40 |
A40-48C |
|
A30-24C |
NVIDIA A16 |
A16-16C |
NVIDIA A10 |
A10-24C |
NVIDIA RTX A6000 |
A6000-48C |
NVIDIA RTX A5500 |
A5500-24C |
NVIDIA RTX A5000 |
A5000-24C |
Board |
vGPU |
|---|---|
Quadro RTX 8000 passive |
RTX8000P-48C |
Quadro RTX 6000 passive |
RTX6000P-24C |
Board |
vGPU |
|---|---|
Tesla V100 SXM2 |
V100X-16C |
Tesla V100 SXM2 32GB |
V100DX-32C |
Hypervisor Platform Support for Multi-vGPU and P2P#
Hypervisor Platform |
NVIDIA AI Enterprise Infra Release |
Supported vGPU Types |
Documentation |
|---|---|---|---|
Red Hat Enterprise Linux with KVM |
All active NVIDIA AI Enterprise Infra Releases |
All NVIDIA vGPU for Compute with PCIe GPUs; on supported GPUs, both time-sliced and MIG-backed vGPUs are supported. |
|
Ubuntu with KVM |
All active NVIDIA AI Enterprise Infra Releases |
All NVIDIA vGPU for Compute with PCIe GPUs; on supported GPUs, both time-sliced and MIG-backed vGPUs are supported. |
|
VMware vSphere |
All active NVIDIA AI Enterprise Infra Releases |
Time-sliced multi-vGPU on supported GPUs. MIG-backed multi-vGPU requires VMware Cloud Foundation (VCF) 9.1 or later (NVIDIA AI Enterprise Infra 7.8 and later). |
Note
P2P CUDA transfers are not supported on Windows. Only Linux OS distros as outlined in NVIDIA AI Enterprise Infrastructure Support Matrix are supported.
Note
MIG-backed multi-vGPU on VMware vSphere requires VMware VCF 9.1 or later. It is not supported on VCF 9.0P01 or on standalone ESXi without VCF 9.1. Linux KVM hypervisors continue to support MIG-backed multi-vGPU on supported GPUs.
Footnotes