Features#

NVIDIA vGPU (Virtual GPU) for Compute virtualizes NVIDIA GPUs for AI, machine learning, and high-performance computing. The subsections below describe MIG (Multi-Instance GPU) partitioning, provisioning, data paths, migration, multi-GPU guests, interconnects, scheduling, power state, and unified memory.

Table 45 Feature capability summary#

Feature

What it covers

Page

MIG-Backed vGPU

Hardware-level GPU partitioning with spatial isolation

MIG-Backed vGPU

Device Groups

Topology-aware detection and provisioning of connected devices

Device Groups

GPUDirect

RDMA and storage paths that reduce CPU overhead

GPUDirect RDMA and GPUDirect Storage

Heterogeneous vGPU

Mixed vGPU profiles on one physical GPU

Heterogeneous vGPU

Live Migration

VM migration with short stun time on supported hypervisors

Live Migration

Multi-vGPU and P2P

Multi-GPU guests, board support, and NVLink P2P (hub)

Multi-vGPU and P2P

NVIDIA NVSwitch

High-bandwidth NVLink fabric between GPUs (includes multicast)

NVIDIA NVSwitch

NVLink Multicast

One-to-many data distribution (requires UVM; see NVSwitch page)

NVLink Multicast

Scheduling Policies

Best Effort, Equal Share, and Fixed Share time-slicing

Scheduling Policies

Suspend-Resume

VM state preservation for resource management

Suspend-Resume

Unified Virtual Memory

Single address space across CPU and GPU (includes board support)

Unified Virtual Memory (UVM)

Inference mode decision

Time-sliced vs MIG-backed vs time-sliced MIG-backed for inference

Choosing a vGPU Mode for Inference