Choosing a vGPU Mode for Inference#
Use this page to pick among Time-Sliced vGPU, MIG-Backed vGPU, and Time-Sliced MIG-Backed vGPU for inference and related multi-tenant workloads. Mode definitions and GPU coverage live on Overview. Profile catalogs live under Types Reference.
If you need… |
Prefer |
Avoid / caution |
|---|---|---|
Maximum density on a GPU without MIG |
Time-sliced vGPU |
Strict per-VM latency SLAs on a busy GPU |
Hard isolation and predictable bandwidth between tenants |
MIG-backed vGPU |
Platforms without MIG-backed vGPU support |
Hard boundaries between tenant groups and more than one light workload per MIG slice |
Time-sliced MIG-backed vGPU (where supported) |
Mixing with heterogeneous vGPU scheduler modes that disallow Fixed Share |
Mixed framebuffer sizes on one GPU |
Heterogeneous vGPU (with Best Effort or Equal Share) |
Fixed Share scheduling with heterogeneous profiles |
Hopper / Blackwell profile sizing for inference vs training |
Architecture reference pages (for example Hopper Architecture vGPU Types, Blackwell Architecture vGPU Types) |
Assuming every profile fits every mode |
Tip
After you choose a mode, confirm hypervisor and guest OS coverage in the Support Matrix, then set scheduling policy on Scheduling Policies for time-sliced guests.
See also
Overview - mode table and platform limitations
MIG-Backed vGPU - MIG-backed behavior
Scheduling Policies - Best Effort / Equal Share / Fixed Share
NVIDIA vGPU Types by Hardware - architecture profile catalogs