Choosing a vGPU Mode for Inference#

Use this page to pick among Time-Sliced vGPU, MIG-Backed vGPU, and Time-Sliced MIG-Backed vGPU for inference and related multi-tenant workloads. Mode definitions and GPU coverage live on Overview. Profile catalogs live under Types Reference.

Table 69 Inference mode decision guide#

If you need…

Prefer

Avoid / caution

Maximum density on a GPU without MIG

Time-sliced vGPU

Strict per-VM latency SLAs on a busy GPU

Hard isolation and predictable bandwidth between tenants

MIG-backed vGPU

Platforms without MIG-backed vGPU support

Hard boundaries between tenant groups and more than one light workload per MIG slice

Time-sliced MIG-backed vGPU (where supported)

Mixing with heterogeneous vGPU scheduler modes that disallow Fixed Share

Mixed framebuffer sizes on one GPU

Heterogeneous vGPU (with Best Effort or Equal Share)

Fixed Share scheduling with heterogeneous profiles

Hopper / Blackwell profile sizing for inference vs training

Architecture reference pages (for example Hopper Architecture vGPU Types, Blackwell Architecture vGPU Types)

Assuming every profile fits every mode

Tip

After you choose a mode, confirm hypervisor and guest OS coverage in the Support Matrix, then set scheduling policy on Scheduling Policies for time-sliced guests.

See also