vGPU Deployment#
NVIDIA vGPU lets multiple virtual machines (VMs) share a single physical GPU, which makes it a common deployment pattern in enterprise data centers and cloud environments.
NIM LLM and VLM supports deployment on vGPU-enabled VMs, so you can run LLM inference workloads in existing virtualized infrastructure.
vGPU deployment is a good fit in the following situations:
Your organization uses VMware vSphere, Red Hat KVM, Citrix Hypervisor, or other supported hypervisors with NVIDIA vGPU.
You need to share GPU resources across multiple workloads or tenants.
Your IT policy requires VM-based isolation rather than container-only deployment.
You are deploying in an NVIDIA AI Enterprise environment.
Prerequisites#
Before you deploy NIM LLM and VLM on vGPU, complete the following hardware and software prerequisites.
Hardware#
Prepare the following hardware:
Provide an NVIDIA GPU that supports vGPU (refer to the NVIDIA vGPU supported products).
Allocate a VM with sufficient GPU frame buffer for the target model (refer to Memory and frame buffer on vGPU in Troubleshooting).
Software#
Install or configure the following software:
Install and license NVIDIA vGPU software on the hypervisor host.
Install the NVIDIA vGPU guest driver inside the VM.
Install NVIDIA Container Toolkit 1.14+ inside the VM.
Install Docker 24.0+ or another OCI container runtime inside the VM.
Use a Linux guest OS (refer to the NVIDIA vGPU guest OS support matrix).
To deploy NIM LLM and VLM on a vGPU-enabled VM, complete the following steps:
Confirm that the vGPU is visible inside the VM:
nvidia-smi
Expected output shows a virtual GPU with allocated frame buffer:
+-----------------------------------------------------------------------------------------+
| NVIDIA-SMI 560.xx.xx Driver Version: 560.xx.xx CUDA Version: 12.x |
|-------------------------------------------+----------------------+----------------------+
| GPU Name Persistence-M | Bus-Id Disp.A | Volatile Uncorr. ECC |
| Fan Temp Perf Pwr:Usage/Cap | Memory-Usage | GPU-Util Compute M. |
|===========================================+======================+======================|
| 0 NVIDIA A100-SXM4-40GB-... On | 00000000:00:10.0 Off | 0 |
| N/A 35C P0 45W / N/A | 0MiB / 40960MiB | 0% Default |
+-------------------------------------------+----------------------+----------------------+
Confirm that the NVIDIA Container Toolkit is configured:
docker run --rm --gpus=all nvidia/cuda:12.6.3-base-ubuntu24.04 nvidia-smi
Deploy NIM on the VM by following the Quickstart instructions.
Troubleshoot vGPU Deployments#
Use the following guidance when list-model-profiles reports insufficient memory or nvidia-smi does not show a GPU inside the VM.
Memory and Frame Buffer on vGPU#
On vGPU, the usable GPU frame buffer inside the VM is less than the configured vGPU profile size because the vGPU software reserves a portion for its own use. For exact usable frame buffer per profile, refer to the NVIDIA vGPU documentation.
Choosing a vGPU profile — Select a vGPU profile (for example A100-40C, H100-80C) that provides enough frame buffer for the model you intend to serve. As a guideline:
Model Size |
Approximate GPU Memory Required |
Example vGPU Profiles |
|---|---|---|
7B–8B (FP16/BF16) |
~16 GB |
A100-20C, H100-20C or larger |
7B–8B (FP8) |
~10 GB |
A100-10C, H100-10C or larger |
13B (FP16/BF16) |
~28 GB |
A100-40C, H100-40C or larger |
70B (FP8) |
~70 GB |
Full GPU profile (A100-40C x2 or A100-80C) |
These are rough estimates. Actual memory usage depends on the model, quantization, context length, and KV cache configuration. Use list-model-profiles inside the VM to see which profiles are compatible with the available GPU memory.
GPU memory utilization — In vGPU environments, NIM can automatically lower the default GPU memory utilization (--gpu-memory-utilization) to account for the reduced available frame buffer. If the memory estimator detects a shared-memory environment (vGPU or UMA), it sets the value to the greater of 0.5 or the estimated minimum required fraction.
You can override this by passing the backend argument directly to the container:
docker run --gpus=all \
<nim_llm_image> \
--gpu-memory-utilization 0.85
Passing the option as a container argument preserves any
NIM_PASSTHROUGH_ARGS value provided by the image. For other deployment
interfaces, refer to Advanced Configuration.
Note
Setting GPU memory utilization too high in a vGPU environment can cause out-of-memory errors because the vGPU software reserves frame buffer that is not visible to the application. Start with a conservative value and increase gradually.
Profile Selection Reports Insufficient Memory#
If list-model-profiles shows all profiles as “Incompatible with system,” the vGPU profile does not have enough frame buffer for any available model profile.
Use the following resolutions:
Increase the vGPU profile size (allocate more frame buffer to the VM) at the hypervisor level.
Use a smaller model or a more aggressively quantized profile (FP8, MXFP4).
Decrease the maximum model length to lower the KV cache memory requirement.
vGPU Not Detected#
If nvidia-smi does not show a GPU inside the VM:
Verify that the NVIDIA vGPU software is installed and licensed on the hypervisor host.
Verify that a vGPU profile is assigned to the VM in the hypervisor management console.
Verify that the NVIDIA vGPU guest driver is installed inside the VM.
Check the hypervisor logs for vGPU allocation errors.
Refer to the NVIDIA vGPU troubleshooting guide for detailed diagnostics.
Next Steps#
Refer to the following resources for more information: