Overview#
When something fails in an NVIDIA AI Enterprise Infrastructure deployment, start here to find the fix quickly.
If a single product is failing (for example, GPU Operator pods, Container Toolkit runtime errors, vGPU install or licensing reachability failures, or Network Operator issues), use the component resource index to open that product’s troubleshooting or known-issues documentation.
Component Resource Index#
Use this table when a failure is isolated to one Infrastructure component. Open Troubleshooting for how to diagnose and fix the issue, and Known issues for documented product limitations.
Component |
Troubleshooting |
Known issues |
|---|---|---|
NVIDIA Data Center GPU Driver |
Data Center drivers documentation (release notes for your version) |
|
NVIDIA DOCA Driver for Networking |
||
NVIDIA Fabric Manager |
Data Center drivers documentation (Fabric Manager ships with the driver set) |
|
NVIDIA DOCA Microservices |
||
NVIDIA vGPU for Compute |
Limitations (per-hypervisor known product limitations) |
|
NVIDIA Container Toolkit |
||
NVIDIA Run:ai |
||
NVIDIA DPU Operator (DPF) |
DPF documentation (open Operational Readiness > Troubleshooting) |
DPF documentation (open Release notes for known issues) |
NVIDIA GPU Operator |
||
NVIDIA Network Operator |
||
NVIDIA NIM Operator |
||
NVIDIA Base Command Manager (BCM) |