Overview#

When something fails in an NVIDIA AI Enterprise Infrastructure deployment, start here to find the fix quickly.

If a single product is failing (for example, GPU Operator pods, Container Toolkit runtime errors, vGPU install or licensing reachability failures, or Network Operator issues), use the component resource index to open that product’s troubleshooting or known-issues documentation.

Component Resource Index#

Use this table when a failure is isolated to one Infrastructure component. Open Troubleshooting for how to diagnose and fix the issue, and Known issues for documented product limitations.

Table 156 Infrastructure component troubleshooting resources#

Component

Troubleshooting

Known issues

NVIDIA Data Center GPU Driver

Data Center drivers documentation

Data Center drivers documentation (release notes for your version)

NVIDIA DOCA Driver for Networking

BlueField software troubleshooting

DOCA documentation

NVIDIA Fabric Manager

Fabric Manager user guide

Data Center drivers documentation (Fabric Manager ships with the driver set)

NVIDIA DOCA Microservices

BlueField software troubleshooting

DOCA services

NVIDIA vGPU for Compute

vGPU for Compute user guide License server reachability

Limitations (per-hypervisor known product limitations)

NVIDIA Container Toolkit

Container Toolkit troubleshooting

Container Toolkit release notes

NVIDIA Run:ai

NVIDIA Run:ai documentation

NVIDIA Run:ai documentation

NVIDIA DPU Operator (DPF)

DPF documentation (open Operational Readiness > Troubleshooting)

DPF documentation (open Release notes for known issues)

NVIDIA GPU Operator

GPU Operator troubleshooting

GPU Operator release notes

NVIDIA Network Operator

Network Operator common issues

Networking documentation

NVIDIA NIM Operator

NIM Operator documentation

NIM Operator release notes

NVIDIA Base Command Manager (BCM)

Base Command Manager documentation

BCM 11.x release notes, BCM 10.x release notes