Abstract#
This Observability Guide for NVIDIA Enterprise Reference Architectures (Enterprise RAs) offers a standardized and production-ready reference as well as step-by-step guidance for implementing observability in enterprise AI or HPC environments.
Built on top of NVIDIA’s AI infrastructure and Kubernetes-native platforms, this version of the guide specifically focuses on advanced custom dashboard and alerting solutions for AI factories, providing administrators and enterprise customers with actionable insights into GPU, CPU, Kubernetes, and applications. This includes guidance on setting up data pipelines and custom dashboards using tools such as NVIDIA Data Center GPU Manager (DCGM), NIM Operator, Prometheus Node Exporter, and other telemetry sources. Dashboards and alerts will ensure improved cluster health visibility as it experiences AI workloads and accelerated troubleshooting.
NVIDIA global system partners, channel partners, and enterprise customers are invited to apply the guidance described in this document to their software stack of choice, and to substitute components (e.g. from NVIDIA ecosystem software partners offering Kubernetes distributions and/or observability and monitoring tools) where needed while retaining the benefits of the observability approach.
Future versions of this guide will extend observability capabilities to include logging and distributed tracing, enabling a complete end-to-end observability stack.