Overview of the Observability Guide for NVIDIA Enterprise Reference Architectures (Enterprise RAs)#

The Observability Guide for NVIDIA Enterprise RAs is designed to provide customers with a methodology to monitor the health, performance and resource utilization of NVIDIA’s AI infrastructure starting from Day 1.

As AI workloads scale and grow in complexity, observability becomes a critical requirement to ensure system reliability, and resource efficiency. This Observability Guide complements the existing hardware-focused NVIDIA Enterprise RAs, and is intended to accelerate time-to-value for customers and partners by reducing the operational setup associated with observability tools from ecosystem software partners.

Along with this guide, there are Grafana JSON dashboard and alerting templates that can be accessed through GitHub.

This Observability Guide is intended to be used as a design reference and practical step-by-step guide for developers building an observability stack on Enterprise RAs. While this guide offers a reference, it is not exhaustively prescriptive.

Organizations can substitute components for functional equivalents and maintain the benefits of enabling observability across the stack.

For example, various Kubernetes distribution options (e.g., Red Hat OpenShift or Canonical Kubernetes) are compatible with the components described[1] and recommendations offered in this guide. This guidance is also supplemented by an ecosystem of validated observability software partners – described in the NVIDIA AI factory validated and reference designs for enterprise and government. NVIDIA global system partners and channel partners can apply the observability approach described below to their customers’ enterprise stack of choice.

_images/observability-stack-overview.png

Figure 1 Observability stack overview for NVIDIA Enterprise Reference Architectures#