Observability Guide# NVIDIA Enterprise Reference Architecture Abstract Overview of the Observability Guide for NVIDIA Enterprise Reference Architectures (Enterprise RAs) Technical Prerequisites for Observability NVIDIA Enterprise RAs – Prerequisites Verifying Kube Prometheus and Grafana Setup Configuring Persistent Storage for Grafana Steps to configure a PVC for Grafana Network Monitoring with NVIDIA NetQ Configuration of Data Sources Base Command Manager – Prometheus DCGM Exporter NIM Metrics – Kube Prometheus NetQ Prometheus Custom Dashboards GPU Dashboard NIMs Dashboard Kubernetes Global Dashboard CPU & Node Storage Dashboard Network Dashboard Alerting Configuring Grafana for Alerting Folder Creation Contact Point Creation Setting up SMTP Server for configuration Automated Alert Uploads Script Overview Alert Rule JSONs Summary Out of Scope Appendix A Notices NVIDIA Enterprise Reference Architecture Notice & Disclaimers