Skip to main content
Ctrl+K
Observability Guide - Home Observability Guide - Home

Observability Guide

Observability Guide - Home Observability Guide - Home

Observability Guide

Table of Contents

NVIDIA Enterprise Reference Architecture

  • Abstract
  • Overview of the Observability Guide for NVIDIA Enterprise Reference Architectures (Enterprise RAs)
  • Technical Prerequisites for Observability
  • Configuration of Data Sources
  • Custom Dashboards
  • Alerting
  • Summary
  • Out of Scope
  • Appendix A

Notices

  • NVIDIA Enterprise Reference Architecture Notice & Disclaimers

Observability Guide#

NVIDIA Enterprise Reference Architecture

  • Abstract
  • Overview of the Observability Guide for NVIDIA Enterprise Reference Architectures (Enterprise RAs)
  • Technical Prerequisites for Observability
    • NVIDIA Enterprise RAs – Prerequisites
    • Verifying Kube Prometheus and Grafana Setup
    • Configuring Persistent Storage for Grafana
    • Steps to configure a PVC for Grafana
    • Network Monitoring with NVIDIA NetQ
  • Configuration of Data Sources
    • Base Command Manager – Prometheus
    • DCGM Exporter
    • NIM Metrics – Kube Prometheus
    • NetQ Prometheus
  • Custom Dashboards
    • GPU Dashboard
    • NIMs Dashboard
    • Kubernetes Global Dashboard
    • CPU & Node Storage Dashboard
    • Network Dashboard
  • Alerting
    • Configuring Grafana for Alerting
      • Folder Creation
      • Contact Point Creation
      • Setting up SMTP Server for configuration
    • Automated Alert Uploads
      • Script Overview
    • Alert Rule JSONs
  • Summary
  • Out of Scope
  • Appendix A

Notices

  • NVIDIA Enterprise Reference Architecture Notice & Disclaimers

next

Abstract

NVIDIA NVIDIA
Privacy Policy | Manage My Privacy | Do Not Sell or Share My Data | Terms of Service | Accessibility | Corporate Policies | Product Security | Contact

Copyright © 2025-2026, NVIDIA Corporation.

Last updated on Jul 14, 2026.