Skip to main content
Ctrl+K
NVIDIA Mission Control Software Administration Guide - Home NVIDIA Mission Control Software Administration Guide - Home

NVIDIA Mission Control Software Administration Guide

NVIDIA Mission Control Software Administration Guide - Home NVIDIA Mission Control Software Administration Guide - Home

NVIDIA Mission Control Software Administration Guide

Table of Contents

NVIDIA Mission Control

  • Overview
  • Mission Control Software Stack
  • Node and Category Management
  • Slurm Workload Management
  • NVIDIA Run:ai Installation
  • Adding and Removing Nodes from Run:ai or Slurm
  • Observability Software
  • Connecting to NVIDIA Mission Control autonomous hardware recovery
  • Out-of-Band Management
  • NVLink Partition Management
  • NVLink Management Software (NMX + NetQ)
  • Leak Detection
  • Backups
  • Autonomous Job Recovery
    • Introduction
    • Accessing Clusters
    • Accessing Dashboards
    • Monitoring and Logs
    • Grafana Cloud Setup
    • Viewing Job Details
    • Accessing the Cockpit
    • AJR Job Monitoring
    • Confirming AJR is Operational
    • How-to: Toggle Dry-Run Mode
    • Debugging Common Issues
  • Power Reservation Steering
    • Introduction
    • Concepts and Components
    • Installation
    • Advanced Configuration
    • Metrics
    • Troubleshooting
    • FAQ
  • Workload Power Profile Solution (WPPS)
    • Introduction
    • Components and Concepts
    • Installation
    • First Slurm Job with WPPS
    • Frequently Asked Questions
  • Power Reservation Steering

Power Reservation Steering#

  • Introduction
    • Motivation
    • What is PRS?
  • Concepts and Components
    • Concepts
    • Components
  • Installation
  • Advanced Configuration
    • Understanding Roles and Overlays
    • Configuring PRS Server
    • Configuring PRS Client
  • Metrics
    • Device-Level Metrics
    • Node-Level Metrics
    • PD-Level Metrics
    • Job-Level Metrics
    • Metrics with cmsh
    • Metrics with BaseView (Web UI)
  • Troubleshooting
    • Jobs stuck in the queue pending
  • FAQ
    • Static vs. dynamic power
    • Who should configure the PDN?
    • Why can’t I submit one large job across all nodes?
    • Example 1: Configuring a PD for a GB200/GB300 NVL72 rack
    • Example 2: Power budget adjustment and verification for two GB200 NVL72 racks
    • Will PRS PDs update when a node is removed from BCM or a category?
    • Will PRS PDs update when a new node is added to a category?
    • What is the lifecycle of updates to the PRS config server?
    • Does PRS reflect autoscaler changes automatically?
    • When is the CPU included in PRS-managed devices?
    • What does it mean to stop or start a PD?

previous

Debugging Common Issues

next

Introduction

NVIDIA NVIDIA

Copyright © 2024-2026, NVIDIA Corporation.

Last updated on Aug 14, 2026.