Skip to main content
Ctrl+K
NVIDIA Mission Control Software Administration Guide - Home NVIDIA Mission Control Software Administration Guide - Home

NVIDIA Mission Control Software Administration Guide

NVIDIA Mission Control Software Administration Guide - Home NVIDIA Mission Control Software Administration Guide - Home

NVIDIA Mission Control Software Administration Guide

Table of Contents

NVIDIA Mission Control

  • Overview
  • Mission Control Software Stack
  • Node and Category Management
  • Slurm Workload Management
  • NVIDIA Run:ai Installation
  • Adding and Removing Nodes from Run:ai or Slurm
  • Observability Software
  • Connecting to NVIDIA Mission Control autonomous hardware recovery
  • Out-of-Band Management
  • NVLink Partition Management
  • NVLink Management Software (NMX + NetQ)
  • Leak Detection
  • Backups
  • Autonomous Job Recovery
    • Introduction
    • Accessing Clusters
    • Accessing Dashboards
    • Monitoring and Logs
    • Grafana Cloud Setup
    • Viewing Job Details
    • Accessing the Cockpit
    • AJR Job Monitoring
    • Confirming AJR is Operational
    • How-to: Toggle Dry-Run Mode
    • Debugging Common Issues
  • Power Reservation Steering
    • Introduction
    • Concepts and Components
    • Installation
    • Advanced Configuration
    • Metrics
    • Troubleshooting
    • FAQ
  • Workload Power Profile Solution (WPPS)
    • Introduction
    • Components and Concepts
    • Installation
    • First Slurm Job with WPPS
    • Frequently Asked Questions
  • Autonomous Job Recovery

Autonomous Job Recovery#

  • Introduction
  • Accessing Clusters
  • Accessing Dashboards
  • Monitoring and Logs
  • Grafana Cloud Setup
  • Viewing Job Details
  • Accessing the Cockpit
    • Monitoring and Performance
  • AJR Job Monitoring
    • Application Log Format Requirements
  • Confirming AJR is Operational
    • Singleton Dependency
    • Requeue
  • How-to: Toggle Dry-Run Mode
  • Debugging Common Issues
    • Issue: Job is not managed by AJR or is not visible on the Cockpit
    • Issue: Job is available in Cockpit but does not receive any anomaly and checkpoint events
    • Issue: The job died but was not requeued
    • Issue: The requeued job or the next dependency job was not released

previous

Backups

next

Introduction

NVIDIA NVIDIA

Copyright © 2024-2026, NVIDIA Corporation.

Last updated on Aug 14, 2026.