MaxLPS Power Pilot Runbooks#

Run this pilot to determine whether four GB200 or GB300 NVL72 racks can deliver more aggregate inference throughput inside the same facility power envelope as a three-rack baseline. DPS manages fleet GPU power while your existing workload scheduler and provisioning systems keep ownership of workload placement and fleet membership.

Start Here#

Pilot at a glance

Get the pilot goal, the five steps, the minimum concepts, and links to the right detailed guide. Start here if you are new to DPS or returning to the pilot after working in another guide.

MaxLPS Power Pilot

Run the Pilot#

Run the MaxLPS power pilot

Complete the end-to-end procedure: prepare DPS, establish the 3-rack baseline, enable the default MaxLPS bundle for the 4-rack managed run, and compare results.

Run the MaxLPS Power Pilot

Supporting Guides#

Collect pilot telemetry

Capture the GPU and CPU power and utilization evidence needed to compare the baseline and managed runs.

Collect Pilot Telemetry: 1-Second GPU and CPU Data with DCGM Exporter
How MaxLPS works

Learn how the topology, resource group, default MaxLPS bundle, and Power Steering work together to dynamically allocate GPU power within the facility envelope.

MaxLPS Power Management
BMC readiness and health

Prepare BMC access, verify Redfish compatibility, run the health gate, and resolve BMC latency or power-control issues before a managed DPS run.

BMC Readiness and Health Guide
Troubleshoot the pilot

Find targeted recovery steps for deployment verification, topology, BMC health, resource groups, and Power Steering.

Troubleshoot the MaxLPS power pilot