MaxLPS Power Pilot Runbooks#
Run this pilot to determine whether four GB200 or GB300 NVL72 racks can deliver more aggregate inference throughput inside the same facility power envelope as a three-rack baseline. DPS manages fleet GPU power while your existing workload scheduler and provisioning systems keep ownership of workload placement and fleet membership.
Start Here#
Get the pilot goal, the five steps, the minimum concepts, and links to the right detailed guide. Start here if you are new to DPS or returning to the pilot after working in another guide.
Run the Pilot#
Complete the end-to-end procedure: prepare DPS, establish the 3-rack baseline,
enable the default MaxLPS bundle for the 4-rack managed run, and compare
results.
Supporting Guides#
Capture the GPU and CPU power and utilization evidence needed to compare the baseline and managed runs.
Learn how the topology, resource group, default MaxLPS bundle, and Power
Steering work together to dynamically allocate GPU power within the facility
envelope.
Prepare BMC access, verify Redfish compatibility, run the health gate, and resolve BMC latency or power-control issues before a managed DPS run.
Find targeted recovery steps for deployment verification, topology, BMC health, resource groups, and Power Steering.