Certify a Cluster
Before you begin
- Install nvcrectl and set up the controller
- Confirm your kubeconfig points at the target cluster
AWS
GB200 (EFA interconnect)
The controller auto-detects AWS + GB200 and applies EFA-specific resources (hugepages-2Mi, vpc.amazonaws.com/efa: 4, EFA hostPath volume) automatically.
GB300 (RoCE interconnect)
Same spec as GB200 with nvidia.com/gpu.product: NVIDIA-GB300. The controller detects GB300 and applies RoCE resource claims (roce-channel) instead of EFA — no hugepages, no EFA volumes.
H100
H100 on AWS uses vpc.amazonaws.com/efa: 32. No hugepages or ComputeDomain.
GCP
Content coming soon.
Azure
Content coming soon.
Monitoring progress
Reviewing results
- Passed — all categories met their thresholds. Cluster is ready.
- Failed — one or more categories failed. See Interpret Results for how to read the failed node list and act on it.