Scaling NICo with machine-a-tron: 100 → 1000 → 4500 simulated hosts
Scaling NICo with machine-a-tron: 100 → 1000 → 4500 simulated hosts
Status: Scale testing validated with up to 13,500 BMC endpoints.
Quick start
Reset the site between runs, check the BMC network against the ServiceCIDR, then install a scale profile. Refer to Teardown and Reset for the reset. Cluster Prerequisites covers the namespace label and pull Secret the install expects. Controller Mode lists the Core values a simulation site needs.
Replicating the 250-Rack Fleet
The quick start above installs host-count fleets. The 250-rack measurement in Large Site Sizing and Settings ran ten machine-a-tron pods of 25 GB200 NVL72 racks each in Controller Mode behind the protocol gateway. To replicate it on a fresh site:
- In
helm-prereqs/values.yaml, setsiteCredentials.enabled: truewith an explicitbmcRoot.password(Site Credentials Secret), and raisepostgresql.resources.limitsto the Postgres limits in Settings Changed From the Defaults beforesetup.shruns. Teardown deletes thepostgresnamespace, so a resize applied to the live cluster is lost on the next install. - Copy
helm-prereqs/values/nico-core-simulation.yaml, fill in the site blanks, and pointnico-api.credentials.file.existingSecret.nameat the Secret from step 1. The copy already carries the[site_explorer]budget,max_concurrency, and the nico-api CPU limit from the sizing page. It also setsmax_database_connections = 900and enables the hardware-health rate limiter, which keep the connection pool and the Redfish session rate within limits at this scale. - In the same copy, widen
[networks.simulated-oob]to a/17before nico-api first starts, as the sizing page’s fleet paragraph explains, and pointnico-api.rms.apiUrlat the protocol gateway (Pointing NICo at It). - Deploy the DOCA Platform Framework (DPF) simulator, then run
./setup.sh --skip-dpf --core-values <copy>fromhelm-prereqs/. The overlay header and the dpf-sim-controller quick start give themake deploycommand and the namespace it creates. - Check the rack profile. nico-api derives an expected rack’s profile id from
the expected rack group that declares it: the group topology in upper case,
an underscore, the compute vendor in upper case, and the suffix
_NO_POWERSHELFwhen the group declares no power shelf. It replaces the requested id with the derived one and rejects the rack unless that id exists under[rack_profiles]. With PR 6995, machine-a-tron declares one group per rack from its rack type, sorack_profile_idin the machine-a-tron values must equal the derived id (GB200_NVL72R1_C2G4_WIWYNNforwiwynn_gb200_nvl72,GB300_NVL72R1_C2G4_LENOVOforlenovo_gb300_nvl72), which the nico-api chart ships; that PR also updates the deployment page and the 10-rack values. On a site installed before it, create the groups withnico-admin-cli expected-rack-groupfirst. Inspect the declarations withnico-admin-cli expected-rack-group showandnico-admin-cli expected-rack show. - Create the labeled
nico-matnamespace and pull Secret (Cluster Prerequisites), runhelm-prereqs/check-mat-service-cidr.pyagainst the machine-a-tron values with--site-configpointing at the copy from step 2, then install machine-a-tron fromhelm-prereqs/values/machine-a-tron-10racks.yamlwith each pod’sidslist extended to 25 racks, itsrack_profile_idset to the id from step 5, and itsresourcesandstartupProbescaled with it. Deploying a 250-Rack Site gives the per-pod sizing and the Service count to expect. - Leave
max_concurrencyat the overlay’s 80 or raise it to 120. Time to Ready Is a Concurrency Setting gives the measured range and how to set it, and the settings table lists the[api_admission_control]value used for the simulated fleet.
What this work delivers
- Controller Mode in the
nico-machine-a-tronchart, with themat-k8s-controllerfor dynamic per-BMC Services, and the scale profiles underhelm-prereqs/values/:machine-a-tron-scale.yaml,machine-a-tron-multipod.yaml,machine-a-tron-scale-4500.yaml, andmachine-a-tron-scale-4500-proxy.yaml.machine-a-tron-10racks.yamladds the ten-rack GB200 NVL72 example with the Rack Management Service (RMS) mock behind the protocol gateway. helm-prereqs/values/nico-core-simulation.yaml, the NICo Core values for a simulation site:allow_insecure_discovery, the site-explorer throughput knobs, the three simulated networks, and the pool sizes.helm-prereqs/check-mat-service-cidr.py, the BMC network versus ServiceCIDR preflight, andhelm-prereqs/ingestion-rate-report.sh, the per-run ingestion curves from the database timestamps.
Architecture: Controller Mode
The mat-k8s-controller dynamically creates one Service per BMC:
- Discovers machine-a-tron pods via
nvidia-infra-controller/mat-service=truelabel - Polls
/machines/statusfrom each pod - Creates Services with the BMC IP (assigned by NICo DHCP) as
externalIPs - Services route to correct pod via
nvidia-infra-controller/pod-nameselector
Requirements:
- The BMC network must lie outside the Kubernetes ServiceCIDR, pod CIDR,
node network, and networks that nodes or pods must otherwise reach
(BMC IPs are Service externalIPs, for which kube-proxy programs forwarding
rules on every node). Run
helm-prereqs/check-mat-service-cidr.pyagainst the values file before every install. It resolves each BMC relay to its[networks.*]prefix and fails on an overlap or when it cannot determine the ServiceCIDR.SCALE_SERVICE_CIDRSreplaces the cluster lookup, andSCALE_BMC_PREFIXESadds networks the site config does not declare yet - NICo siteConfig needs
allow_insecure_discovery = trueand a network covering the BMC IP range - Leave
site_explorer.bmc_proxyunset - NICo dials each BMC IP directly
NICo siteConfig (helm-prereqs/values/nico-core-simulation.yaml):
The file also declares simulated-admin and simulated-underlay (the
underlayDhcpRelayAddress target) and sizes [pools.lo-ip] and
[pools.fnn-asn] for 4,500 hosts with 2 DPUs.
Complete issue log
Every issue below was found live on a 3-node development cluster. The Fix column names the file or section that now carries each fix.
Baseline (override-mode) end-to-end
Scale mode (100 hosts × 2 DPUs and up)
A note on the verification loop
The retired setup script’s final phase actively shepherded the pipeline. It
re-cleared AvoidLockout and Unauthorized latches (they are one-way by
design) and unparked healthy endpoints. With pinned BMC passwords the latch
clearing was a no-op for the whole stage-3 run, so the chart path carries no
equivalent. An operator clears a latched endpoint with
nico-admin-cli site-explorer refresh <bmc-ip> or
nico-admin-cli site-explorer clear-error <bmc-ip>.
Where we are today
Stage-3 observations worth reviewers’ attention:
- The ingestion pipeline is fully autonomous once configured. Client connectivity to the cluster dropped twice for extended periods during stage 3; ingestion continued unattended both times (e.g. +720 machines through one outage, +4,000 through another). The shepherd loop’s latch clearing, critical in earlier iterations, was a no-op for the entire stage-3 run thanks to pinned credentials.
- Measured stage-3 rates on the 3-node development cluster: DHCP
~110 interfaces/min; exploration ~120 to 360 endpoints/cycle; creation
40 to 240 machines per completed explore cycle, sawtoothing with cycle
phasing (identification rebuilds
explored_managed_hostseach cycle). - Per-MAC Vault credential lifecycle needs batching at scale (issue 18
below): site-explorer stores one
machines/bmc/<mac>/rootentry per BMC, 13,500 entries. Deleting them one API round-trip at a time takes hours, batched server-side it takes seconds. expected_machinesauto-registration worked at stage 3 (9,890+ registered by machine-a-tron via the API), confirming the RBAC grant path.
Additional issue found at stage 3:
Open Questions - Feedback Wanted
- RBAC:
Machineatron→AddExpectedMachineis granted on main (commit9a9ba072a), and machine-a-tron registers expected machines through the API. Resolved. - Seed-once reconcile semantics: networks, and resource-pool
definitions are all create-once; config changes on established sites are
silently ignored (or warn-only). The chart path declares them before first
start and uses
nico-admin-cli network-segment createandresource-pool growon established sites. Should NICo support declarative updates for these instead? - AvoidLockout at scale: one-way latches are right for real BMCs, but
simulation fleets guarantee latch storms during resets. Worth a
site-config escape hatch (e.g.
site_explorer.lockout_protection = false) instead of operator-side clearing? - Mock fidelity: the mock returns to its configured password after a BMC reset. Real BMCs persist a rotated password across resets. Should bmc-mock persist rotated credentials so the rotation path can be exercised at scale without pinning?
- lo-ip per machine: is one loopback IP per machine the intended allocation at 13.5k machines, and is there guidance for sizing this pool in production site templates (dev templates ship 3)?
- Cycle economics: identification/creation only run at the end of a
completed
explore_sitecycle, soexplorations_per_runtrades sweep throughput against creation latency in a non-obvious way. Worth documenting (or decoupling creation from the exploration cycle)?