What is it
Nodewright and nodewright-customizations are two halves of the integration. Nodewright is a Kubernetes Operator that applies nodewright packages with consistent, repeatable, and tested lifecycles within a cluster. Nodewright-customizations are instances of the NodeWright Custom Resource (nodewright.nvidia.com/v1alpha1) that define one or more nodewright packages to deploy. Operator v0.18.0 renamed the kind from Skyhook, and AICR’s manifests now declare NodeWright; the registry pins v0.19.0. On a cluster that predates the rename the operator mirrors each legacy Skyhook into a NodeWright of the same name and writes status only on the NodeWright, so applying these manifests adopts the mirrored object rather than creating a second one. The legacy kind is read-only from v0.18.0 on: its admission webhook rejects any spec, pause or disable change, which is why the manifests moved. See the upstream migration guide.
The registry default namespace is nodewright. A deployment made before that moved still runs in skyhook, and Helm cannot relocate a release, so pass aicr recipe --inherit-from <prior recipe or bundle> to keep it where it is; the health-check assertions follow the inherited namespace. These packages were selected to provide two main functions:
- Optimize a node for inference or training workloads via grub, sysctl and systemd service settings.
- Be able to install all of the necessary software to bring a vanilla Kubernetes node to the AICR spec.
References
Optimizer
Uses tuned to apply a sequence of profiles to optimize primarily grub and sysctl settings. Your mileage may vary depending on the particulars of the virtualization if not running baremetal.
Package: nvidia-tuned
Configuration documentation:
A full configuration supplies: intent, accelerator, service. A minimal configuration is just accelerator
Supported accelerators: see the generated table below.
Integration notes:
- If you provide a service it MUST exist in the profiles service directory
- If you are integrating a new service beware that even tested paths may not fully work due to limitations in that service. For example you will notice that
ekshas overrides to remove settingkernel.sched_latency_nsandkernel.sched_min_granularity_nsas these are not available on AWS kernels. They cannot fail silently as the package will test to make sure the changes asked for actually happens and error if it does not.
Secondary optimizer
A second, more stripped down, optimizer is available for operating systems that are mostly read only such as GKE’s ContainerOptimizedOS. In this case the nvidia-tuning-gke is available to directly perform sysctl writes. Also note the change in Nodewright configuration to write to a different directory tree in order to have a writable FS and to re-apply changes every boot: recipes/overlays/gke-cos.yaml
Versioning and extension notes
Both of these packages (nvidia-tuned and nvidia-tuning-gke) extend other nodewright packages (tuned and tuning) and as such could directly use those and provide the configuration via configmaps. The choice was made to go with specific versioned packages in order to provide a more clear path for upgrades and understanding differences. However, the base packages are still useful to quickly iterate on configurations without requiring new versions of the extended packages used in AICR.
Setup
Uses a set of bash scripts to do the necessary actions to bring an ubuntu worker to the desired AICR spec.
Package: nvidia-setup
See the Tuning status table below for the current service + accelerator coverage. Each service must be added explicitly; the documentation for how to make this update is in the nvidia-setup README.
The version overview has all of the information about what each version for a service + accelerator pair will install or configure.
Manifests
Tuning and Setup
Tuning are typically alterations to sysctl, kernel boot parameters and service drop ins to make the system better optimized for AI workloads.
Setup generally is any configuration needed to make AI workloads work in that service that is not already provided by that service. In EKS for example this is kernel and EFA installation. For BCM it is symlinks to support gpu operator.
The table below is generated from the recipes by make tuning-docs — do not edit it by hand. Setup and Tuning are the pinned nodewright package versions applied for each service + accelerator. Profile is the tuning profile accelerator when it differs from the selected accelerator (for example, h200 and a100 nodes use the h100 profile), and - when identical. In the Setup and Tuning columns, - means no such package is applied for that service + accelerator. * means the recipe pins no value for that dimension.
Note: the generated table lists the packages pinned in the manifests and
cannot see per-recipe value gates. The aks rows show nvidia-tuned 0.10.0,
but AKS recipes disable it by default via nodewright-customizations
tuningEnabled: false under the Azure-managed driver profile; see
AKS GPU setup
for the rationale and re-enable path.
The tuningEnabled gate (default true; only an explicit false disables)
applies uniformly across the tuning manifests: on the shared tuning.yaml it
omits the nvidia-tuned package while nvidia-setup keeps running, and on the
single-package manifests (tuning-gke.yaml, tuning-generic.yaml,
tuning-rke2.yaml, tuning-gb300.yaml) it suppresses the whole tuning CR,
since the tuning package is that CR’s only content. No recipe sets it outside
AKS today, so default renderings are unchanged elsewhere.
The tuning CR is a normal release-managed resource (no Helm hooks), so
flipping tuningEnabled from true to false retracts it on all deployers —
under Helm/Argo (which already stripped the hooks at bundle time) and under
Flux (where a hook resource would otherwise have been orphaned on the flip,
since before-hook-creation never fires when the CR leaves the render).
Ordering after the Skyhook CRD install is handled by component dependency
ordering (nodewright-customizations depends on nodewright-operator), not an
intra-release hook.
One-time Flux migration: a cluster that was already running a pre-hook-removal
bundle created the tuning Skyhook as a Helm hook (outside the release
lifecycle). The first upgrade to a bundle carrying this change must keep
tuningEnabled=true so Helm adopts the now-managed CR into the release; only
then does a later upgrade with tuningEnabled=false prune it. If the first
upgrade both removes the hooks and sets tuningEnabled=false, nothing
renders for Helm to adopt and the legacy hook-created Skyhook is left behind —
delete it explicitly in that case (kubectl delete skyhook tuning -n skyhook).
Helm/Argo clusters are unaffected (they never created it as a hook).
Readiness: the deployment validator’s Go check requires every CR the recipe
renders to be complete by name, reading NodeWright and falling back to the
legacy Skyhook on operators older than v0.18.0, and requires every node to
have shed the operator’s runtime-required taint. The taint is read from the
operator Deployment’s RUNTIME_REQUIRED_TAINT env, so a --workload-gate
value passed at bundle time is what the gate waits on, together with the legacy
skyhook.nvidia.com=runtime-required:NoSchedule the operator still removes
during its deprecation window. That check derives its expected names from the
effective values, so a CR the values suppress is never waited on.
Value-gated readiness: the chainsaw health check asserts a NodeWright CR
reaches status.status: complete and cannot read effective values itself. The
deployment validator renders those values and suppresses the assert only when
they produce no CR at all — tuningEnabled: false on a
tuning-gke.yaml/tuning-generic.yaml/tuning-rke2.yaml/tuning-gb300.yaml
recipe, or enabled: false anywhere — so a deliberately untuned cluster passes
rather than failing on an intentionally absent CR
(#1844). When a CR does render,
completion is still required. The suppression is fail-closed: a render, read or
discovery error propagates instead of being read as “nothing to assert”. (AKS
is unaffected by tuningEnabled: false: its tuning CR still renders with the
nvidia-setup packages.)
Workload-gate taint key: operator v0.18.0 changed the chart default
runtimeRequiredTaint from skyhook.nvidia.com=runtime-required:NoSchedule
to nodewright.nvidia.com=runtime-required:NoSchedule. The operator recognizes
both until v0.20.0 (tolerates, removes on completion, never double-taints), but
auto-taints newly joined nodes with the configured key only. Node pools that
pre-taint with the legacy key should pass
--workload-gate skyhook.nvidia.com=runtime-required:NoSchedule to
aicr bundle so the cluster carries a single key.
See recipes/components/nodewright-customizations/manifests for the specifics on packages and their configuration.
Rollout pacing
Every reboot-carrying tuning manifest (tuning.yaml, tuning-gb300.yaml,
tuning-generic.yaml, tuning-rke2.yaml) pins
spec.interruptionBudget.count: 1, so nodewright reboots one matched node at a
time.
The budget is load-bearing, not a style preference. An omitted
interruptionBudget is not “unset”: the operator’s
createLegacyDefaultCompartment substitutes percent: 100, so the whole
matched GPU fleet reboots together. Nothing confines a generated bundle to
build time — it can be applied to a cluster already running work — so the safe
value has to be the default.
count: 1 costs N x (reboot time) to converge, and runtimeRequired: true
keeps workloads gated on the last node for the whole rollout. On an initial
cluster build, before any workload has landed, that serialization buys
nothing. Two equivalent ways to opt out of it for bringup:
- Delete the
interruptionBudgetblock from the rendered CR, or - set
percent: 100, which reproduces the pre-budget behavior exactly (countandpercentare mutually exclusive — set one or the other, never both).
Restore the budget before the cluster starts taking work.
GB300 host kernel granule
nvidia-gb300-performance sizes its hugepage pools for a 64k-page ARM64 kernel
(2M + 512M, and no 1G — a 64k granule has no PUD level). On EKS,
nvidia-setup-kernel installs the pinned 64k kernel, so profile and host agree
by construction.
On --service generic there is no nvidia-setup to pin one: it dispatches on
<service>-<accelerator> and ships no generic-* config, so a bare-metal GB300
node takes its kernel from the host image. A 64k-granule host kernel is the
recommended image there.
A 4k-granule kernel is not fatal. Linux rejects the invalid hugepagesz=512M
clause and silently drops the hugepages= count paired with it, so the node
boots and runs — it just comes up without the pool the profile intended, at
reduced performance. aicr bundle emits CheckGB300HostKernelGranule at
severity: info for that combination, so the tradeoff is visible at generation
time rather than only in a manifest comment.
Tuning-gke
A GKE + Container Optimized OS (COS) specific tuning that only sets some of the sysctl settings and does NOT require any interrupts due to being able to configure seamlessly while workloads are running.
See recipes/components/nodewright-customizations/manifests/tuning-gke.yaml
No-op
A no-op package may be used as a place holder until a full package suite can be tested. See recipes/components/nodewright-customizations/manifests/no-op.yaml