CLI Reference
Complete reference for the aicr command-line interface.
For details on which CLI verbs and critical user journeys are exercised by tests, on what hardware, and at what cadence, see the Coverage Matrix.
Version numbers in examples (component, chart, and driver versions) are illustrative and may not match any released recipe. The authoritative, current versions live in the Component Catalog and the Container Images BOM.
Overview
AICR provides a four-step workflow for optimizing GPU infrastructure:
Step 1: Capture system configuration Step 2: Generate optimization recipes Step 3: Validate constraints against cluster Step 4: Create deployment bundles
Global Flags
Available for all commands:
Logging Modes
AICR supports three logging modes:
-
CLI Mode (default): Minimal user-friendly output
- Just message text without timestamps or metadata
- Error messages display in red (ANSI color)
- Example:
Snapshot captured successfully
-
Text Mode (
--debug): Debug output with full metadata- Key=value format with time, level, source location
- Example:
time=2025-01-06T10:30:00.123Z level=INFO module=aicr version=v1.0.0 msg="snapshot started"
-
JSON Mode (
--log-json): Structured JSON for automation- Machine-readable format for log aggregation
- Example:
{"time":"2025-01-06T10:30:00.123Z","level":"INFO","msg":"snapshot started"}
Examples:
Commands
aicr snapshot
Capture comprehensive system configuration including OS, GPU, Kubernetes, and SystemD settings.
Synopsis:
Flags:
Output Destinations:
- stdout: Default when no
-oflag specified - File: Local file path (
/path/to/snapshot.yaml) - ConfigMap: Kubernetes ConfigMap URI (
cm://namespace/configmap-name)
Output Formats: --format applies to every destination.
yaml(default) delivers the agent’s document byte-for-byte to a file or stdout, so fields emitted by a newer--imagethan the CLI survive.jsonre-encodes that document with the same keys, which is whataicr diff --target snapshot.jsonandjqexpect. Sinceaicr diffpicks its decoder from the file extension, pair--format jsonwith a.jsonpath.tableis a flattenedFIELD/VALUErendering for humans and cannot be read back byaicr diff,aicr validate --snapshot, oraicr recipe --snapshot.- ConfigMap destinations store the rendering under the
snapshot.yaml,snapshot.json, orsnapshot.txtdata key alongside aformatkey; AICR’s ConfigMap readers follow that key, soyamlandjsonare both consumable fromcm://. Because those keys and the resource labels are derived from the document, acm://destination re-serializes it (deterministically, and without dropping unmodeled fields) rather than storing the agent’s exact bytes — use a file or stdout when you need byte-identical YAML. --templatesupplies its own rendering and therefore requires--format yaml(or no--format).
What it captures:
- SystemD Services: containerd, docker, kubelet configurations
- OS Configuration: grub, kmod, sysctl, release info
- Kubernetes: server version, images, ClusterPolicy
- GPU: driver-free PCI/NFD hardware detection — GPU presence, GPU count, accelerator model/SKU (resolved from the PCI device ID), nvidia kernel-module loaded state, and detection source. Does not require the NVIDIA driver or
nvidia-smi, and does not emit driver version, CUDA version, or MIG settings. - NodeTopology: node topology (cluster-wide taints and labels across all nodes)
- NetworkTopology: per-hardware-group network topology (PFs, capabilities, kernel modules, machine/GPU type, fabric type) — emitted only when
--cluster-configor--discover-networkis set
Examples:
Snapshot Config File Mode
Drive aicr snapshot from an AICRConfig document so the snapshot inputs version-control alongside the recipe, bundle, and validate steps in an end-to-end workflow.
Precedence: a CLI flag always wins over the matching config field. Selectors and tolerations omitted entirely inherit the snapshotter’s compiled-in defaults (tolerations defaults to tolerate all taints); an explicit empty list (tolerations: []) clears the tolerate-all default — the same nil-vs-empty semantics used by spec.validate.agent.
Custom Templates
The --template flag enables custom output formatting using Go templates with Sprig functions. Templates receive the full Snapshot struct:
Example template extracting key cluster info:
See examples/templates/snapshot-template.md.tmpl for a complete example template that generates a concise cluster report.
Agent Deployment
When running against a cluster, AICR deploys a Kubernetes Job to capture the snapshot. For the RBAC the agent creates, the in-cluster Job lifecycle, ConfigMap storage, and GPU-node auto-targeting (proactive selector injection plus the reactive placement-mismatch warning), see Agent Deployment.
ConfigMap Output
When using ConfigMap URIs (cm://namespace/name), the snapshot is stored directly in Kubernetes:
Snapshot Structure:
aicr recipe
Generate optimized configuration recipes from query parameters or captured snapshots.
Synopsis:
Modes:
aicr recipe resolves a recipe from criteria — service, accelerator, os, intent, platform. (nodes is accepted but is advisory metadata — it does not select or filter overlays.) You can supply those criteria three ways, composed with the precedence CLI flags > --config file > --snapshot:
- Snapshot (
--snapshot) — criteria are auto-detected from a captured cluster snapshot: accelerator from thenvidia.com/gpu.productGFD label (primary — cluster-wide, so it surfaces heterogeneous clusters) or the per-node PCI device ID (fallback — maps the device ID to a SKU, e.g.h100, with no driver or GFD label required); service from the node’s cloud-provider ID; OS from the node’s OS release; node count from cluster topology. Use this when you already have a running cluster: you don’t hand-specify the hardware, AICR reads it. (A detected SKU fillscriteria.acceleratoronly when it’s in the supported accelerator set; an unsupported GPU is recorded descriptively but not as a recipe criterion.) On the snapshot path AICR also reads the sampled GPU node’sdriver-loadedreading; on unprofiled/legacy compositions whose resolved values already carry the coordinated preinstalled-driver configuration, it injectsgpu-operator.overrides.driver.enabled=falseinto the resolved recipe, while on a profiled AKS composition the snapshot only qualifies the explicit-or-defaultgpuStackselection — the injector skips profile-owned paths and never mutates them; the inverse mismatch — a preinstalled-driver overlay resolved against a snapshot with no driver loaded — is handled per family: on profiled AKS it either fails closed at the profile constraint (pools not readingInstall) or recordsgpuDriverState: absentfor the bundle-time gate whose remedy is pool repair/recreation + recapture +--profile; on non-profiled families and legacy artifacts it warns with the bundle-time override set. See Component Catalog › GPU Operator Driver Auto-Detect for the gate, warnings, and the pre-deploy-snapshot requirement. - Config file (
--config) — criteria (and bundle settings) from anAICRConfigdocument; good for reproducible, version-controlled workflows. - Query flags (
--service,--accelerator, …) — state the criteria directly. Use this when there is no cluster to snapshot yet — the common case when you generate a recipe in order to provision a cluster, or build one offline/ahead of time for hardware you can’t reach.
Why you sometimes specify criteria the snapshot could detect: recipe generation does not collect live cluster state or deploy/modify workloads — it reads inputs (criteria, embedded/--data catalog, an optional snapshot) and does not snapshot a cluster for you. Its only cluster interaction is explicit cm:// ConfigMap reads/writes (snapshot input or recipe output). This keeps generation hermetic and reproducible (same inputs → same recipe) and lets it run before a target cluster even exists. So you state criteria explicitly when there’s no cluster to read; when a cluster is available, capture it first and let detection fill them in:
When both a snapshot and explicit criteria are given, the explicit values win (e.g. --snapshot … --service gke overrides the detected service). This is also the only way to select a generic recipe from a snapshot: a bare-metal cluster fingerprints as its provisioner (metal3, rke2), never as generic, so pass the override explicitly:
Configuration profiles: an overlay composition may declare one named
configuration choice. Select a non-default value with
--profile name=value; omitting the flag applies the declaration’s required
default. A selection against a composition with no declaration, a wrong name,
or an unknown value fails closed. Profile-bearing output records
metadata.selectedProfile, uses recipe apiVersion aicr.run/v1beta2, and
locks every declared owned path: divergent aicr bundle/aicr mirror
static overrides are rejected (identical values accepted), and
argocd-helm install-time values are rejected on key presence alone —
even when the value is identical to the selected one. The AKS family is the first embedded
adopter: gpuStack with values azure-managed (default) and operator-managed —
see AKS GPU setup.
The GKE family declares gpuStack with values gke-default (default; GKE’s
managed plugin stays the advertiser — recorded as advertiser: external —
for default-provisioned clusters with no node label) and bundle-installer
(the GPU Operator’s device plugin owns nvidia.com/gpu; GPU node pools carry
gke-no-default-nvidia-gpu-device-plugin=true and, because that label
forfeits GKE’s managed driver install, are created
gpu-driver-version=disabled — the bundle’s gcp-driver-installer
component supplies the driver with a recipe-pinned version); because the GKE
values govern advertisement, the #1327 allocation-policy paths are
closure-locked in addition to the declared owned paths — see
GKE GPU setup and
Component Catalog › GKE Device-Plugin Ownership.
The OKE family declares gpuStack with values oci-managed (default;
Oracle’s GPU node image supplies the driver and OKE’s NvidiaGpuPlugin
add-on advertises — advertiser: external) and operator-managed
(bring-your-own driverless image with the add-on removed; the operator
installs driver, toolkit, and plugin, with the DRA driver root in
lockstep). Each value is qualified by the add-on’s control-plane state,
supplied as an oci ce cluster list-addons --cluster-id <cluster-ocid> --all --output json dump via
--oke-addons on aicr snapshot and aicr validate — see
OKE GPU setup.
Profiles can also be exercised through a versioned external overlay.
Selection and verification are independent: --profile (or the default)
always decides the selected value — never the snapshot — and a supplied
--snapshot always verifies the selection against the cluster’s recorded
readings (fail-closed on mismatch or a missing reading; on AKS, the pool
mode from --aks-gpu-pools). Without a snapshot no check can run at
generation; the recorded constraint is enforced at aicr validate
readiness instead. See the with/without matrices in
AKS GPU setup and
GKE GPU setup.
Every stated criteria dimension must be honored: for service, accelerator, intent, os, and platform, resolution now enforces a coverage post-condition — if you state a value for one of these dimensions, at least one applied overlay must carry that exact value, or the request fails with an actionable INVALID_REQUEST error instead of silently returning a recipe that ignores what you asked for. The error names the uncovered dimension and, where possible, the additional criteria that would make the request resolvable (e.g. platform 'kubeflow' ... requires os (valid: ubuntu)). --nodes is exempt from this check — it is advisory only, since no overlay gates on node count, and stating it never requires matching node-count-specific recipe content to exist.
Snapshot-detected dimensions are advisory, not strict: on the --snapshot path, service, accelerator, and os can be auto-detected from the cluster fingerprint rather than stated by you. If the coverage post-condition above fails and every uncovered dimension came from that detection (none of them were set via a CLI flag or --config), AICR relaxes those dimensions back to unstated and retries resolution once, logging a warning that names each relaxed dimension and its detected value — this handles overlay trees that are deliberately agnostic to a dimension the snapshot still reports (e.g. an OS-agnostic Kind overlay tree observing os=ubuntu on the node). Dimensions you passed explicitly via a flag are never relaxed: if any uncovered dimension was explicitly stated, the error still fails the request.
Config File Mode (Recommended)
Generate recipes using an AICRConfig document. The same file format also drives the bundle command, so a single file can describe an end-to-end recipe-to-bundle workflow.
Flags:
The config file uses a Kubernetes-style envelope:
Individual CLI flags always override config file values. For slice/map flags, presence on the CLI replaces the file’s value (no append).
For a composition that declares a profile, the equivalent config field is
spec.recipe.profile. The CLI flag takes precedence. For example, on the
AKS family (the embedded adopter):
--config accepts a local file path or an HTTP/HTTPS URL. ConfigMap (cm://) sources are not supported; export the data with kubectl get cm <name> -o yaml and pass the resulting file.
Query Mode
Generate recipes using direct system parameters:
Flags:
Accelerator values name a GPU model, not a machine type. A provider
usually offers several machine types for the same GPU, and the machine type —
not the GPU — determines the fabric, the NIC count, and which components a
recipe can use. --accelerator h100 therefore does not, on its own, say which
node shape the resolved recipe targets. See
Qualified Machine Types below.
Service / Accelerator / OS / Intent / Platform value listings above are the OSS-embedded set. When
--dataregisters additional values (e.g., undisclosed providers, proprietary platforms), the CLI admits them at runtime through the criteria registry — see Data Extension.--criteria-strictrestores the OSS-only set regardless of what--datacontributes.
--service rke2and--accelerator vr200are Preview. They publish an early-adopter recipe path without the full production support and lifecycle qualification required for Supported status. See the published validation evidence at validation.aicr.run for current coverage.
--service k0sis Preview, covering the singlek0s / h200 / ubuntu / trainingcoordinate. See k0s H200 Setup for its prerequisites and known gaps.
Examples:
The AICR-provided example above is criteria-only: because it does not use
--snapshot, the resulting recipe carries no target-cluster MariaDB Operator
conflict evidence. Bundling warns that conflicts were not evaluated and
proceeds for compatibility. Use a current snapshot before deployment when you
need conflict detection. See
Conflict detection requires snapshot evidence.
Qualified Machine Types
Each recipe is qualified against a specific node shape. Criteria resolution does not reject another machine type of the same GPU model — there is no axis to reject it on — so a recipe always resolves. What differs by family is what happens afterwards: on some, deployment validation fails; on others it succeeds and only the performance gates are affected.
A row that names no intent applies to every intent for that accelerator and service. Where a row names a family rather than a machine type, the entry is a family-level statement — either because no component in that family binds to a machine type, or because the specific shape is not recorded in-repo. The row says which.
Two distinct failure modes are worth separating:
- Component-level (hard). Two families fail deployment outright, by
different mechanisms. The GKE H100 training lineage pins artifacts to a
machine type —
h100-gke-cos-trainingand the leaves inheriting it — so on a non-matching shape the DaemonSets have nowhere to land and a Chainsaw health check fails. The AKS H100 training lineage instead wires an RDMA fabric unconditionally, and a Go readiness gate in the deployment phase fails closed when no node advertises the shared RDMA resource — which is every non-IB NCads shape. Neither is a degradation; both stop the deployment phase. - Performance-gate (soft). Elsewhere the recipe deploys normally, but the NCCL and inference floors are fixed absolute values calibrated on full, high-bandwidth nodes. They are not normalized for GPU count or fabric class, so a smaller shape can fail a gate while being perfectly healthy. This is the EKS and GB200 case; on AKS the deployment gate above bites first. See Validation › Node-shape assumption. Normalizing these floors per GPU or per fabric class was considered and declined (#1256, #1254, both closed as not planned) — the floors are deliberately fixed absolute full-node values, so running a qualified shape is the supported way to pass them.
The table lists the accelerator/service pairs that have a qualified shape; a pair or a shape absent from it is undocumented rather than known-broken. It has not been qualified, and the criteria model has no axis that would distinguish it from one that has. Whether AICR should gain one — finer-grained accelerator values, a machine-type axis, or a fabric class — is tracked in #2377.
Snapshot Mode
Generate recipes from captured snapshots:
Snapshots derive infrastructure criteria such as service, accelerator, and OS,
but intent and platform remain explicit choices. K8s.slinky-slurm
reports declared Slinky Controller presence and, for a single Controller, a
secret-safe associated-resource topology; it does not select Slurm or
reconstruct the installed chart’s Helm values. K8s.mariadb-operator records
official MariaDB Operator API/CR conflict evidence but does not infer database
availability, accounting intent, or database source. Pass --platform slurm
to resolve a Slurm leaf. Older snapshots without either subtype remain
compatible. For aicr-provided, missing MariaDB Operator conflict evidence
causes bundling to warn that conflicts were not evaluated and proceed without
target-cluster conflict detection.
Flags:
Snapshot Sources:
- File: Local file path (
./snapshot.yaml) - URL: HTTP/HTTPS URL (
https://example.com/snapshot.yaml) - ConfigMap: Kubernetes ConfigMap URI (
cm://namespace/configmap-name)
Examples:
Output structure:
A profiled result applies the selected value’s constraints and component overrides to the normal recipe fields. It also uses a profile-aware recipe apiVersion and records the selected identity and declaration-wide lock surface:
aicr recipe list
Enumerate overlay recipes in the catalog. Useful for discovering which criteria combinations have a dedicated leaf overlay versus an intermediate shared recipe.
Each leaf overlay also carries a structural-health verdict (ADR-009 §4): a rolled-up status and a per-phase declared-coverage summary, computed by resolving the recipe and inspecting the result. Intermediate (non-leaf) overlays are not scored — only leaf overlays resolve to a concrete combination.
Synopsis:
Flags:
Filter flags narrow the output to overlays whose criteria carry that exact value. Unspecified flags match all overlays for that dimension. Multiple filters are combined with AND.
Output fields:
In table format the health axis is rendered as two extra columns: STATUS
(the rolled-up verdict) and COVERAGE, a compact per-phase named-check summary
of the form R:2 D:4 P:1 C:10 (Readiness / Deployment / Performance /
Conformance). Non-leaf overlays render - in both columns. A dimension whose
grader cannot reach a confident verdict surfaces as unknown, and the status
column still renders.
Resolving every leaf overlay to compute its health verdict adds latency on each
invocation. For purely interactive “what overlays exist?” lookups — or scripted
callers that only consume name/criteria/is_leaf/source — pass
--no-health (alias --skip-health) to skip the health computation entirely.
With the flag set the table format omits the STATUS/COVERAGE columns and
the json/yaml formats omit the health block, leaving only the enumeration
columns/fields.
Examples:
Example table output:
Intermediate (non-leaf) overlays render - in the STATUS/COVERAGE
columns; only leaf overlays are scored.
Example JSON output:
The criteria keys are capitalized because the criteria struct carries no
field tags; the structured output mirrors the Go field names. The health
block is present only for leaf overlays — non-leaf overlays omit it.
aicr recipe verify-catalog
Verify the embedded recipe catalog (registry.yaml + validators/catalog.yaml)
against its Sigstore bundle. The bundle is distributed as recipe-catalog.sigstore.json
release asset alongside each tagged aicr binary.
aicr recipe verify-catalog recomputes a deterministic SHA-256 over the embedded
catalog content using a length-prefixed encoding of the two raw files, then
verifies the digest against the Sigstore bundle using NVIDIA CI identity pinning.
Exit code is 0 on success, non-zero on any verification failure.
Synopsis:
Flags:
Examples:
Output:
On success, prints the verified content digest and the Fulcio certificate identity (the GitHub Actions workflow that signed the release):
aicr query
Query a specific value from the fully hydrated recipe configuration. Resolves a recipe
from criteria (same as aicr recipe), merges all base, overlay, and inline value
overrides, then extracts the value at the given dot-path selector.
Because aicr query resolves criteria the same way aicr recipe does, it is subject
to the same coverage post-condition: a stated service/accelerator/intent/os/platform
value that no overlay honors fails the query with an actionable error rather than
silently extracting a value from an unrelated recipe. --nodes remains advisory and
is never required to be covered.
Synopsis:
Flags:
All aicr recipe flags are supported, including --profile and --inherit-from, plus:
Selector Syntax
Uses dot-delimited paths consistent with Helm --set and yq:
Leading dots are optional (yq-style): .components.gpu-operator.chart and
components.gpu-operator.chart are equivalent.
Output:
- Scalar values (string, number, bool) are printed as plain text — no YAML wrapper
- Complex values (maps, lists) are printed as YAML (default) or JSON (
--format json)
Examples:
Advanced Examples:
aicr validate
Validate a system snapshot against the constraints defined in a recipe to verify cluster compatibility. Supports multi-phase validation with different validation stages.
For a task-oriented walkthrough (capture snapshot → generate recipe → run each phase, with worked training and inference examples), see Validation.
Synopsis:
Flags:
Input Sources:
- File: Local file path (
./recipe.yaml,./snapshot.yaml) - URL: HTTP/HTTPS URL (
https://example.com/recipe.yaml) - ConfigMap: Kubernetes ConfigMap URI (
cm://namespace/configmap-name)
Validation Phases
Validation can be run in different phases to validate different aspects of the deployment:
Note: Readiness constraints (K8s version, OS, kernel) are always evaluated implicitly before any phase runs. If readiness fails, validation stops before deploying any Jobs and exits 2 (
INVALID_REQUEST). This gate always fails closed —--fail-on-error=falsescopes to phase check results and does not downgrade a readiness failure.Declared-check pre-flight: Every check named under a phase’s
checkslist must resolve to exactly one catalog validator in that phase. Before any Job is deployed,validatefails closed with exit 2 (INVALID_REQUEST) if a declared check matches no validator (a typo, or a check missing from the loaded--datacatalog), exists only under a different phase, or is declared more than once — reporting every offender at once. This runs in--no-clustermode too, and like readiness it is independent of--fail-on-error. It replaces the previous warn-and-continue behavior, which let a phase with only unresolved checks reportskippedand exit0.Version skew: Snapshots and recipes record the
aicrversion that produced them. When the recipe, the snapshot, and the running binary report different release versions,validatelogs a single advisory warning (version skew detected across validate inputs) naming all three. This is a debugging breadcrumb — mixing artifacts from different versions can surface as confusing failures — and does not fail the command. Dev (dev) and pre-release (-next) builds are ignored to avoid noise.apiVersion gate: As of v0.22, the ADR-022 emitter switch, AICR emits
aicr.run/v1for snapshots and default recipes,aicr.run/v1beta1for config and ordinary catalog inputs, andaicr.run/v1beta2for profile-bearing recipes. Readers additionally still accept the supersededaicr.run/v1alpha2andaicr.run/v1alpha3, so artifacts produced by v0.21 or earlier keep loading. Unsupported artifact headers fail fast; raw external catalog headers are checked before merge or hydration. Recapture, regenerate, or update the authored header with a version supported by the running AICR release. See ADR-011 and ADR-022. v1.0.0 stops accepting the alpha values, along with the empty header that the snapshot, recipe, and criteria readers still tolerate. Reading either now logs a deprecation warning naming the file.AICRConfigand external catalog headers already reject an empty value, so they have no tolerance to retire. Catalog and binary compatibility has the release-by-release table.
Phases run sequentially with --phase all and all phases run by default, producing results regardless of earlier failures; use --fail-fast to stop after the first failing phase. For what each phase actually checks (deployment-phase readiness signals, graceful-skip semantics, RBAC, Day-N re-verification, and evidence), see Validation.
Constraint paths and operators
Constraints use fully qualified measurement paths: {Type}.{Subtype}.{Key}
Constraint paths are validated against the measurement catalog when recipe data
is loaded, so a typo such as K8s.server.verison fails immediately — naming the
file, the field, and the nearest matching key — instead of silently evaluating
as a missing reading. See
Recipe Development › Constraints for the
addressing rules.
Supported operators:
Examples:
Validate Config File Mode
aicr validate --config <path> reads inputs from an AICRConfig YAML/JSON file
under spec.validate. CLI flags always override values loaded from --config;
override events are logged at INFO so users can see which input won. The OIDC
identity token used for --push signing stays out of the schema by design
(short-lived tokens must not be committed); the CLI resolves it at sign time
through the precedence chain described on --identity-token.
Supported schema:
Examples:
The --node-selector and --toleration flags control scheduling for the inner validation workloads (NCCL benchmark workers, conformance test pods). When --snapshot is omitted, they also configure the preliminary live snapshot agent. They do not configure the validator orchestrator Job. For when to use them with non-standard GPU labels or taints, see Validation.
Output Structure (CTRF JSON):
Results are output in CTRF (Common Test Report Format) — an industry-standard schema for test reporting.
Note: The
testsarray above is truncated for brevity. A full validation run produces one entry per check across all phases. Each entry includesstdoutwith detailed diagnostic output.
Test Statuses:
Exit Codes:
aicr diff
Compare two snapshots field-by-field to surface configuration drift between cluster states. Reports added, removed, and modified readings across every measurement type (K8s, GPU, OS, SystemD, NodeTopology, NetworkTopology).
Synopsis:
Flags:
Inputs:
- File paths (
./baseline.yaml,/tmp/snap.json) - ConfigMap URIs (
cm://gpu-operator/aicr-snapshot) - Both inputs may mix freely; e.g., a local baseline file vs. a live ConfigMap target.
Output Semantics:
- A nil reading is rendered as the literal
<nil>so it cannot be confused with an empty-string value (""). Both forms surface as drift when one side is nil and the other is a concrete value. - Changes are emitted in deterministic order (sorted by
Path) so the diff is reproducible across runs and machines. - The
Resultenvelope includesbaselineSourceandtargetSource(the supplied paths), achangesarray, and asummarywithadded,removed,modified, andtotalcounts.
Items are ordered and compared by zero-based position. Reordering distinct records is drift. A list-length change emits items.length and field-level additions or removals. Ordinary existing Data paths remain unchanged for compatibility. Data keys equal to context or items, or beginning with context., items., or items[, use JSON-style bracket quoting so they cannot collide with structured paths; for example, subtype Data key context.node is emitted as <type>.<subtype>["context.node"], and the same key in Item Data is emitted as <type>.<subtype>.items[<index>].data["context.node"].
Examples:
Exit Codes:
Note on CI gating: A non-zero exit identifies that drift was detected, but doesn’t by itself distinguish drift from malformed input — both map to exit
2. To differentiate without relying on stderr format (text by default; JSON only with--log-json), inspect the diff payload directly: write the result with--output drift.json --format jsonand branch on the presence of the file plus itssummary.totalfield. That signal is format-stable regardless of logging mode.
aicr upgrade-check
Compare two recipes or bundles component by component and report, for each component whose version or namespace changed, whether moving between them is safe to apply. Verdicts come from the transition records the running aicr release ships. No cluster state is inspected, which makes this the CI and GitOps path: the comparison reads two artifacts and nothing else. A cm:// path is an artifact location like a file path, so reading or writing one does contact that cluster’s API for the ConfigMap itself.
Synopsis:
Flags:
Note the two defaults that differ from sibling commands. --format defaults to table rather than yaml, because the report’s payload is a list of operator steps that folded YAML scalars make unreadable. --fail-on-error defaults to true, the opposite of aicr diff --fail-on-drift: you chose to run this check, so its exit code is what makes running it worth something in a pipeline.
The two questions it answers:
--from X --to Y asks “is this specific move safe?” and presumes you already know your target.
Omitting --to asks “am I behind, and does catching up hurt?”, which is usually the real question: what an operator holds is an old artifact, not a chosen destination. The --from artifact’s embedded criteria are re-resolved against the running binary’s pins to synthesize the target.
Verdicts:
Anything other than safe exits non-zero. unknown is included deliberately: a transition nobody assessed is not a transition anyone approved, and the distance moved does not change that.
The report still reports a breaking boundary (a major bump, a minor bump while the major version is 0, or a changed prerelease identifier over an otherwise unchanged major.minor.patch) in the NOTES cell and the JSON breaking field, because the size of a move tells you how hard to look. It no longer affects the exit code.
Four kinds of unknown. They differ in what would close the gap, so they carry different reason codes and different report text:
blocked and unknown say opposite things. blocked means AICR has something to tell you and a version to stop at: read it and act on it. unknown means AICR has nothing for you: read the component’s own upstream release notes and decide. Neither is a pass.
Rollout note: expect red today. Only two registry components ship a transition record so far, so most components that change version report unknown and the check exits non-zero on most comparisons. That is a coverage problem being worked (#2535 makes records mandatory per pin bump), not a tool limitation, and it shrinks as records are authored. Use --fail-on-error=false if you want the report without the gate in the meantime.
Components whose version and namespace are identical on both sides produce no row. Added components are reported with nothing to do; removed components are reported and stay installed, because AICR does not uninstall them.
Namespaces are compared too. A recipe is regenerated from scratch on every AICR upgrade, and each component’s namespace comes from recipes/registry.yaml, so a moved registry default lands in the regenerated recipe. Helm cannot move a release between namespaces, so applying the resulting bundle installs a second copy of the component beside the running one, and nothing reconciles the two. A version-only comparison reports that as no change at all, so the namespace is compared as its own axis:
unknown rather than blocked is deliberate. blocked is an author’s judgement recorded against a version boundary, and the record vocabulary has no way to express one about where a release lives, so no author can record it. --format json and --format yaml carry the move as identityChanges, a list of {field, from, to} objects, on both kinds of row.
The remedy is not an upgrade step. Either move the release deliberately and re-run the check, or resolve the target recipe with aicr recipe --inherit-from so it keeps the namespaces the prior artifact deployed into and the relocation never enters the hop. See Upgrading a Deployed Stack.
The four routes to blocked. A record is crossed when your source version sits below the boundary its to names and your target reaches it. That is a property of the jump alone, so a record still counts even when the jump flies straight over it:
A crossed safe boundary does not count toward the second row: it carries no steps by construction, so crossing one composes nothing and skips nothing, and a jump whose only substantive boundary describes the whole move reports that record rather than stopping. The first row renders its record’s steps, deployer-scoped, exactly as a manual row does: the author marked the move blocked and then wrote what to do instead. The other three render none, because the record that carries them describes a different move than the one you asked about. All four name a stopping point.
The fourth exists because a record vouches only as far as its own to ceiling. A record claiming >=0.18.0 <0.19.0 says nothing about 0.25.0, and letting it lend its verdict there would report eight minors as safe on the strength of a two-minor claim. Stop at the assessed ceiling and re-run, or have the record widened. This is also how a component deliberately held below a breaking release reports: its record’s ceiling sits at that release, so any target above it is told to stop there and take the boundary on its own.
Every row states its reason in the detail block under the table, and --format json carries the same thing as reason (a stable code: recorded, record-blocks, multiple-boundaries, undefined-origin, beyond-record-ceiling, no-record, no-boundary-crossed, downgrade, identity-changed, not-comparable) plus explanation, the sentence naming your versions.
Why --deployer is required rather than defaulted:
Steps are deployer-scoped. Showing an Argo CD operator an imperative “delete the legacy CRDs” step is the exact failure deployer-scoping exists to prevent, so the command asks rather than guessing, and never renders every deployer’s path. It is only required when some component actually carries steps. A bundle now records the deployer that built it in bundle-info.yaml, so upgrade-check can stop asking once it reads that record (#2528).
Example:
More examples:
Exit Codes:
Note on CI gating: as with
aicr diff, a bad invocation and a failing check both exit2. To tell them apart without parsing stderr, write the report with--format json --output report.jsonand branch on the file’s presence plus itssummary.failingcount.
Limitations:
- A bundle is read through the
recipe.yamlat its root. Every deployer writes one as of #2759; a directory without it is neither a recipe nor a bundle and is rejected rather than misread. - Coverage starts near zero. Every transition without an authored record reports
unknown. See the authoring guide. - A namespace move is seen only when both artifacts state one. An absent namespace is read as a field the artifact did not carry, not as the default, so a component that gains or loses an explicit namespace between the two artifacts produces no relocation row. Reading it the other way would report a move nobody performed for every component the moment one side stopped carrying the field.
- No cluster comparison yet.
--from cluster, which reads installed Helm release inventory, is tracked in #2531.
aicr bundle
Generate deployment-ready bundles from recipes containing Helm values, manifests, scripts, and documentation.
Synopsis:
Flags:
--image-refs writes the published digest through a mode-0600 temporary
file and an anchored same-directory rename. Its target may be absent or an
existing retained regular file; directories, symlinks, other non-regular
files, and aliases of the bundle tree are rejected. The final validation and
rename are ordered but are not one atomic identity-conditioned filesystem
operation, so no other process should mutate the target directory while the
command runs.
For every CLI OCI publication, AICR revalidates a private bundle snapshot and publishes only its closed-world inventory; it never packages the mutable caller tree directly.
For Argo CD Helm OCI output, AICR keeps the raw Distribution tag in the
registry reference and derives a strict Helm semantic version for chart
metadata and consumers. For example, registry tag 1.2.3_build.5 remains
unchanged, while Chart.yaml, Argo CD targetRevision, and Helm’s
--version use 1.2.3+build.5. The OCI repository basename must equal the
generated chart name.
Bundle Config File Mode
The bundle command accepts the same AICRConfig format used by aicr recipe. A single file can populate both spec.recipe and spec.bundle, capturing an end-to-end workflow that can be committed to git, fetched from CI, or shared across environments.
When both spec.recipe.output.path and spec.bundle.input.recipe are set, they must reference the same path; otherwise loading fails fast.
CLI flags always override values loaded from --config. For slice/map flags (--set, --dynamic, --system-node-selector, etc.), CLI presence replaces the config’s value rather than appending. Override events are logged at INFO so users can see which input won.
Secrets: the cosign identity token is never read from a config file; supply it via --identity-token or COSIGN_IDENTITY_TOKEN.
Node Scheduling
The --accelerated-node-selector and --accelerated-node-toleration flags control scheduling for GPU-specific components:
NFD (Node Feature Discovery) workers must run on all nodes (GPU, CPU, and system) to detect hardware features. This matches the gpu-operator default behavior where NFD workers also run on control-plane nodes. The --accelerated-node-selector is intentionally not applied to NFD workers so they are not restricted to GPU nodes.
Note: When no
--accelerated-node-tolerationis specified, a default toleration (operator: Exists) is applied to both GPU DaemonSets and NFD workers, allowing them to run on nodes with any taint.
Example:
Cluster node requirements: This example assumes the cluster has nodes labeled
nodeGroup=system-workerwith taintsdedicated=system-workload:NoSchedule,NoExecutefor system infrastructure, and GPU nodes labelednodeGroup=gpu-workerwith taintsdedicated=worker-workload:NoSchedule,NoExecute.
This results in:
- gpu-operator operand DaemonSets (driver, device-plugin, toolkit, dcgm): no
nodeSelectoris applied — the gpu-operator chart and its ClusterPolicy CRD have nodaemonsets.nodeSelectorfield, so the selector cannot constrain them; the operator places these operands on GPU nodes itself via its GFD/NFD-drivennvidia.com/gpu.deploy.*labels. They do receive the tolerations fordedicated=worker-workload(bothNoScheduleandNoExecute) through the realdaemonsets.tolerationsvalue. - NFD workers: no nodeSelector (runs on all nodes) + tolerations for
dedicated=worker-workloadwith bothNoScheduleandNoExecute - System components (gpu-operator controller, NFD gc/master, dynamo grove, agentgateway proxy):
nodeSelector=nodeGroup=system-worker+ tolerations fordedicated=system-workloadwith bothNoScheduleandNoExecute
Behavior:
- All components from the recipe are bundled automatically
- Each component creates a subdirectory in the output directory
- Components are deployed in the order specified by
deploymentOrderin the recipe
DRA Driver Upgrade Eviction
This is opt-in. By default AICR injects nothing here, so the DRA kubelet plugin runs on every accelerated node with no extra node label required. Set --dra-eviction-node-label (or scheduling.draEvictionNodeLabel) to opt in.
What the opt-in does. When a recipe includes both nvidia-dra-driver-gpu and gpu-operator and a label is configured, AICR merges that key=value into kubeletPlugin.nodeSelector and sets the GPU Operator driver.manager.env entry NODE_LABEL_FOR_GPU_POD_EVICTION to the same key, so GPU Operator’s Driver Manager can deschedule the plugin ahead of a driver container restart. The same applies to the -ocp components. Existing accelerated-node selectors and unrelated Driver Manager environment variables are preserved.
The label is derived, not provisioned. The same opt-in keeps the dra-node-labeler component in the bundle: a small DaemonSet that runs on every node GFD reports as nvidia.com/gpu.present=true and applies the configured key=value once, after GPU Operator is up. It never rewrites an existing value, so the Driver Manager’s paused-for-driver-upgrade and a hand-set false opt-out both survive a labeler restart. Nodes added later by autoscaling, replacement, or a pool scaled from zero are labeled the same way as soon as GFD labels them. Without the flag, or when a bundlers filter leaves out the labeler, the labeler is left out of the bundle entirely (a bundlers selection that names the labeler without the flag or without both components it serves is rejected as invalid); node labeling is then required only if both components remain and eviction is still opted in. OpenShift recipes are not wired for the labeler yet (#2828) and keep the provisioning workflow. To provision the label yourself (node-pool labels, Karpenter NodePool labels) pass --set dra-node-labeler:enabled=false; the requirements in the next paragraphs then apply.
What you give up by not opting in. The plugin is not descheduled before a driver container restart. On a driver upgrade the module unload can fail with failed to uninstall nvidia driver components, leaving the replacement uninstalled and the old driver possibly partially torn down. On a restart with unchanged driver configuration the stale driver rootfs is unmounted underneath the running plugin; upstream documents this as leaving NodePrepareResources unable to build CDI specs, with no error at restart time. That second shape concerns the full-GPU allocation path, and we have not observed it in AICR’s default configuration — see the note below. aicr bundle warns about this when GPU Operator manages the driver.
This does not apply where the driver is provider-installed (driver.enabled=false — AKS azure-managed, GKE COS, OKE). Those deploy no GPU Operator driver pod and therefore no Driver Manager, so there is nothing to coordinate with and no warning is emitted.
If you disable the labeler, every GPU node must carry the label. Set it in the node pool definition — an EKS managed nodegroup labels entry, a Karpenter NodePool spec.template.metadata.labels entry, or the equivalent for your provisioner. An ad hoc kubectl label node is a repair, not a configuration: it does not survive node replacement, recycling, autoscaling, or a nodegroup scaled from zero. With the labeler in the bundle (the default once opted in) none of this is required; the remaining gap is a GPU node GFD has not labeled that does not already carry the configured key=value (for example from a node pool that still provisions it): such a node runs no kubelet plugin until nvidia.com/gpu.present appears.
The opt-in failure mode is silent, and partial coverage is the dangerous shape. An unlabeled GPU node then runs no DRA kubelet plugin and publishes no ResourceSlices for itself. If no GPU node carries the label, the nvidia-dra-driver-gpu-kubelet-plugin DaemonSet sits at DESIRED=0. If some do, those nodes work normally while the rest silently lack DRA. Helm and the bundle’s deploy.sh report success in every case, so check the DaemonSet after applying.
Recovering from the two opt-out failures. Both are recoverable, and the procedures are different — only the first needs the plugin suppressed.
Driver upgrade stopped with failed to uninstall nvidia driver components. The plugin still holds the driver, so the replacement is not installed and the old driver may be partially torn down. Recovery means clearing the node’s DRA claim holders and then removing the plugin before retrying the driver. Whether the removal can be confined to the failed node depends on the deployed tolerations, below.
kubectl cordon does not keep the plugin off the node. The DaemonSet controller adds a node.kubernetes.io/unschedulable:NoSchedule toleration to its pods, so a cordoned node still gets one and deleting the pod simply recreates it. Suspending a GitOps controller does not help either; the DaemonSet controller is what recreates the pod.
Do not patch the DaemonSet to exclude a node. Two properties of the shipped DaemonSet rule out per-node template edits: its affinity is five OR-ed nodeSelectorTerms, so excluding a node requires editing every term rather than one; and it uses RollingUpdate with maxUnavailable: 100%, so any change to .spec.template rolls pods on every node it covers — and reverting the change rolls them again. A node-scoped-looking patch is therefore cluster-wide in effect.
A node taint is the exception, and it is worth checking for. A taint acts on the node, not the DaemonSet template, so it rolls nothing. Whether it can work depends on the tolerations of your deployed DaemonSet — read them before choosing a procedure:
Expect a wildcard, because that is AICR’s default. When --accelerated-node-toleration is not passed, AICR applies a keyless {operator: Exists} toleration, which accepts every taint and leaves no key to exclude a node with. A default bundle therefore has no node-scoped option and must use the cluster-wide sequence below. This was the deployed state on an EKS GB300 cluster whose bundle was generated without toleration flags.
A taint works only when the deployed list is narrow — a bundle built with explicit --accelerated-node-toleration flags, or the upstream chart default of nvidia.com/gpu alone. Then a taint whose key appears nowhere in that list excludes the plugin from exactly that node.
“Node-scoped” describes the plugin suppression, not the blast radius of the whole procedure. The taint affects one node, and the DaemonSet keeps running everywhere else. But step 2 quiesces the controllers that own claim holders, and a controller is rarely node-scoped: scaling a Deployment, StatefulSet or multi-replica NodeSet to zero terminates its Pods on every node it runs on, not just the failed one. Scope that step as narrowly as your workloads allow — suspending a single Job, or cordoning and draining only what holds claims here — and treat wider quiescing as a maintenance window, not a node-local repair.
Three invariants govern the order, and every step below exists to preserve one of them:
- Fence before you drain. Terminating a claim holder while a controller can still schedule onto the node re-fills what you just cleared.
- The plugin outlives its claim holders. The kubelet routes
NodeUnprepareResourcesthrough the plugin, so a holder still terminating after the plugin is gone will hang. - The driver recovers before the plugin returns. A plugin that comes back early reopens the driver mid-recovery, reintroducing the original failure.
-
Fence the node. This is the very first operation — before touching any workload.
NoScheduleblocks new pods while leaving running ones alone, so the taint fences the node without disturbing the plugin, which is still needed. Every later step assumes this fence is up. -
Quiesce the controllers that own claim holders on that node. The invariant is that no owning controller can create a replacement Pod; confirm reconciliation is actually quiesced before relying on it. The action is workload-specific — set
.spec.suspend: trueon Jobs, scale Deployments, StatefulSets and standalone ReplicaSets to zero, scale operator-owned workloads (a Slinky SlurmNodeSet, for instance) to zero replicas. A Deployment-owned ReplicaSet must be controlled through its Deployment, which will otherwise recreate it.kubectl rollout pauseis not sufficient: it halts rollout progression while the ReplicaSet keeps reconciling replicas, so a deleted claim holder is recreated immediately.These actions terminate Pods — Job suspension deletes active Pods, scaling to zero removes them — and, as noted above, they do so wherever that controller runs. That is why the fence goes up first: the terminations on this node cannot then be undone by a controller rescheduling onto it.
-
Confirm every claim holder on that node has completed
NodeUnprepareResources. Allocated ComputeDomain claims can legitimately persist cluster-wide, so a cluster-wide claim listing does not establish that this node is clear — resolve holders to the node before proceeding. This is stated as an invariant rather than a command because the holder-to-node mapping depends on your workload shape and no single command was verified here. -
Only then delete the plugin on that node.
Select on
nvidia-dra-driver-gpu-component=kubelet-plugin, the DaemonSet’s own selector. The chart-wideapp.kubernetes.io/name=nvidia-dra-driver-gpulabel also matches the DRA controller, which can be colocated on a GPU node. -
Confirm at the node’s runtime that the plugin container is actually gone, not merely that the Pod object was deleted. Aggregate
kubectl get dscounts prove nothing about this node, and a stuckFailedKillPodleaves the container — and the driver handle — alive after the API object disappears. -
Retry the driver, and wait for it to finish. Driver Manager must report success and the driver Pod must reach
Ready. Do not proceed while it is still retrying. -
Only now remove the taint.
-
Confirm the plugin returns
Readyon that node and itsResourceSlicesare republished, then restore the controllers quiesced in step 2. Until the slices are back the node cannot serve DRA allocations, even though the pod is running.
NoSchedule does not evict running pods, so the driver pod stays put and its driver-manager init container keeps retrying while the plugin is gone. The DRA plugin on other nodes is untouched; whether their workloads are depends entirely on how wide step 2 had to reach.
The suppression mechanism in steps 1, 7 and 8 was verified on a two-GPU-node AKS cluster built with explicit tolerations, whose plugin tolerated only nvidia.com/gpu=present:NoSchedule: tainting one node took the DaemonSet from DESIRED=2 to DESIRED=1 with no replacement pod on the tainted node, while the second node’s pod kept its original start time — no roll. Removing the taint returned it to DESIRED=2 READY=2 with both ResourceSlices restored. That cluster uses a host-installed driver with no GPU Operator Driver Manager, so steps 3, 5 and 6 — the driver-recovery half — were not exercised there.
Cluster-wide sequence — required for the default wildcard-toleration bundle, and for any cluster where the check above shows no usable taint key:
- Stop workload reconciliation across every node the DaemonSet covers. Same invariant as the node-scoped path — no owning controller may create a replacement Pod, confirmed quiesced before you drain — and the same caveat:
kubectl rollout pauseleaves the ReplicaSet reconciling replicas, so scale replica controllers to zero and suspend Jobs instead. This matters more here than in the node-scoped path, because there is no taint fence to fall back on: quiescing the controllers is the only lever. - Clear DRA claim holders on every node the DaemonSet covers, not only the failed one, and confirm none remain. The kubelet needs the plugin to complete
NodeUnprepareResources, so any claim holder still terminating when the plugin goes away will hang. - Suppress the DaemonSet by deleting it. A DaemonSet has no replica count to scale, and editing its
nodeSelectorto match nothing is itself the cluster-wide template rollout — acceptable here only because this sequence has already accepted cluster-wide impact. Suspend any GitOps reconciliation first, or it will recreate the object underneath you. - Confirm every plugin pod has terminated. Pod-object deletion is not proof the container is gone; check the nodes.
- Retry the driver on the affected node and wait for it to finish — Driver Manager reporting success and the driver Pod
Ready. Restoring the DaemonSet while the driver is still retrying recreates the plugin and lets it reopen the driver mid-recovery, which is the same race the node-scoped path guards against. - Only then restore the DaemonSet, and confirm the plugin reaches
Readywith itsResourceSlicesrepublished. - Finally restore the workload controllers quiesced in step 1.
If a GitOps controller reconciles the DaemonSet, suspend it for the duration and resume afterwards.
Claim preparation failing after a restart with unchanged driver configuration. The stale rootfs was unmounted underneath a plugin that kept running, so upstream’s comment describes it as no longer able to build CDI specs. This one needs no DaemonSet suppression and no claim-holder clearing: once the replacement driver pod is Ready, restart the plugin pod on the affected node and the recreated pod binds the new rootfs.
Not reproduced in AICR’s default configuration. An attempt on an EKS GB300 cluster running the default ComputeDomain-only DRA setup (resources.gpus.enabled: false) reached the exact conditions — driver-manager logged skipping the uninstallation and Unmounting NVIDIA driver rootfs while the kubelet plugin was never evicted — and then a fresh ComputeDomain claim still prepared successfully, with the plugin not restarted. An existing claim holder also terminated cleanly, so NodeUnprepareResources was still being serviced.
The degradation upstream describes is about CDI specs, which is the full-GPU allocation path that AICR disables by default, so this shape may be unreachable in the default configuration. Treat that as one negative result in one configuration, not proof: full-GPU DRA was not tested, and only the unchanged-config restart path was exercised. The procedure above is retained because it is the correct response if the degradation does occur.
Known limitation of the opt-in. Under k8s-driver-manager v0.12 the configured label is paused in the same batch as other GPU operands, and no wait covers the standalone DRA kubelet plugin, so ordering against DRA claim holders and completion of plugin teardown are not guaranteed. The mechanism is the one NVIDIA documents, but it is best-effort here. See NVIDIA/k8s-driver-manager#250.
Support posture by GPU Operator version. The 26.3 DRA installation guide requires this label for its full-GPU allocation workflow; its ComputeDomain-only procedure does not. AICR’s default disables full-GPU DRA (resources.gpus.enabled: false) while keeping ComputeDomains, so an unlabeled default bundle does not satisfy the 26.3 full-GPU workflow — it does not put every bundle outside the guide. GPU Operator 26.7 documents the GPUCluster stack instead, which needs no separate label; AICR has not adopted that path.
The flag accepts exactly one Kubernetes label in key=value form; AICR uses the full pair for DRA placement and the key for GPU Operator:
GPU Operator’s Driver Manager receives only the label key; it does not receive or compare the configured value. The cluster’s node-labeling convention must therefore preserve the configured key/value pair when the label is restored.
The wiring is absent when either component is disabled, and when no eviction label is configured. Once opted in, direct value overrides for the managed selector key or NODE_LABEL_FOR_GPU_POD_EVICTION are overwritten so the cross-chart contract cannot drift, and a --dynamic declaration intersecting kubeletPlugin.nodeSelector or driver.manager.env is rejected because install-time editing would split the same contract. Without the opt-in AICR owns neither path, so both remain freely overridable and declarable. See NVIDIA’s GPU Operator DRA installation guide for the upstream driver-upgrade requirement.
Storage Class
The --storage-class flag injects a Kubernetes StorageClass name into components at bundle time. StorageClass is a cluster infrastructure detail — the right value depends on what the target cluster has provisioned, not on the recipe.
When provided, the value is written to all Helm value paths declared in the component registry under storageClassPaths, overriding any storageClassName set in recipe overlays. If a per-component --set <component>:<path>=<value> explicitly targets the same path, that value takes precedence over --storage-class.
Example:
When --storage-class is not set, any storageClassName values already defined in the recipe overlays are preserved as defaults. When it is set, --set <component>:<path>=<value> on the same path still wins — --storage-class only fills in paths that were not explicitly overridden.
If a rendered component creates a PVC at a registry-declared storageClassPaths entry and no usable storageClassName is set after overlay, --storage-class, and --set precedence is resolved, aicr bundle emits a non-blocking warning. The bundle still relies on the target cluster’s default StorageClass in that case.
aicr bundle reports cluster-state dependencies it cannot verify as non-blocking warnings of this kind. Two more concern DRA eviction. Without --dra-eviction-node-label, and only where GPU Operator manages the driver, the bundle warns that automatic eviction was not configured and that a driver restart carries the documented risks. With the flag set, it reports that dra-node-labeler derives the label from nvidia.com/gpu.present and that a GPU node GFD has not labeled, and that does not already carry the configured label, runs without DRA; with the labeler disabled it warns instead that every GPU node must carry the configured label, that the label belongs in the node pool definition rather than an ad hoc kubectl label, and that unlabeled nodes silently run without DRA — no ResourceSlices, and DESIRED=0 if no GPU node matches at all. These describe state AICR deliberately does not own — StorageClasses and node labels are cluster infrastructure. See DRA Driver Upgrade Eviction.
--shared-storage-class is a separate input for registry-declared
sharedStorageClassPaths. It is used by opt-in Slinky Slurm PVCs mounted at
/home and /scratch/fsw, which request ReadWriteMany; it does not affect
existing generic storage paths. Enable the feature with
--set slinkyslurm:storage.enabled=true. If either shared PVC has no effective
class after per-component overrides and --shared-storage-class are applied,
bundle creation fails instead of falling back to a potentially RWO-only class.
See Slurm Shared Storage.
In contrast, when a bundle includes the agentgateway component with an empty or unset allowedSourceRanges, aicr bundle is private by default: it injects the RFC1918 private ranges (10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16) into the inference-gateway’s loadBalancerSourceRanges and records a bundle note, so the deployed gateway is reachable from inside the cluster/VPC but denied to the public internet — it is never emitted open to 0.0.0.0/0. (Kubernetes treats an empty loadBalancerSourceRanges as allow-all, so the safe default has to be a real list.) An invalid value — a bare-string --set, a non-list, an unparseable CIDR, or a non-canonical CIDR such as 1.2.3.4/24 — is rejected with an error. To admit specific clients (e.g. a corporate VPN, which egresses from a public IP not covered by the default), scope it via a recipe componentRef override or the list-aware --set-json flag (agentgateway:allowedSourceRanges='["<cidr>"]'); to deliberately expose it publicly, opt in explicitly with '["0.0.0.0/0"]', which generates with a loud warning. See Inference Gateway Network Exposure.
Deployment Methods
The --deployer flag controls how deployment artifacts are generated:
Note:
--dynamicis not supported with--deployer argocd. Use--deployer argocd-helminstead, which produces a Helm chart where all non-profile-owned values are overridable at install time (a profiled recipe ships a lock template that rejects install-time values on profile-owned paths).
Note:
--dynamicdeclarations targeting the GPU allocation-policy keys (nvidia-dra-driver-gpuresources.gpus.enabled/gpuResourcesEnabledOverride,gpu-operator(-ocp)devicePlugin.enabled, or those components’enabledtoggle) are rejected: validators verify the recipe-resolved allocation policy, so its value cannot be deferred to install time. See Configured GPU allocation policy.
Note:
--dynamicis rejected on a path that equals, contains, or is contained by a path a component’srequireNodeSelector(or its conditionalrequireNodeSelectorIfStorageClassSetcounterpart, e.g.kube-prometheus-stack) marks as required (e.g.slinky-slurm,slurm-accounting-mariadb), or by a declaredstorageClassPaths/sharedStorageClassPathspath that conditionsrequireNodeSelectorIfStorageClassSet. This holds regardless of whether a storage class is configured yet, since one could be added later without rebuilding the bundle. The flag defers the value to install time, the same unpinned staterequireNodeSelectorexists to reject. Supply--system-node-selector/--accelerated-node-selectorinstead. SeenodeScheduling.systemvsaccelerated.
Deployment Order:
Ordering follows each component’s declared dependencies (dependencyRefs), not its position in the bundle. Components with no dependency between them deploy in parallel; a component with a real dependency waits for that dependency to become healthy. The NNN-<name>/ folder numbers reflect a valid serialization for readability only.
- Helm:
deploy.shinstalls one component at a time in deployment order (intentionally serial). - Argo CD / Argo CD (Helm):
argocd.argoproj.io/sync-waveannotation assigned by dependency depth. Independent components share a wave and sync together; a dependent lands in a later wave band that Argo starts only after the prior tier (including any readiness gate) is healthy. - Flux:
dependsOnreferences in eachHelmReleasemirror the component’s declareddependencyRefsdirectly (a component with no dependencies has nodependsOnand reconciles in parallel). Pre/post manifests preserve the per-component chain<name>-pre → <name> → <name>-post. The bundle’s rootkustomization.yamlis a plain Kustomize file (not a Flux Kustomization CR). - Helmfile: emits one
level-N.yamlsub-helmfile per dependency tier, processed in sequence (so each tier’s CRDs register before the next tier renders). Within a tier,needs:chains only a component’s own-pre → primary → -postreleases; independent components carry no edge, so helmfile applies them concurrently.
Pass --serial to force every deployer to install strictly one component at a time in deployment order (reverts argocd/argocd-helm/flux/helmfile to a linear chain; helm is already serial). For the full model and per-deployer rationale, see Deployment ordering.
Value Overrides
Override any value in the generated bundle files using dot notation:
Format: bundler:path=value where:
bundler- Bundler name (e.g.,gpuoperator,networkoperator,certmanager,nodewright-operator,nvsentinel)path- Dot-separated path to the fieldvalue- New value to set
Behavior:
- Duplicate keys: When the same
bundler:pathis specified multiple times, the last value wins - Overlapping scalar paths: Scalar paths are applied shallowest-first. If
one
--setassigns a non-map parent and another assigns one of its descendants (for example,comp:driver=truepluscomp:driver.enabled=false), the request is rejected regardless of flag order. Use one coherent object through--set-json/--set-filewhen a parent and its nested keys must be supplied together. - Array values: Individual array elements cannot be overridden (no
[0]index syntax).--setis scalar-only — pointing it at a list/object field writes a bare string and produces type-invalid output. To replace an entire array or object from the CLI, use--set-json/--set-file; recipe-level overrides incomponentRefs[].overridesare the alternative. - Type conversion: String values are automatically converted to appropriate types (
true/false→ bool, numeric strings → numbers) - Component enable/disable: The special
enabledkey controls whether a component is included in the bundle.--set <component>:enabled=falseexcludes a component the recipe enabled. A component the recipe disabled (overrides.enabled: false) cannot be re-enabled this way —--set <component>:enabled=trueon such a component is rejected, since re-enabling a platform-provided component would install a conflicting second copy. Theenabledkey is consumed by the bundler and not passed to Helm chart values. - Aliases merge: overrides supplied under both a component’s canonical name and a registered alias (e.g.
gpu-operatorandgpuoperator) are combined, not dropped; the canonical name wins on any shared path. (Same alias-merge behavior as--set-json/--set-file.) - GPU allocation-policy keys are deprecated at bundle time: static overrides of the nested policy values —
nvidia-dra-driver-gpuresources.gpus.enabled/gpuResourcesEnabledOverrideandgpu-operator(-ocp)devicePlugin.enabled— still work via--set,--set-json, or--set-filebut log a deprecation warning; the component-levelenabledtoggle of those components is honored only via scalar--set(the typed--set-json/--set-filepath rejectsenabledfor every component, as described above) and likewise warns. Validators verify the recipe-resolved allocation policy, so a bundle-time change surfaces as recipe/cluster drift; move the allocation mode to a recipe overlay.--dynamicon any of these keys is rejected (the value would be unknowable when the policy is resolved). This boundary covers only what the bundler renders: post-generation changes —argocd-helm/ Argo CD parameter overrides, install-timehelm --set, and manual edits to generated bundles — cannot be intercepted by AICR and are outside the guarantee; they surface later as recipe/cluster drift when validation verifies the recipe-resolved policy. See Configured GPU allocation policy. - Profile-owned paths are locked: on a recipe carrying
metadata.selectedProfile(ADR-015; the AKS and GKEgpuStackfamilies), a static override on a profile-owned path is accepted only when identical to the selected value — a divergent value is rejected,--set-json/--set-fileare always rejected for an owned component’senabledpresence key, and--dynamicis rejected on any intersection with an owned path. On GKE the lock additionally covers the closure-locked allocation-policy paths (devicePlugin.enabled, DRAresources.gpus.enabled/gpuResourcesEnabledOverride) — divergent overrides there are rejected rather than deprecation-warned; see GKE GPU setup. - Repeat to add; commas are literal: To supply multiple overrides, repeat the flag (
--set a:x=1 --set b:y=2). On thebundlecommand, commas inside a single slice-flag value are taken literally (not treated as a value separator), so a value containing a comma — and the comma-heavy JSON passed to--set-json— is preserved intact. This applies to all repeatablebundleflags (--set,--set-json,--set-file,--dynamic,--*-node-selector,--*-node-toleration,--workload-selector).
Examples:
Argo CD Deployer Options
The deployer prefix is reserved: with --deployer argocd or --deployer argocd-helm, --set deployer:<key>=<value> configures the generated Argo CD Applications instead of component chart values. Unknown deployer: keys are rejected, and the prefix is rejected entirely with any other --deployer type (helm, flux, helmfile).
Child Applications only: namePrefix, destinationServer, and project affect the per-component child Applications, not the parent app-of-apps. Application CRs are reconciled only from the cluster running Argo CD, so the parent stays on the control-plane cluster in project default — see the Argo CD cluster bootstrapping guide. The parent’s name is set with --app-name.
Prerequisites for remote destinations: AICR only writes deployer:destinationServer and deployer:project into the generated Applications — it does not configure Argo CD. The destination cluster must already be registered with Argo CD (argocd cluster add <context> or a declarative cluster Secret), and the Argo CD project referenced by deployer:project must permit the emitted destinations and source repositories; otherwise the child Applications fail to sync with a permission or unknown-cluster error.
Name limits: the composed child Application name (namePrefix + component name) must be at most 53 characters — Argo CD uses the Application name as the Helm release name, which Helm caps at 53 — and must not equal the parent Application’s name (set with --app-name). Both violations are rejected at bundle time, and argocd-helm bundles re-check them at helm template/helm install time for install-time --set deployer.namePrefix=... overrides.
Install-time overrides with argocd-helm: for --deployer argocd-helm, the three string keys ship as defaults in the bundle chart’s root values.yaml and can also be overridden at install time via helm install --set deployer.<key>=... (note the dot, not colon). The bundle ships a values.schema.json that Helm applies on install/upgrade/template/lint: unknown deployer.* keys (e.g. a destinationSever typo) and malformed values are rejected at install time instead of silently falling back to defaults. cascadeDelete is bundle-time only — it adds the resources-finalizer.argocd.argoproj.io finalizer (a list field, not overridable via --set) so deleting an Application also deletes its deployed resources.
Children-only rendering (deployer.includeRootApp, install-time only): argocd-helm bundles render the parent app-of-apps Application as a chart template, so an externally managed root Application pointed at the published chart — for example one created by a controller such as a Cluster API addon provider — would otherwise render a second parent (aicr-stack) that owns the same child Applications and fights the external root through its automated prune/selfHeal sync policy. Set deployer.includeRootApp: false in that root Application’s spec.source.helm.valuesObject (or helm install --set deployer.includeRootApp=false) to render children-only. The default (true) keeps the standalone helm install flow unchanged; the same published bundle serves both consumption modes. This is a boolean key — pass it with plain --set (not --set-string), and note it is not a bundle-time deployer: key.
Use --set-string for values Helm would type-infer: apart from the boolean deployer.includeRootApp, the schema is intentionally string-typed, and Helm’s plain --set parses booleans and bare numbers into their inferred types — helm install ... --set deployer.project=true delivers a boolean, which the schema rejects. Pass such values with --set-string so they stay strings: helm install ... --set-string deployer.project=true.
List and Object Value Overrides
--set is scalar-only: it cannot express a list or object value. Pointing it
at a list field such as agentgateway.allowedSourceRanges writes a bare
string at that path, which the bundler rejects with an invalid-request error
(it must be a list of CIDR strings). Use --set-json (inline) or --set-file
(from a file) for any list or object override.
Format: component:path=<value> where:
component/path— same component name and dot-separated path as--set<value>— for--set-json, a JSON-encoded value; for--set-file, a path to a regular file containing a single JSON or YAML value (one document — in a multi-document YAML file only the first document is read). A non-regular path (directory, FIFO/named pipe, device, socket) is rejected with a fast validation error rather than hanging or failing later.
Behavior:
- Object values deep-merge into any existing map at the path (partial
object overrides compose with recipe/base values), matching how inline recipe
componentRefs[].overridesmerge. Within a merged object, a JSONnulldeletes that key from the result (same explicit-null semantics as recipe overrides). - Lists and scalars replace the value at the path.
- Precedence: applied after
--set, so a--set-json/--set-fileentry wins over a scalar--seton the same path. Between the two typed flags, an inline--set-jsonwins over a--set-fileon the samecomponent:path(mirroring Helm’s--settaking precedence over-fvalue files). Within a single flag, the last entry for a givencomponent:pathwins. - Overlapping nested paths: when one override targets a parent object and
another targets a key beneath it (e.g.
comp:driver.env=<object>pluscomp:driver.env.HTTPS_PROXY=<value>), the deeper, more-specific path wins on the keys they share — regardless of the order the flags are given. The parent object’s other keys are preserved. - Aliases merge: overrides supplied under both a component’s canonical name
and a registered alias (e.g.
gpu-operatorandgpuoperator) are combined, not dropped; the canonical name wins on any shared path. - Node-scheduling paths deep-merge with CLI injection (asymmetric with
--set): a typed override on a node-scheduling path (e.g.--set-json gpu-operator:nodeSelector=<object>) deep-merges into the selectors and tolerations injected by--accelerated-node-selector/--system-node-selector/--*-node-toleration, rather than suppressing that injection. This is intentional — deep-merge is the point of the typed path — but it differs from scalar--set:--set comp:nodeSelector.x=ymarks that path as user-populated and suppresses CLI injection on it, whereas a typed override composes with the injected keys (so system-injected selector keys remain present alongside it). Use scalar--set, or omit the node-scheduling flags, when you need to fully replace an injected selector instead of merging into it. - CLI-only:
--set-json/--set-filehave noAICRConfig(spec.bundle.deployment.set) or HTTP API (?set=) equivalent — those surfaces remain scalar-only. To set a list or object value outside the CLI, use a recipe overlay orcomponentRefs[].overrides. - Not for
enabled: the specialenabledcomponent toggle is honored only via scalar--set; passing it through--set-json/--set-fileis rejected with an error (it would not toggle the component and would leak a strayenabled:value into the chart).
Examples:
Vendoring Charts for Air-Gap
The --vendor-charts flag pulls upstream Helm chart bytes into the bundle at bundle time. With the flag set, every Helm-typed component becomes a local chart inside the generated bundle and the resulting artifact deploys end-to-end with zero registry egress. Without the flag, deploy-time helm upgrade --install calls fetch from the upstream repository — which works for connected clusters but breaks in air-gapped environments.
Bundle-time requirement: the helm binary must be on $PATH when aicr bundle --vendor-charts runs. Authentication for private chart registries flows through Helm’s own conventions:
- HTTP(S) repositories —
HELM_REPOSITORY_USERNAME/HELM_REPOSITORY_PASSWORDenvironment variables. - OCI registries — standard docker config (
~/.docker/config.jsonor$DOCKER_CONFIG); rundocker login <registry>ahead of time.
Tradeoff: CVE-yank fail-loud signal is lost. Non-vendored bundles fail loudly when an upstream chart version is yanked at registry time, which prompts a rebundle with a fixed recipe. Vendored bundles freeze the chart bytes at bundle creation and silently install the frozen version even after upstream yank. Treat provenance.yaml (below) as the audit surface for cross-referencing yank lists.
Bundle-time costs. Vendoring adds bundle-time network egress (the chart pull), bundle-time auth surface (private registries need credentials at the bundle host), and bundle size (typically 0.5–5 MB unpacked per chart). Users who don’t need air-gap shouldn’t set --vendor-charts and shouldn’t pay these costs.
Bundle layout with --vendor-charts — every Helm component emits a wrapper folder holding the vendored tarball. Mixed components keep the primary + -post split they have on the non-vendored path, so recipe-side manifests stay tracked members of their own Helm release:
provenance.yaml sits at the bundle root and lists one entry per vendored chart, using the same K8s-style apiVersion/kind shape as the rest of AICR’s persisted formats:
The sha256 field is the digest of the bytes copied into charts/, suitable for yank-list lookups and cross-bundle drift comparisons. Pipe through yq -o=json provenance.yaml if your scanner expects JSON.
Examples:
Readiness Gates
The --readiness-hooks flag makes a deploy block on component-specific readiness signals rather than just the chart’s own resources reporting Ready. A component opts in by shipping a recipes/components/<name>/readiness.yaml Chainsaw test that asserts the signal that actually means “ready” — for example, gpu-operator waits for its ClusterPolicy to reach status.state: ready, which Helm and Argo CD cannot assess natively.
With the flag set, the bundler emits an extra folder, NNN-<name>-readiness/, immediately after each opted-in component. The folder is a small chart containing a Kubernetes Job (plus the ServiceAccount/RBAC and a ConfigMap holding the Chainsaw test). The Job runs the gate CLI (ghcr.io/nvidia/aicr-gate), which polls the test until it passes continuously for a stability window or a --max-wait ceiling elapses. The deploy blocks on that Job:
helm—deploy.shruns the readiness folder withhelm upgrade --install --wait. The gate Job is apost-install,post-upgradehook, and--waitblocks on hook completion regardless of--wait-for-jobs, so the latter is not needed. Helm’s own--timeoutis derived by the bundler from the gate’s--max-waitplus a buffer, so the gate owns the deadline (Helm never preempts it).argocd/argocd-helm— the readiness folder inherits the next sync-wave after its component, and Argo CD blocks that wave on the gate Job via its built-inbatch/Jobhealth (Progressing → Healthy on success, Degraded on failure). No custom health Lua and no directClusterPolicywatch — the readiness logic stays encapsulated in the Chainsaw test the Job runs.
flux and helmfile are not yet supported and --readiness-hooks is rejected for them. Components without a readiness.yaml are unaffected.
The gate evaluates the test in-process: it reads cluster state through its own ServiceAccount and applies the assertions itself, using the same executor aicr validate --phase deployment uses. The image ships no Chainsaw binary.
Two independent controls keep a readiness test read-only. First, the executor honors only the assert and error operations — every state-changing or side-effecting operation is rejected before evaluation, and the check fails. Second, the gate’s ServiceAccount is bound to a ClusterRole granting only get, list, and watch, so a mutating call would be denied by the API server even if one were somehow issued.
Dynamic Install-Time Values
The --dynamic flag declares value paths that are cluster-specific and should be provided at install time rather than baked into the bundle at build time. This enables building a single bundle that can be deployed to multiple clusters with different configurations.
Use --dynamic for values that genuinely vary per cluster — cluster names, subnet IDs, endpoint URLs, region-specific settings. For values that are static per bundle but differ from the recipe default (e.g., a specific driver version), use --set instead.
Attestation scope: Dynamic values are supplied at install time and are not covered by
--attest. Attestation binds the generated closed-world inventory, includingrecipe.yaml, not operator-provided overrides. If you need to constrain dynamic values at deploy time, use admission control or Argo sync hooks — see Attestation Scope.
Format: component:path where:
component- Component name or override key (same keys as--set, e.g.,gpuoperator,kubeprometheusstack)path- Dot-separated path to the value that varies per cluster
Helm deployer behavior:
Dynamic paths are removed from values.yaml and written to a separate cluster-values.yaml per component. The generated deploy.sh passes both files to Helm:
Before deploying, fill in cluster-values.yaml with cluster-specific values.
Argo CD deployer behavior:
The --deployer argocd-helm generates a Helm chart app-of-apps where all non-profile-owned values are overridable at install time. When the recipe carries a selected profile, templates/aicr-profile-lock.yaml fails the install if a value is supplied for a profile-owned path. Static values are baked into the chart as files; dynamic overrides are merged on top at render time. Use --dynamic to pre-populate specific paths in the root values.yaml for components that resolve to remote Helm charts. Local-chart components (Helm with no upstream Source, common on OCP overlays) and non-Helm components have no install-time stub surface — --dynamic naming them is rejected rather than silently dropped.
Examples:
Bundle structure with --dynamic (Helm deployer):
Bundle structure with --dynamic (Flux deployer):
The --deployer flux bundle uses Flux’s native spec.valuesFrom to reference ConfigMaps containing dynamic values. Dynamic paths are removed from the inline spec.values and placed into a ConfigMap per component. Flux merges valuesFrom first, then inline values on top — since dynamic paths are stripped from inline values, the ConfigMap values take effect without conflicts.
Before applying the bundle to your cluster, edit each configmap-values.yaml with the correct per-cluster values:
Bundle structure with --dynamic (Helmfile deployer):
The --deployer helmfile bundle references both values.yaml (static) and cluster-values.yaml (dynamic stubs) per release. helmfile merges value files in declaration order, so cluster-values.yaml overrides on top of the generated values.yaml. Edit cluster-values.yaml per component before helmfile apply:
Argo CD Helm chart structure with --dynamic:
The --deployer argocd-helm bundle is itself a Helm chart whose templates/ create per-component Argo Applications. Each application’s helm.values block merges static values (loaded via .Files.Get for upstream-helm components, or read from the wrapped chart’s own values.yaml for local-chart components) with dynamic overrides from the parent chart’s .Values.
The same uniform NNN-<component>/ folder layout used by --deployer argocd is included at the bundle root so that path-based Argo Applications (manifest-only, kustomize-wrapped, mixed -post) can resolve their path: references against the OCI-published bundle.
Manifest-only components and mixed-component raw manifests are supported by --deployer argocd-helm via the path-based Application shape.
static/ holds per-component values for upstream-helm Applications only. A recipe whose components are all local charts contributes no such files, and the directory is omitted from the bundle entirely rather than emitted empty. Most OpenShift (OCP) overlay components are local charts (OLM installer + CR pairs), but a few — components with no certified OCP operator, such as prometheus-adapter-ocp and nvidia-dra-driver-gpu-ocp — reuse their upstream Helm chart and do contribute a static/ values file.
The bundle’s repoURL defaults to the registry it was pushed to. No --repo flag is needed (and is ignored if passed with --deployer argocd-helm). When pushed to an OCI registry, the parent namespace is baked into values.yaml as the default repoURL — a plain helm install works with no --set repoURL needed. Override with --set repoURL=oci://mirror when deploying from a different registry.
Recommended deploy flow:
The chart’s templates/aicr-stack.yaml renders the parent Argo Application with .Values.repoURL and .Values.targetRevision substituted in. The parent Application then triggers Argo to render the chart again from the OCI source, creating the per-component child Applications with sync-wave ordering preserved. Child Applications whose source is path-based (manifest-only and mixed-component -pre / -post folders) inherit .Values.repoURL and append .Chart.Name so they pull from the same published artifact as the parent.
Argo CD OCI prerequisites. Path-based child Applications use Argo CD’s generic OCI artifact source type (introduced in Argo CD v2.13). The argocd-helm bundle therefore requires:
- Argo CD ≥ v2.13 on the target cluster.
- A registry that serves Helm-pushed OCI artifacts through the generic OCI manifest fetch path (most modern registries — ECR, GHCR, GAR, Harbor, Artifactory, plain
oras-compatible registries — support this).
If the recipe is pure-Helm (no manifest-only / mixed components), path-based children are not exercised and the bundle can work on Argo CD versions older than v2.13. If path-based children are present, Argo CD v2.13+ is required. See the troubleshooting section below if Failed to load target state appears on aicr-stack or any <component>-pre / <component>-post Application.
helm install ./bundle from a local directory also works, but with a caveat: child Applications whose source is path-based require Argo’s repo-server to fetch the bundle from a remote (git or OCI) — there is no local-filesystem source type for an Argo Application. Local helm install is therefore end-to-end only when the recipe contains pure-Helm components. For everything else, publish first.
Bundle structure (with default Helm deployer):
Folder layout rules:
- Folders are numbered
NNN-<component>/(1-based, zero-padded). Numbering is regenerated on every bundle. - Each folder is one of two kinds, distinguished by the presence of
Chart.yaml:- upstream-helm — no
Chart.yaml;upstream.envcarriesCHART/REPO/VERSION;install.shinstalls the upstream chart. - local-helm —
Chart.yaml+templates/;install.shinstalls the local chart (helm upgrade --install <name> ./).
- upstream-helm — no
- Mixed components (Helm chart + raw manifests) emit two adjacent folders: a primary upstream-helm
NNN-<name>/and an injected(NNN+1)-<name>-post/local-helm wrapper carrying the raw manifests. Subsequent components shift by one. - Manifest-only components (no upstream Helm chart, just raw manifests) become a single local-helm wrapped chart.
- Kustomize-typed components run
kustomize buildat bundle time; the output becomes a singletemplates/manifest.yamlinside a local-helm folder.
Breaking change vs. earlier releases:
Previous releases used a flat <component>/ layout with manifests/ siblings and a --deployer helm script that branched on component kind. The new format is uniform:
- All folders carry a rendered
install.sh. The top-leveldeploy.shis a generic loop with no per-component branching — name-matched special-case blocks (nodewright-operator taint cleanup, kai-scheduler async timeout, orphan-CRD scan, DRA kubelet-plugin restart) live around the loop, not inside it. - Raw manifests for mixed components now apply post-install only, via the injected
-postwrapped chart. The earlier pre-apply mechanism with a CRD-race retry wrapper is gone — Helm now owns CRD ordering for mixed components natively. - Tooling that parsed bundle paths by bare component name must account for the
NNN-prefix.
Argo CD bundle structure (with --deployer argocd):
The argocd deployer uses the same uniform NNN-<component>/ folder layout as --deployer helm. Each folder carries an application.yaml whose Application shape is decided by the folder kind:
Chart.yamlabsent (KindUpstreamHelm — pure Helm components): today’s multi-source Application pointing at the upstream Helm repository plus a values $ref to the user’s git repo. Unchanged for current users.Chart.yamlpresent (KindLocalHelm — manifest-only, kustomize-wrapped, mixed-post): single-source path-based Application withsource.path: NNN-<name>against the user’s repo.
The argocd deployer emits only what Argo CD’s repo-server consumes: application.yaml, values.yaml (multi-source helm.valueFiles for upstream-helm, or local-chart Helm rendering for KindLocalHelm), and Chart.yaml/templates/ for KindLocalHelm. The helm-deployer orchestration files (install.sh, upstream.env, cluster-values.yaml) are stripped — Argo doesn’t run shell scripts or source shell env, and --dynamic is rejected with --deployer argocd (use --deployer argocd-helm for install-time values).
Manifest-only components (e.g., nodewright-customizations) and mixed-component raw manifests (the -post injection) are now deployed by --deployer argocd. Previously they were silently dropped. Set --repo <user-git-or-oci> to populate the repoURL on path-based Applications so Argo can resolve them.
Day 2 Options:
The --workload-gate and --workload-selector flags are day 2 operational options for cluster scaling operations:
-
--workload-gate: Specifies a taint for nodewright-operator’s runtime required feature. This ensures nodes are properly configured before workloads can schedule on them during cluster scaling. The taint is configured in the nodewright-operator Helm values file atcontrollerManager.manager.env.runtimeRequiredTaint. When the flag is omitted the chart default applies:nodewright.nvidia.com=runtime-required:NoSchedulefrom operator v0.18.0 (previouslyskyhook.nvidia.com=runtime-required:NoSchedule, which the operator still recognizes and removes until v0.20.0). Node pools that pre-taint with the legacy key should pass--workload-gate skyhook.nvidia.com=runtime-required:NoScheduleso auto-tainted and pre-tainted nodes carry the same key.aicr validate --phase deploymentreads the configured taint from the operator Deployment, so whatever value is passed here is what the readiness gate waits to see cleared. For more information about runtime required, see the Nodewright documentation. -
--workload-selector: Specifies a label selector for nodewright-customizations to prevent nodewright from evicting running training jobs. This is critical for training workloads where job eviction would cause significant disruption. The selector is set in the Skyhook CR manifest (tuning.yaml) in thespec.workloadSelector.matchLabelsfield.
Estimated node count (--nodes):
The --nodes flag is a bundle-time option: it is applied when you run aicr bundle, not when you run aicr recipe. The value is written to each component’s Helm values at the paths declared in the registry under nodeScheduling.nodeCountPaths.
- When to use: Pass the expected or typical number of GPU nodes (e.g. size of your node pool). Use
0(default) to leave the value unset. - Where it goes: Components that define
nodeCountPathsin the registry receive the value at those paths in their generatedvalues.yaml. - Example:
aicr bundle -r recipe.yaml --nodes 8 -o ./bundleswrites8to every path listed in each component’snodeScheduling.nodeCountPaths.
Component Validation System:
AICR includes a component-driven validation system that automatically checks bundle configuration and displays warnings or errors during bundle generation. Validations are defined in the component registry and run automatically when components are included in a recipe.
How Validations Work:
- Automatic Execution: When generating a bundle, validations are automatically executed for each component in the recipe
- Condition-Based: Validations can be configured to run only when specific conditions are met (e.g., intent, service, accelerator)
- Severity Levels: Each validation can be configured as a “warning” (non-blocking) or “error” (blocking)
- Custom Messages: Each validation can include an optional detail message that provides actionable guidance
Validation Warnings:
When generating bundles with nodewright-customizations enabled, validation warnings are displayed for missing configuration:
- Workload Selector Warning: When nodewright-customizations is enabled with training intent, if
--workload-selectoris not set, a warning will be displayed:
- Accelerated Selector Warning: When nodewright-customizations is enabled with training or inference intent, if
--accelerated-node-selectoris not set, a warning will be displayed:
Viewing Validation Warnings:
Validation warnings are displayed in the bundle output after successful generation:
Resolving Validation Warnings:
To resolve the warnings, include the appropriate flags when generating the bundle:
Examples:
Argo CD Applications use multi-source to:
- Pull Helm charts from upstream repositories
- Apply values.yaml from your GitOps repository
- Deploy additional manifests from component’s manifests/ directory (if present)
Flux OCI Mode
When using --deployer flux with OCI output (--output oci://...), AICR generates ArtifactGenerator and ExternalArtifact CRs instead of GitRepository sources for local-chart components. This allows Flux to reconcile HelmReleases directly from OCI artifacts without a Git repository.
Prerequisites (Flux v2.7+):
- source-watcher controller must be deployed (
source.extensions.fluxcd.io). This controller watches ArtifactGenerator CRs and creates ExternalArtifact objects. - ExternalArtifact=true feature gate must be enabled on helm-controller. This allows HelmRelease CRs to reference ExternalArtifact objects via
spec.chartRef.
Without both prerequisites, bundles generate successfully but HelmReleases will not reconcile at deploy time.
Configuration flags:
The generated ArtifactGenerator CRs extract per-component chart directories from the outer OCIRepository into ExternalArtifact objects. Each HelmRelease then references the ExternalArtifact via spec.chartRef instead of the traditional spec.chart.spec.sourceRef pointing at a GitRepository.
Bundle Attestation
Prerequisite:
--attestneeds the binary attestationaicr-attestation.sigstore.jsonbeside theaicrbinary. The install script and the release archives both include it (keep it next toaicrwhen installing manually); binaries fromgo installlack it and cannot use--attest.
When --attest is passed, the bundle command performs five steps:
- Verifies the binary attestation file exists — The running
aicrbinary must have a valid SLSA provenance file (aicr-attestation.sigstore.json) alongside it, included by the install script from a release archive. If missing, the command fails immediately with guidance on how to install correctly. - Acquires a signing credential — in the default keyless mode this is an OIDC token (see OIDC Token Sources below); with
--signing-keythis step instead resolves the KMS key and no OIDC token is acquired (see KMS-Backed Signing). - Verifies the binary’s own attestation — Cryptographically verifies the SLSA provenance binds to the running binary and was signed by NVIDIA CI. This ensures only NVIDIA-built binaries can produce attested bundles.
- Signs the bundle — Creates a SLSA Build Provenance v1 in-toto statement binding the creator’s identity to the generated closed-world inventory, including
recipe.yaml, and the binary that produced it. - Writes attestation files —
attestation/bundle-attestation.sigstore.jsonandattestation/aicr-attestation.sigstore.jsonare added to the bundle output.
Attestation is opt-in; bundles are unsigned by default. By default, signing uses Sigstore keyless signing (Fulcio CA + Rekor transparency log) and records the entry in Rekor v2 (the signing config is fetched from Sigstore’s TUF repository, so shard rotation is handled automatically; a cold cache is fetched on demand). Verifying such bundles with aicr verify needs only the aicr binary; verifying them with cosign verify-blob-attestation needs Cosign v3.0.1+. For CI/CD environments without OIDC, pass --signing-key to sign with a KMS key instead; see KMS-Backed Signing below. For verification, see aicr verify.
Rekor v1 / private Sigstore: pass --rekor-url to sign to Rekor v1 at a specific URL — a private instance, or the public-good v1 URL — instead of the v2 default. Organizations running their own Fulcio CA can also redirect with --fulcio-url (both must be absolute https:// URLs with no embedded credentials); the two are independent. For a fully custom v2 setup, --signing-config takes a Sigstore signing config JSON (mutually exclusive with --rekor-url).
Verification: these flags redirect signing only. Verify the resulting bundles with
aicr verify --trust-root <trusted_root.json>, supplying thetrusted_root.jsonyour self-hosted Fulcio/Rekor emits. That root is unioned with AICR’s built-in public-good root, so privately-signed and NVIDIA-signed bundles both verify; see theaicr verify--trust-rootflag.
OIDC Token Sources
--attest resolves an OIDC identity token from the first matching source, in
order:
--identity-tokenflag (orCOSIGN_IDENTITY_TOKENenv) — a pre-fetched token. Use this when a token is obtained out of band (e.g., from a cloud workload-identity exchange or anothercosigninvocation). On shared hosts prefer the env var: a flag value is visible inpsand/proc/<pid>/cmdlineto any user on the same machine.ACTIONS_ID_TOKEN_REQUEST_URL+ACTIONS_ID_TOKEN_REQUEST_TOKEN— the ambient GitHub Actions OIDC credential. Used automatically in CI.--oidc-device-flowflag (orAICR_OIDC_DEVICE_FLOWenv) — OAuth 2.0 Device Authorization Grant (RFC 8628). The CLI prints a verification URL and short code; the user enters the code in a browser on a separate device. Use on headless hosts (bastions, remote build boxes) where the default browser callback cannot reach the machine runningaicr. The host still needs outbound network access to Sigstore’s OIDC and signing endpoints.- Interactive browser flow — opens the default browser and listens on a
random
localhostport for the redirect. Default on workstations.
Both interactive flows time out after 5 minutes.
Attestation works with all deployers (helm, argocd, argocd-helm, flux,
helmfile). External --data files copied into the bundle are included in
checksums.txt and listed as resolved dependencies in the attestation.
Privacy: identity in keyless signatures
Keyless signing trades a long-lived key for a short-lived Fulcio certificate minted from your identity — and that identity becomes part of the public record. Before any of the interactive sources above (device-code or browser) opens a login, understand what is published:
- The Fulcio certificate embeds the authenticated OIDC identity — typically your email address plus the OIDC issuer — in the certificate’s subject alternative name.
- With the public-good Sigstore defaults (
fulcio.sigstore.dev/rekor.sigstore.dev), the signature and certificate are recorded in the Rekor transparency log, which is public, append-only, and permanent — entries cannot be deleted, so the identity is globally searchable forever. - The same identity-bearing Sigstore bundle is attached to the pushed OCI artifact as a referrer, visible to anyone who can pull it.
This applies equally to aicr bundle --attest, aicr validate --emit-attestation --push,
and aicr evidence publish --push.
Controlling exposure. To avoid publishing a personal identity:
- Sign from a service / CI ambient identity (e.g. GitHub Actions OIDC) rather than a personal browser login.
- Supply a pre-fetched
--identity-tokenminted from a non-personal identity (a workload-identity exchange, a CI bot account). - Point
--fulcio-url/--rekor-urlat private Sigstore infrastructure so the identity stays inside your organization rather than the public commons. (Note: thevalidate --emit-attestation/evidence publishpaths always use the public-good endpoints today; onlybundle --attestexposes these flags.) - Use
--signing-key(KMS-backed signing) instead of keyless — no OIDC identity is embedded. See KMS-Backed Signing.
Interactive confirmation gate. Because the consequence is irreversible on
public Sigstore, the CLI gates the interactive login behind an explicit
confirmation. When aicr is about to open a browser or device-code flow it
prints a disclosure banner naming the Fulcio/Rekor endpoints in effect and what
will be published, then:
- On a TTY, it pauses for a
y/Nconfirmation (default no) and aborts cleanly — no browser opens — if you decline. - On non-interactive stdin (CI, pipes), it prints the banner and proceeds without blocking, so scripted/CI signing is never wedged.
- Pre-fetched
--identity-token, ambient GitHub Actions OIDC, and--signing-keypaths are not gated — they neither open a browser nor publish a surprise identity.
Pass --yes (alias --assume-yes, env AICR_ASSUME_YES) to skip the prompt for
trusted interactive automation; the banner is still printed.
KMS-Backed Signing
Keyless signing depends on an OIDC identity provider. Some CI/CD environments
(Jenkins, internal pipelines, air-gapped build hosts) have no OIDC issuer that
Sigstore trusts. For those, --signing-key signs the --attest bundle with a
KMS-backed key instead of a short-lived Fulcio certificate. The flag takes
a KMS URI; the supported schemes are:
awskms://: AWS Key Management Servicegcpkms://: Google Cloud KMSazurekms://: Azure Key Vaulthashivault://: HashiCorp Vault (Transit secrets engine)
Because a KMS key reference is durable and non-secret, it can live in a
version-controlled --config file as spec.bundle.attestation.signingKey
instead of on the command line (the --signing-key flag wins when both are
set). As with the flag, signing only runs when attestation is enabled, so a
config-only KMS workflow must also set spec.bundle.attestation.enabled: true
(the config equivalent of --attest); signingKey alone with enabled: false
produces no attestation. See Bundle Config File Mode.
--signing-key is mutually exclusive with the keyless-only flags
--identity-token, --oidc-device-flow, and --fulcio-url. Passing
--signing-key together with any of them is a validation error, since they
select incompatible signing modes (KMS key versus Fulcio-issued certificate).
Like keyless signing, KMS signs to Rekor v2 by default. Opt out with
--rekor-url to log to a Rekor v1 instance (private or public-good), or
--signing-config to use a custom signing config; the two opt-outs are mutually
exclusive with each other but both compose with --signing-key.
For a fully offline / air-gapped signing host that cannot reach any Rekor
instance, add --tlog-upload=false. KMS signing then skips the transparency-log
upload entirely and the bundle carries no Rekor entry:
--tlog-upload=false requires --signing-key; it is rejected on the keyless
path, because keyless OIDC signing needs Fulcio and Rekor network access to mint
a verifiable certificate. Verify an air-gapped bundle offline with
aicr verify --key <public-key.pem> --insecure-ignore-tlog, which skips the
transparency-log lookup that would otherwise require network access. Use a local
PEM public key (exported once with cosign public-key --key <kms-uri>) for a
fully offline verify: a KMS --key URI still makes a live GetPublicKey call.
The resulting bundle uses the same Sigstore bundle format as keyless signing, but its verification material is the signing key’s public key rather than a Fulcio certificate.
Verification: verify a KMS-signed bundle with
aicr verify --key <uri>, supplying the same KMS URI used to sign (or a local PEM public-key file). See theaicr verifyflags below. cosign’s public-key path (cosign verify-blob-attestation --key <same-kms-uri> ...) also works, since the bundle uses the standard Sigstore bundle format.
The hashivault://<transit-key-name> scheme signs through HashiCorp Vault’s Transit secrets engine (or an API-compatible server such as OpenBAO). It reads the server address and token from the standard VAULT_ADDR and VAULT_TOKEN environment variables (the OpenBAO equivalents BAO_ADDR and BAO_TOKEN are honored as fallbacks), so no additional flags are required; set TRANSIT_SECRET_ENGINE_PATH if the Transit engine is mounted somewhere other than the default transit/. Verify the resulting bundle with the same URI via aicr verify --key hashivault://<transit-key-name>.
Attestation Scope
Attestation binds a closed-world bundle inventory. checksums.txt contains one
SHA256 entry for every regular payload file — including recipe.yaml,
defaults, dynamic-value stubs, and external --data files copied into the
bundle. Verification derives the required directories and rejects every
additional file or directory, symlink, and other non-regular object. Only
checksums.txt, attestation/bundle-attestation.sigstore.json, and
attestation/aicr-attestation.sigstore.json may exist outside the manifest;
they remain part of the verified inventory and the attestation files are
verified separately. Manifest parsing is order-independent and AICR generates
entries sorted by canonical slash-relative path, but reordering an already
signed checksums.txt changes the signed bytes and invalidates its existing
attestation.
The attestation does not bind install-time values supplied via helm --set, a user-provided -f extra.yaml, or Argo
Application.spec.source.helm.parameters. That boundary is intentional:
dynamic values are the operator’s domain by design.
If you need to enforce specific install-time values (e.g., pinning driver.version), that is a policy concern, not an attestation one. Use admission control (Kyverno, Gatekeeper) or Argo sync hooks to reject deployments that violate the policy. aicr verify checks bundle integrity and provenance; it does not evaluate install-time value constraints.
Deploying a bundle
Note:
deploy.shis a convenience script — not the only deployment path. EachNNN-<component>/folder contains a renderedinstall.shthat runs the exacthelm upgrade --installcommand for manual or pipeline-driven deployment. For teardown, bundles delegate to the deployer-native uninstall path (see Bundle Uninstall below).
Deploy Script Behavior (deploy.sh)
The deploy script installs components in the order specified by deploymentOrder in the recipe.
Flags:
Unknown flags are rejected with an error to catch typos (e.g., --bes-effort or --retires N).
Note on install completion vs. workload readiness. By default,
deploy.shwaits on Helm chart readiness where AICR useshelm --wait. Some components are intentionally installed without Helm chart-level waiting, and the script does not wait for bundle-level workload readiness such as Nodewright node tuning, GPU operator operand rollout (driver, toolkit, device-plugin DaemonSets), or NVIDIA DRA kubelet plugin registration. Those continue asynchronously after the script exits. When--best-effortis used, the script may also finish with non-fatal component failures; check warning lines and logs before treating the install/apply pass as fully successful.--no-waitonly skips the Helm chart-level wait where AICR uses it; it does not affect bundle-level convergence.
Retry behavior:
The deploy script retries failed helm upgrade --install and kubectl apply operations with exponential backoff. By default, each operation is retried up to 5 times (6 total attempts). The backoff delay increases quadratically: 5s, 20s, 45s, 80s, 120s (capped) between retries.
Use --retries 0 to disable retries (fail-fast behavior). When --best-effort is also set, retries are exhausted first before falling through to best-effort handling.
Pre-install manifests and CRD ordering:
Some components have pre-install manifests (CRDs, namespaces, ConfigMaps) that must exist before helm install. The script applies these with kubectl apply before the Helm install. On first deploy, CRD-dependent resources may produce no matches for kind warnings because the CRD hasn’t been registered yet — these warnings are suppressed. All other kubectl apply errors (auth failures, webhook denials, bad manifests) fail the script immediately.
After helm install, the same manifests are re-applied as post-install to ensure CRD-dependent resources are created.
Async components:
Components that use operator patterns with custom resources that reconcile asynchronously (e.g., kai-scheduler) are installed without --wait to avoid Helm timing out on CR readiness.
DRA kubelet plugin registration
After installing nvidia-dra-driver-gpu, the script automatically restarts the DRA kubelet plugin DaemonSet. This is a best-effort mitigation for a known issue: after uninstall/reinstall, the kubelet’s plugin watcher (fsnotify) may not detect new registration sockets, causing DRA driver gpu.nvidia.com is not registered errors.
If DRA pods fail with this error after redeployment, the DaemonSet restart alone may not be sufficient — a node reboot is required to reset the kubelet’s plugin registration state. To reboot GPU nodes:
Bundle Uninstall
AICR bundles do not ship a generated undeploy.sh. Teardown is delegated
to the deployer-native uninstall path; AICR’s role ends at design-time
generation. Pick the walkthrough that matches the deployer used to generate
your bundle.
helm
Uninstall releases in reverse deployment order — the same order the
generated README.md lists under ## Uninstall:
Helm intentionally does not delete CRDs declared under crds/. PVC lifecycle
depends on how the claim is created: StatefulSet-created claims normally
outlive the StatefulSet, but standalone PVCs rendered as release resources are
deleted unless the chart marks them with helm.sh/resource-policy: keep.
Review the chart and the bound PersistentVolume reclaim policy before removing
storage:
Resources created by Helm hooks are not tracked as ordinary release resources
and can remain after helm uninstall. Review the component-specific cleanup
notes in the Component Catalog and the hook lifecycle
guidance in Bundling before deleting them manually.
If a release is stuck in pending-install or pending-upgrade (interrupted
deploy), retry with --no-hooks:
See Helm 3 uninstall docs for the full flag reference.
argocd
Delete the parent Application that owns the bundle’s child Applications
(app-of-apps). By default AICR does not set the
resources-finalizer.argocd.argoproj.io finalizer on generated
Applications, so a plain kubectl delete removes only the Application CR
and leaves the managed resources running. Bundles generated with
--set deployer:cascadeDelete=true (see
Argo CD Deployer Options) are the exception:
the finalizer is baked onto the parent and every child Application, so a
plain kubectl delete on the parent already cascades to the managed
resources. For default bundles, use one of the cascade-aware flows
instead:
If you can only use kubectl, add the finalizer first so the controller
performs the cascade for you:
The CRD and PVC notes from the helm walkthrough above still apply, but
Argo CD does not run helm uninstall for Helm-templated children. It renders
manifests with helm template and prunes rendered resources directly.
For standalone PVCs, Delete=false (or the equivalent
helm.sh/resource-policy: keep) prevents cleanup during Application deletion,
while Prune=false prevents pruning during manual or automated sync after a PVC
disappears from the desired manifests. StatefulSet-created claims are not
rendered as Application resources and normally remain.
See Argo CD app deletion docs for finalizer behavior, cascade modes, and selective deletion.
argocd-helm
Same path as plain argocd: Argo CD uses Helm only to render charts into
Kubernetes manifests (via helm template) and then manages those resources
itself. Deleting the Application with cascade enabled prunes the resources
Argo CD tracks; it does not run helm uninstall, and helm ls will
not show the bundle’s releases.
The kubectl + finalizer-patch fallback from the argocd walkthrough applies here too, and CRD / PVC cleanup follows the helm notes above.
See the Argo CD Helm user guide
and the Argo CD FAQ entry on helm ls
for why Helm CLI tools don’t see Argo-deployed releases.
flux
AICR’s flux bundle emits one HelmRelease per component (plus the
HelmRepository / OCIRepository source objects). Deleting each
HelmRelease from the cluster triggers helm-controller to run
helm uninstall for the underlying release, honoring the chart’s
spec.uninstall settings (disableHooks, keepHistory, etc.):
Delete the bundle’s source objects (HelmRepository / OCIRepository)
after the releases are gone. The CRD / PVC notes from the helm
walkthrough above still apply: helm-controller honors Helm resource-policy
annotations, while unannotated standalone PVCs are release resources and are
deleted during uninstall.
See the Flux helm-controller uninstall reference
for spec.uninstall field semantics.
helmfile
AICR’s helmfile bundle emits a single helmfile.yaml release graph.
The upstream helmfile CLI handles teardown:
CRD / PVC cleanup follows the helm walkthrough above. See the
Helmfile destroy documentation
for flags and behavior.
aicr mirror list
Discover container images and Helm charts referenced by a recipe for air-gapped
mirroring. Renders each component’s Helm chart with recipe-resolved values and
scans referenced manifests to produce a deduplicated image and chart list. When
the recipe was resolved with --data <dir>, both values and manifests are read
through the overlay so overlay-shadowed paths take precedence over embedded.
Recognized mapping-valued image descriptors with null, empty, or non-scalar
members are rejected so the command cannot succeed with a known-incomplete
image set.
For an end-to-end walkthrough covering Hauler and Zarf workflows, see Air-Gapped Mirroring.
Synopsis:
Flags:
Examples:
aicr verify
Verify the complete closed-world inventory and attestation chain of a bundle.
aicr verify . is the full verification command when run from the bundle
root. It rejects any additional file or directory, symlink, or other
non-regular filesystem object, except the three exact inventory metadata
paths. By default verification is offline and makes no network calls. The one
exception is --key with a KMS URI, which reaches the KMS provider to fetch
the public key (see the --key network behavior note below).
Synopsis:
Flags:
Verify Config File Mode
aicr verify --config <path> reads verification policy from an AICRConfig
YAML/JSON file under spec.verify. CLI flags always override values loaded from
--config; override events are logged at INFO so users can see which input won.
This is the one consumer-side section of the schema, so a single committed document can carry both the settings that build an artifact and the trust floor a downstream consumer enforces against it.
Supported schema:
Every field is a durable, non-secret reference or policy value, so the whole section is safe to commit. No private key material is part of the schema.
Three aicr verify flags are deliberately not in the schema:
- The bundle directory, which is a positional argument rather than a flag.
--format, which is presentation rather than policy.--insecure-ignore-tlog, which weakens the trust floor by dropping the transparency-log requirement. Keeping it command-line-only means a committed file can never silently disable that check, and an air-gap override stays an explicit operator act. It still composes with a config-suppliedkey.
Values are validated when the document loads, so a typo fails with its spec path
(for example invalid spec.verify.policy.minTrustLevel) rather than after a full
verification run. One limit is worth knowing: cliVersionConstraint is checked
at the operator level only, so ">=" with no version is rejected at load time
while ">= not-a-version" is accepted and fails later when the constraint is
evaluated.
Trust Levels
Verification steps
- Closed-world inventory — verifies every regular payload file is listed in
checksums.txt, all listed digests match, required directories are present, and no additional filesystem entries exist beyond the three allowed metadata paths - Bundle attestation — cryptographic signature verified against Sigstore trusted root
- Binary attestation — provenance chain verified with identity pinned to NVIDIA CI (
on-tag.yamlworkflow)
Inventory validation accepts valid manifest entries in any order, while AICR
generates canonical slash-relative entries in sorted order. Reordering a
signed checksums.txt invalidates its existing attestation. Legacy bundles
with incomplete manifests report unknown trust and must be regenerated.
Examples:
--keynetwork behavior: Resolving a KMS URI (awskms://,gcpkms://,azurekms://,hashivault://) makes network calls to the KMS provider to fetch the public key, so credentials for that provider must be available in the environment. A local PEM public-key file is read from disk with no provider calls; export it once withcosign public-key --key <kms-uri>(or your provider’s console) and verify anywhere.Resolving the key is only part of verification: by default the bundle’s Rekor transparency-log entry is also checked. Its inclusion proof is embedded in the bundle, so no live Rekor call is made, but the check needs the Sigstore trusted root. That root is loaded from the local cache and falls back to the embedded trusted root on a cache miss, so no network fetch happens on the verify path and
aicr trust updateis not required for offline use. For a bundle signed without a transparency-log entry at all (bundle --signing-key ... --tlog-upload=false), pass--insecure-ignore-tlogalongside--keyto drop the transparency-log requirement entirely; verification then runs with no network calls and no trusted timestamp, so use it only for true air-gapped bundles you signed yourself.Stale root: If verification fails with certificate chain errors, run
aicr trust updateto refresh the Sigstore trusted root.
aicr evidence digest
Print the canonical sha256 of a resolved recipe — byte-for-byte the same value recorded in predicate.recipe.digest by aicr validate --emit-attestation. The input is resolved through the same recipe builder path as aicr validate -r, so overlays and mixins are hydrated before hashing.
Use this to detect drift between a signed evidence pointer and the current recipe on a PR branch without pulling the OCI artifact.
Synopsis:
Flags:
Exit codes:
Examples:
aicr evidence publish
Sign, push, and write the pointer for a recipe-evidence bundle (predicateType v3; v1 and v2 remain verifiable for already-signed evidence) that was produced earlier by aicr validate --emit-attestation without --push (which leaves an unsigned bundle on disk).
This decouples the cluster-bound validate step from the Fulcio/Rekor-bound signing step so they can run on different networks: validation must run where the cluster is reachable (often a corporate VPN), but keyless signing must reach fulcio.sigstore.dev + rekor.sigstore.dev, which corporate networks frequently block. Run validate --emit-attestation on the VPN, then evidence publish from a host with Sigstore egress (CI runner, jump box, hotspot).
The unsigned subject/predicate — and therefore the OCI bundle digest — is identical regardless of which host ran which leg, because the predicate (including its baked-in attestedAt) is signed verbatim from the bundle on disk. The Sigstore signature, Fulcio certificate, and Rekor entry differ per signing run, so the signed bytes themselves are not byte-for-byte reproducible.
Synopsis:
The positional <bundle-dir> is either the directory --emit-attestation wrote (holds summary-bundle/ and receives pointer.yaml) or the summary-bundle/ directory itself.
Flags:
Identity disclosure:
evidence publishsigns unless--no-signis set. On the interactive (browser / device-code) keyless paths it publishes the signer’s identity (email + issuer) to the public Rekor log, so on a TTY it pauses for confirmation first (--yesskips it).--no-signruns no OIDC flow, so the prompt is skipped entirely. See Privacy: identity in keyless signatures.
Exit codes:
Examples:
aicr evidence sign
Complete the signing leg for a bundle that was already pushed unsigned (via aicr evidence publish --no-sign or validate --emit-attestation ./out --push <ref> --no-sign). It reads a local pointer file (e.g. ./out/pointer.yaml), pulls the bundle it references (bundle.oci + bundle.digest — no recipe-name or bundle-ref input needed), signs the predicate with keyless OIDC, attaches the Sigstore Bundle as an OCI referrer of the existing artifact, and patches the pointer’s signer block in place.
Signing is the only leg that needs Fulcio/Rekor egress, so this command runs wherever Sigstore is reachable (a jump box, CI runner, or hotspot) while the cluster-bound validate/push legs run wherever the cluster lives. The bundle is not re-emitted: the predicate is read verbatim from the pulled bundle, so the signature binds the same bytes the unsigned push produced.
The fork CI signing leg runs this with --relocate on the committed flat pending pointer — signing it and moving it to the nested per-source path the blocking Evidence Pointer Contract gate requires; you can also run it directly on a Sigstore-reachable host. See Publishing Recipe Evidence.
Synopsis:
The positional <pointer> is the committed flat pending pointer recipes/evidence/<recipe>.yaml. The pointer must carry exactly one attestation that is already pushed (bundle.oci/bundle.digest set) and not yet signed (empty signer); otherwise the command fails closed (an already-signed pointer is never re-signed, except under --relocate, which moves an already-signed pointer without re-signing).
With --relocate, the now-signed pointer is moved from its flat pending path to its canonical per-source path recipes/evidence/<recipe>/<src>/<digest>.yaml — the layout the per-source contract gate requires. A flat pointer is the only committable state for an unsigned pointer, because the <src> segment derives from the signer it does not yet have. This is the step the fork-based CI signing leg runs.
Flags:
Exit codes:
Examples:
aicr evidence verify
Verify a recipe-evidence bundle (predicateType v3 for newly produced evidence; v1 and v2 remain verifiable) produced by aicr validate --emit-attestation. When the bundle carries a signature, verifies it against the Sigstore trusted root and extracts the cryptographically anchored predicate. Recomputes every manifest-listed payload file’s sha256 against manifest.json (which the predicate’s manifest.digest field anchors), and surfaces the predicate’s fingerprint, phase counts, and BOM info.
Inline constraint replay is reserved for a follow-up PR.
Synopsis:
The positional argument is auto-detected as one of:
recipes/evidence/<recipe>/<src>/<digest>.yaml— pointer file (preferred). The verifier pulls by digest —registry/repo@<bundle.digest>, with the registry/repo taken frombundle.ociand the digest as the pin — so it fetches the exact attested bytes even if thebundle.ocitag has since been moved to a different artifact. This is the input to use in nearly all cases.ghcr.io/<owner>/aicr-evidence@sha256:...oroci://...@sha256:...— a digest-pinned OCI reference. A tag-only ref (such as thebundle.ocivalue copied from a pointer, e.g....aicr-evidence:h100-eks-ubuntu-training-3f9a1c2b4d5e) is refused by default because tags are registry-rewritable; see--allow-unpinned-tag../out/summary-bundle/(or a parent containing it) — unpacked directory.
Do not extract
bundle.ocifrom a pointer and pass it toverifyas a raw OCI argument. As a raw ref it carries no companionbundle.digest, so a tag-only ref is refused (tags are registry-rewritable). Pass the pointer file itself — the verifier readsbundle.digestfrom it and pullsregistry/repo@<digest>, ignoring the tag. If you must verify a raw OCI ref, use the digest form (...@sha256:<hex>), not the tag.
Flags:
Verdict codes:
These are the values of the exit field in the JSON/Markdown output, mirroring VerifyResult.Exit from the library API. They are not the process exit code — see the table below them.
Process exit codes:
Verdicts 1 and 2 are indistinguishable at the process level because both map through pkg/errors to the same code. To tell them apart, branch on the JSON exit field via jq '.exit' rather than on $?. Verdict 3 is deliberately given its own process code so a CI gate can distinguish an infrastructure fault from an invalid attestation without parsing JSON.
Pending signature. An unsigned bundle whose pointer carries no signer (e.g. one published with --no-sign, awaiting the signing leg) is not a verify failure: it verifies at exit 0 with pending: true in the JSON output and a “pending signature” verdict in the Markdown summary. This is a useful local check on a not-yet-signed bundle. The committed flat pending pointer is signed and relocated to its nested per-source path by the fork CI signing leg (aicr evidence sign --relocate); the blocking Evidence Pointer Contract gate (pkg/evidence/verifier/discover.go) requires that final signed, nested pointer. See Publishing Recipe Evidence.
Failure cause. On verdict 2 or 3, the JSON output carries a structured failureCause object — class (one of registry-forbidden, not-found, registry, signature, integrity, schema, transient, canceled, unknown), an optional httpStatus, and an actionable hint. For example, a private fork registry returns class: registry-forbidden, httpStatus: 403 with a hint to make the package public — so the reason is self-serviceable rather than a bare “invalid”. The Markdown summary renders the same as Cause/Hint lines.
Examples:
See demos/evidence.md for a full producer-and-consumer walkthrough.
Stale root: If verification fails with certificate chain errors, run
aicr trust updateto refresh the Sigstore trusted root.
aicr trust update
Fetch the latest Sigstore trusted root and Rekor v2 signing config from the TUF CDN and update the local cache at ~/.sigstore/root/. This is needed when Sigstore rotates signing keys (a few times per year). The trusted root is verification material (which signatures are valid); the signing config is sign-side (which Rekor/timestamp endpoints signing writes to — AICR signs to Rekor v2 by default).
Synopsis:
This command contacts tuf-repo-cdn.sigstore.dev, verifies the update chain against the embedded TUF root, and writes the result to ~/.sigstore/root/.
Flags:
When to run:
- After initial installation (the install script runs this automatically)
- When
aicr verifyreports a stale or expired trusted root - When Sigstore announces key rotation
Example:
aicr skill
Generate an AI agent skill file that teaches a coding agent how to use the AICR CLI. The generated file is written to the agent’s standard configuration directory.
Synopsis:
Flags:
Install Locations:
Behavior:
- Without
--stdout: writes the file to disk and prints the path - With
--stdout: prints the generated content to stdout - If the target file already exists: prompts
overwrite? [y/N]when stdin is a terminal; aborts on non-interactive stdin unless--forceis set - Creates parent directories as needed
Examples:
Complete Workflow Examples
File-Based Workflow
ConfigMap-Based Workflow (Kubernetes-Native)
E2E Testing
Validate the complete workflow:
Shell Completion
Generate shell completion scripts:
Installation:
Bash:
Zsh:
Environment Variables
AICR respects standard environment variables:
Exit Codes
Common Usage Patterns
Quick Recipe Generation
Save All Steps
JSON Processing
Multiple Environments
Troubleshooting
Snapshot Fails
Recipe Not Found
Bundle Generation Fails
“helm CLI not found on PATH” with --vendor-charts — the bundle-time vendoring path shells out to helm pull. Install Helm v3 or later (brew install helm / package manager) and re-run, or drop --vendor-charts for a registry-referencing bundle. See Vendoring Charts for Air-Gap.
“failed to load manifest <path> for component <name>” — the recipe references a manifest path that does not exist in the current AICR binary’s embedded data. This usually means the recipe was generated by an older binary and a referenced manifest has since been removed or relocated. Regenerate the recipe with the current binary (aicr recipe ...) and re-bundle. AICR recipes are a point-in-time artifact of the binary that produced them; bundling a stale recipe against a newer binary is not supported.
--deployer argocd-helm: aicr-stack or <component>-pre / <component>-post Application stuck at Unknown sync status / “Failed to load target state: … <registry>/<path>:<tag>: not found” — Argo CD cannot resolve the OCI artifact the parent or path-based child Application points at. Common causes:
-
Chart name doubled in
--set repoURL. Under the current contract,--set repoURLcarries the parent namespace only (e.g.,oci://ghcr.io/myorg). The parent Application appends.Chart.Nameinto its OCIsource.repoURL, and path-based children append it directly into their renderedsource.repoURL. For non-OCI Helm repositories, the parent usessource.chartinstead. Passing--set repoURL=oci://ghcr.io/myorg/aicr-bundleproduces a double-suffixed reference (.../aicr-bundle/aicr-bundle:<tag>) that does not exist. Drop the trailing chart segment. -
Argo CD older than v2.13. Path-based children rely on Argo CD’s generic OCI artifact source type, added in v2.13. Older Argo treats the source as Git and fails to resolve. Check with
kubectl -n argocd get deploy argocd-repo-server -o jsonpath='{.spec.template.spec.containers[0].image}'. Upgrade Argo, or use--deployer helmif Argo upgrade is not an option. -
Tag missing from the registry. Verify the published artifact exists at the exact tag the parent expects:
oras manifest fetch <registry>/<path>/<chart>:<tag>. Ifaicr bundleis invoked without a tag (oci://<registry>/<path>/<chart>with no:<tag>suffix), the CLI version is used as the default — make sure--set targetRevision=<chart-version>at install time matches. -
Private registry credentials keyed to a different source URL. Problem: Argo CD matches repository credentials against the source URL it dereferences.
Failure case: For this deployer, path-based OCI Applications render full
oci://<registry>/<path>/<chart>source URLs even though--set repoURLis the parent namespace. A Secret keyed only to<registry>/<path>or to a scheme-less Helm-OCI URL may let localhelm installsucceed while Argo’s repo-server still returns 401.Solution: Key the Argo CD repository credential to the rendered
oci://.../<chart>prefix, or to a broader matching prefix allowed by your cluster’s credential policy, such asoci://<registry>/oroci://<registry>/<path>/.
External Data Directory
The --data flag enables extending or overriding the embedded recipe data with external files. This allows customization without rebuilding the CLI.
Overview
AICR embeds recipe data (overlays, component values, registry) at compile time. The --data flag layers an external directory on top, enabling:
- Custom components: Add new components to the registry
- Override values: Replace default component values files
- Custom overlays: Add new recipe overlays for specific environments
- Registry extensions: Add custom components while preserving embedded ones
Directory Structure
The external directory must mirror the embedded data structure:
Requirements
- registry.yaml is required: The external directory must contain a
registry.yamlfile - Security validations: Symlinks are rejected, file size is limited (10MB default)
- No path traversal: Paths containing
..are rejected
Merge Behavior
Usage Examples
Example: Adding a Custom Component
- Create external data directory:
- Create registry.yaml with custom component:
- Create values file for the component:
- Create overlay that includes the component:
- Generate recipe with external data:
Debugging External Data
Use --debug flag to see detailed logging about external data loading:
Debug logs include:
- External files discovered and registered
- File source resolution (embedded vs external)
- Registry merge details (components added/overridden)
Example Files
The examples/ directory contains reference files for testing and learning:
Recipes (examples/recipes/)
Usage:
Templates (examples/templates/)
Usage:
See Also
- Installation Guide - Install aicr
- Agent Deployment - Kubernetes agent setup
- API Reference - Programmatic access
- Architecture Docs - Internal architecture
- Data Architecture - Recipe data system details