Agent Deployment
Deploy AICR as a Kubernetes Job to automatically capture cluster configuration snapshots.
Overview
The agent is a Kubernetes Job that captures system configuration and writes output to a ConfigMap.
Deployment: Use aicr snapshot to deploy and manage the Job programmatically.
What it does:
- Runs
aicr snapshot --namespace gpu-operator --output cm://gpu-operator/aicr-snapshoton a GPU node - Writes snapshot to ConfigMap via Kubernetes API (no PersistentVolume required)
- Exits after snapshot capture
What it does not do:
- Recipe generation (use
aicr recipeCLI or API server) - Bundle generation (use
aicr bundleCLI) - Continuous monitoring (use CronJob for periodic snapshots)
Use cases:
- Cluster auditing and compliance
- Multi-cluster configuration management
- Drift detection (compare snapshots over time)
- CI/CD integration (automated configuration validation)
ConfigMap storage
Agent uses ConfigMap URI scheme (cm://namespace/name) to write snapshots:
The agent’s namespaced Role grants ConfigMap write access only in its
deployment namespace (--namespace, default default). The cm:// target
namespace must match --namespace — otherwise the Job’s ServiceAccount has
no permission to create the ConfigMap and the snapshot write fails.
This creates:
Prerequisites
- Kubernetes cluster with GPU nodes
- aicr CLI installed
- GPU Operator installed (or appropriate namespace configured via
--namespace) - Permission to create and delete the run’s Job and RBAC in the target
namespace, plus the cluster-scoped
ClusterRole/ClusterRoleBindingthe agent needs. Every run starts by verifying this and stops before touching the cluster if anything is missing — see Pre-flight permission gate. Pointing--service-account-nameat a ServiceAccount you provisioned yourself requires no RBAC permissions at all
Quick Start
1. Deploy Agent with Single Command
This single command:
- Creates RBAC resources (ServiceAccount, Role, RoleBinding, ClusterRole, ClusterRoleBinding)
- Deploys Job to capture snapshot
- Waits for Job completion (5m timeout by default)
- Retrieves snapshot from ConfigMap
- Writes snapshot to stdout (or specified output)
- Cleans up Job and RBAC resources (use
--no-cleanupto keep for debugging)
2. View Snapshot Output
Snapshot is written to specified output:
3. Customize Deployment
Target specific nodes and configure scheduling:
Available flags:
--kubeconfig: Custom kubeconfig path (default:~/.kube/configor$KUBECONFIG)--namespace: Deployment namespace (default:default)--image: Container image (default: matches the CLI version, e.g.ghcr.io/nvidia/aicr:v0.19.0; dev and snapshot builds use:latest)--image-pull-secret: Secret name for pulling the agent image from a private registry (repeatable)--job-name: Job name prefix (default:aicr); the run ID is always appended (<prefix>-<run-id>)--service-account-name: ServiceAccount the agent pod runs as. Exact-if-exists — an existing ServiceAccount of exactly this name in--namespaceis used verbatim and the run creates no RBAC; otherwise it is a name prefix (default:aicr) and the run ID is appended (<prefix>-<run-id>). See Using an existing ServiceAccount--add-roles-to-service-account: Writes manifests and applies nothing. Renders the RBAC that grants the agent’s permissions to the named ServiceAccount into./snapshot-rbac-<run-id>/and exits without taking a snapshot. No cluster is contacted. You review the files, then apply and later delete them yourself. See Using an existing ServiceAccount--node-selector: Node selector (format:key=value, repeatable)--toleration: Toleration (format:key=value:effect, repeatable). Default: all taints are tolerated (usesoperator: Existswithout key). Only specify this flag if you want to restrict which taints the Job can tolerate.--timeout: Wait timeout (default:5m)--no-cleanup: Skip removal of Job and RBAC resources on completion. Warning: leaves the run-scopedaicr-node-reader-<run-id>ClusterRole and ClusterRoleBinding active. By default these grant only read access to nodes, pods, ClusterPolicy CRDs, Slinky Controller/NodeSet/LoginSet/RestApi/Accounting CRs, and official MariaDB CRs (not cluster-admin); however, when combined with--discover-networkthe retained ClusterRole also carries the cluster-scoped mutating discovery rules (CRD/namespace/DaemonSet create-delete,pods/exec,nodes/patch,NicClusterPolicypatch — see Security Considerations), so it is not read-only in that case.--privileged: Run agent in privileged mode (default: enabled; required for GPU/SystemD collectors). Set tofalsefor PSS-restricted namespaces.--require-gpu: Fail the snapshot if no GPU is found. In agent mode also requests annvidia.com/gpuresource for the pod (required in CDI environments).--runtime-class: SetruntimeClassNameon the agent pod fornvidia-smiaccess without consuming a GPU. Use with--node-selectorto target GPU nodes.--os: Node OS family (ubuntu,rhel,cos,amazonlinux,ol,talos). Selects the per-OS pod configuration and service collector backend.--requests/--limits: Override agent container resource requests/limits (comma-separatedname=quantitypairs).--cluster-config: Path to a pre-existing k8s-launch-kit cluster-config.yaml to ingest network topology (local agent mode only).--oke-addons: Path to anoci ce cluster list-addons --cluster-id <cluster-ocid> --all --output jsondump, projected into theK8s.oke-addons.nvidia-gpu-pluginreading. Controller-side: works in agent Job mode too — the file never enters the pod; the CLI merges the projection into the returned snapshot.--aks-gpu-pools: Path to anaz aks nodepool list -o jsondump, projected into theK8s.aks-gpu-pools.gpu-driverreading. Controller-side: works in agent Job mode too — the file never enters the pod; the CLI merges the projection into the returned snapshot.--discover-network: Enable live l8k discovery to populate the NetworkTopology measurement. Not read-only — writesnvidia.kubernetes-launch-kit.*node labels and may patchNicClusterPolicy.
4. Check Agent Logs (Debugging)
If something goes wrong, check Job logs:
Customization
Node Selection
Target specific GPU nodes using --node-selector:
Common node selectors:
Tolerations
By default, the agent Job tolerates all taints using the universal toleration (operator: Exists without a key). Only specify --toleration flags to restrict which taints are tolerated.
Common tolerations:
Image Version
Pin to a specific version:
Finding versions:
- GitHub Releases
- Container registry: ghcr.io/nvidia/aicr
Using an existing ServiceAccount (IRSA and Workload Identity)
By default the agent creates its own ServiceAccount for each run and deletes it
at cleanup, so no two runs share an identity. That does not work when the
ServiceAccount must carry cloud IAM credentials: EKS IRSA
(eks.amazonaws.com/role-arn) and GKE Workload Identity
(iam.gke.io/gcp-service-account) both pin trust to the ServiceAccount name.
IRSA’s role trust policy conditions on
system:serviceaccount:<namespace>:<name>, and a GKE IAM binding names the KSA
as PROJECT.svc.id.goog[<namespace>/<name>] and accepts no wildcard. A
per-run name can never be trusted by either, and copying the annotations onto a
run-scoped ServiceAccount does not help.
--service-account-name therefore behaves as exact-if-exists:
An unset --service-account-name is never probed for existence, so a stray
ServiceAccount named aicr cannot silently capture a run.
Migrating a pre-created ServiceAccount
What changed. Before run isolation, passing --service-account-name at a
ServiceAccount you had created out of band got that ServiceAccount adopted, and
aicr attached its Role and RoleBinding to it. Run isolation made every
run-owned object run-scoped, which turned the flag into a prefix — so a
pre-created ServiceAccount stopped being used at all, and an agent pod that
had been running with IRSA or Workload Identity credentials silently started
running without them. Exact-if-exists restores the pre-created ServiceAccount
as the one the pod runs as; what does not come back is aicr managing its
permissions.
Supported flow. Generate the RBAC manifests, read them, apply them, then take snapshots normally:
Step 2 does not check that the ServiceAccount exists — it contacts no
cluster at all, so it cannot. A mistyped name yields manifests naming a subject
that does not resolve; Kubernetes accepts such a binding and it simply grants
nothing. The generated 02-rolebinding.yaml tells you how to verify the name,
and aicr never creates a ServiceAccount for you.
What the manifests contain
aicr snapshot --add-roles-to-service-account <sa> writes a new directory in
your current working directory, one object per file:
Every file opens with a YAML comment header naming what the object grants and
why the agent needs each rule, so you can decide rule by rule whether to apply
it. The numeric prefixes exist because kubectl apply -f <dir>/ visits a
directory in lexical order — they keep each Role ahead of the binding that
references it. You can also apply or read the files individually.
The rules are the same ones a run-scoped grant carries, so a shared
ServiceAccount is never less capable than a run-owned one. The names end in
-rbac, which no run-scoped name can (a run-scoped name always ends in a run
ID whose last segment is hexadecimal), so the two name spaces cannot collide.
Nothing is applied for you, and nothing is removed for you. The objects you
apply carry no aicr.run/run-id label, no run enters them into its cleanup
list, and no aicr snapshot or aicr validate invocation deletes them.
Teardown is one command, which is why keeping the directory is worth it:
If you no longer have the directory, delete the objects by name instead:
The directory name carries a fresh run ID on every invocation, so generating
twice never overwrites a set you are still reviewing; a directory that already
exists fails the command with CONFLICT. After an aicr upgrade, re-generate
and kubectl apply -f the new directory to refresh the rules in place.
The cluster-scoped pair is named aicr-agent-<namespace>.<service-account>-rbac.
The two segments join on ., which a namespace can never contain, so no other
namespace and ServiceAccount combination can produce the same name — applying
one grant cannot retarget another’s.
Trade-off: per-run permission isolation is waived
Using an existing ServiceAccount is an opt-in exchange, and it is worth understanding before you choose it:
- Concurrent runs share one identity. Two
aicr snapshotinvocations using the same ServiceAccount hold exactly the same grants. Run isolation’s guarantee that one run’s permissions cannot reach another run’s does not apply to them. --discover-networkgrants become permanent. With a run-owned ServiceAccount, the cluster-scoped mutating discovery rules (nodes: patch,pods/exec: create, CRD / namespace / DaemonSet create-delete — see Security Considerations) exist for one run and are revoked at cleanup. Rendered withaicr snapshot --add-roles-to-service-account <sa> --discover-networkand applied, they sit on that ServiceAccount until you remove them.
Generate without --discover-network unless you need live network discovery;
that grant is read-only. When you do use it, 03-clusterrole.yaml carries a
warning header naming every mutating rule and the discovery step it exists for
— read it before applying, which is precisely what writing manifests instead of
applying them is for. If you need discovery only occasionally, prefer a
run-owned ServiceAccount for those runs, or keep a separate ServiceAccount used
only for discovery.
Post-Deployment
Retrieve Snapshot
Generate Recipe from Snapshot
Complete Workflow
Integration Patterns
CI/CD Pipeline
Multi-Cluster Auditing
Drift Detection
Troubleshooting
Job Fails to Start
Check RBAC permissions. The ServiceAccount name is run-scoped (aicr-<run-id>), so look it up first.
Select it by run ID, not by position: with concurrent snapshot runs the namespace
holds one ServiceAccount per run, and .items[0] would pick an arbitrary one — so
the checks below could report on a healthy run while you are debugging a failed one.
The CLI logs the run ID when it starts (snapshot agent run: runID=...), and the
Job, its pods, and its RBAC resources all carry it as the aicr.run/run-id
label. (The staging ConfigMap is the exception — it is written from inside the
pod and carries only app.kubernetes.io/name, app.kubernetes.io/component and
app.kubernetes.io/version; find it by its run-scoped name instead, see
Job Completes but No Output.)
Job Pending
Check node selectors and tolerations:
Job Completes but No Output
Check ConfigMap and container logs:
Permission Denied
A run that stops with missing required permissions never reached the cluster
— read the list it printed, which names every missing verb, its scope, and
whether you or the agent ServiceAccount lacked it. See
Pre-flight permission gate for the full matrix
and for why exact-ServiceAccount mode needs fewer of them.
If the run got past the gate, verify the RBAC it deployed:
Cleanup Left a Resource Behind
Cleanup removes only the objects the run itself created, and pins every delete to the UID the apiserver returned when it created them. If a create’s response is lost in flight, the run knows the name it used but never learned the object’s identity — so cleanup re-reads the object and deletes it only when it still carries that invocation’s own labels. When something else has taken the name over, cleanup keeps the object and says so:
The deciding label there is aicr.run/invocation-id, not aicr.run/run-id.
Every object a run creates carries both, but the run ID is yours to set — aicr validate gives one ID to its snapshot agent and its validation Jobs, and
pinned CI runs reuse one on purpose — so two commands can legitimately stamp
the same one. The invocation ID is generated inside the process for each
Deployer and is not configurable, so it is what separates “the object I
created” from “an object another invocation created at the same name”. Do not
select on it: its value changes every invocation.
Inspect the named object and delete it yourself if it is a stale leftover:
The staging ConfigMap is warned about the same way, but proved differently. It
is written by the in-pod agent rather than by the CLI, so it carries neither of
those labels (see Job Completes but No Output)
and the agent image may be a different aicr version than the CLI. Cleanup
therefore takes its evidence from the Job whose pod wrote it: it sweeps the
ConfigMap only when this run created that Job successfully AND that same Job —
matched by UID, not by name — is still the one standing in the cluster, and
only when the object looks like an aicr artifact
(app.kubernetes.io/name=aicr). If the Job was deleted or replaced first, the
ConfigMap is kept:
Delete it by hand once you have confirmed it is not a live run’s:
Security Considerations
RBAC Permissions
The agent requires these permissions (created automatically by the CLI):
- ClusterRole (
aicr-node-reader-<run-id>, run-scoped): Read access to nodes and pods;get/listaccess to ClusterPolicy CRDs (nvidia.com); cluster-widelistaccess to Slinky Controller, NodeSet, LoginSet, RestApi, and Accounting CRs (slinky.slurm.net); and cluster-widelistaccess to official MariaDB CRs (k8s.mariadb.com) - Role (
aicr-<run-id>, run-scoped): Create/update ConfigMaps and list pods in the deployment namespace
The baseline ClusterRole above is read-only (get/list only). Slinky
detection projects only allowlisted identity, association, and boolean fields;
it omits free-form configuration, status, pod templates, and Secret/ConfigMap
references or contents. MariaDB detection records only official API-group and
CR presence; it does not inspect database configuration, Services, operator
Deployments, pods, or external databases.
Additional privileges with --discover-network. When --discover-network
is set, the CLI appends a set of cluster-scoped mutating rules to the
ClusterRole so k8s-launch-kit’s live discovery can run. These grant far more
than read access:
apiextensions.k8s.ioCustomResourceDefinitions:get,list,create,update,patchnamespaces:get,create,delete(l8k creates and tears down a bootstrap namespace)apps/daemonsets:get,list,watch,create,deleteserviceaccounts,configmaps:get,create,deleterbac.authorization.k8s.ioroles, rolebindings:get,create,deletepods/exec:create(l8k exec’s into the discovery DaemonSet pods to read VPD / link state)nodes:patch(writesnvidia.kubernetes-launch-kit.*node labels)configuration.net.nvidia.comnicdevices:get,listmellanox.comnicclusterpolicies:get,patch
Use --discover-network only against clusters where this mutation and the
broader RBAC grant are acceptable.
Pre-flight permission gate
Every run begins by verifying the permissions it will actually use and stops
before writing anything if any are missing. The gate is read-only: an access
review is an authorization query that persists nothing, and the one object it
reads — the ServiceAccount named by --service-account-name — is read with
get. A failed gate therefore leaves the cluster exactly as it found it.
All failures are reported together, so one run tells you everything to fix. Each line names the verb, the resource, the scope, and which of the two identities lacked it:
What the caller must be able to do
Checked with SelfSubjectAccessReview against your kubeconfig identity. The
first group is required in both ServiceAccount modes:
serviceaccounts: get is not optional. Without it the run cannot tell whether
--service-account-name names an existing ServiceAccount or is a prefix, and
guessing “prefix” would run the agent under a generated ServiceAccount carrying
none of the named account’s IRSA or Workload Identity annotations. The run
stops instead of guessing.
The second group is required only in prefix mode — when the run creates its own run-scoped ServiceAccount:
delete is required alongside create because cleanup always runs, including
on the failure path. An identity that can create but not delete would leave a
ServiceAccount, Role, RoleBinding, ClusterRole and ClusterRoleBinding behind on
every single run.
Exact-ServiceAccount mode requires fewer caller permissions. When
--service-account-name names a ServiceAccount that already exists, the run
creates and deletes no RBAC at all, so none of the five kinds above is
demanded of you. What it requires instead is that the ServiceAccount was
actually provisioned — see below.
What the agent ServiceAccount must be able to do
The agent pod runs as a ServiceAccount, not as you, so its permissions are
checked separately with SubjectAccessReview naming
system:serviceaccount:<namespace>:<name> as the subject. The questions are
derived from the same rule set the run-scoped Role and ClusterRole grant
(and that --add-roles-to-service-account renders), so the gate cannot fall
behind what the agent needs: namespaced configmaps and pods access, plus
cluster-scoped nodes, pods, nvidia.com ClusterPolicies, the Slinky CRs
and the MariaDB CRs — widened to the full mutating set when
--discover-network is passed.
This check runs in exact-ServiceAccount mode only. In prefix mode the
ServiceAccount does not exist yet and the run is about to grant it exactly
those rules, so there is nothing to verify. In exact mode aicr grants nothing,
and the most common failure is that the manifests from
--add-roles-to-service-account were rendered but never applied. The gate
catches that up front and tells you how to fix it, rather than letting a Job
start and fail inside the pod minutes later.
When the ServiceAccount’s permissions cannot be verified. Creating a
SubjectAccessReview is itself a privilege. If you do not hold it, the run
does not silently skip the check — it says so and continues, because the
agent will still fail visibly in-pod if a rule is missing:
Pod Security Context
The agent requires elevated privileges to collect system configuration from the host:
hostPID,hostNetwork,hostIPC: Required to read host system configurationprivileged+SYS_ADMIN: Required to access GPU configuration and kernel parameters/run/systemdmount: Required to query systemd service states
See Also
- CLI Reference - aicr CLI commands
- Installation Guide - Install CLI locally
- API Reference - REST API usage
- Kubernetes Deployment - API server deployment