High Availability
This guide explains how to run the self-hosted NVCF control plane in a high-availability (HA) topology so that it survives the loss of a single node or availability zone (AZ). It covers the cluster prerequisites you must provide, how to turn HA on, what each service does under HA, and how to validate the result.
HA is on by default (highAvailability.mode: preferred). Single-node installs
(local, CI, small proofs of concept) should turn it off by setting
highAvailability.mode: none.
Overview
HA is controlled by a single switch in your Helmfile environment values,
highAvailability.mode. There is no per-component HA configuration beneath
it: sizing and placement are derived uniformly from the mode for every
in-scope release, in deploy/stacks/self-managed/global.yaml.gotmpl.
preferred is the recommended starting point: it gives you the full HA
topology while degrading gracefully on a constrained cluster. Move to
enforced once you have confirmed your node pools have capacity in every AZ
and you want hard placement guarantees.
For the replicated services, preferred guarantees continuity after a single
pod failure but provides only best-effort node/zone separation — under
capacity pressure, replicas can still end up co-located. enforced guarantees
continuity after a single pod or node failure when the prerequisites below
are met. invocation-service and grpc-proxy are still single-replica in both
modes, so they are briefly unavailable while their Pod is rescheduled. Neither mode alone
claims arbitrary availability-zone or site-loss tolerance; see
Zone spread for the quorum services
and Recovery objectives for what
it takes to survive an AZ loss specifically.
Cluster prerequisites
HA depends on infrastructure the operator provides. The stack cannot create nodes or AZs for you — it only schedules against what you label.
1. Node count
Both HA modes require at least 3 schedulable nodes in the pool(s) that host
control-plane and quorum workloads for the intended spread to actually
schedule — the quorum services (Cassandra, NATS, OpenBao) run 3 replicas.
Under preferred, fewer nodes degrade to co-located (but still Ready) pods
rather than blocking the install; under enforced, insufficient nodes leave
replicas Pending.
2. Availability-zone labels
To spread replicas across failure domains, the scheduler uses the standard Kubernetes well-known label on your nodes:
You (the operator) must ensure this label is present on every node. Managed Kubernetes services (EKS, AKS, GKE) apply it automatically. On bare-metal or custom clusters, set it yourself, for example:
If the label is absent, zone topology spread has nothing to spread across. In
preferred this silently degrades to node-level spread only; in
enforced zone-constrained pods can stay Pending. Aim for capacity in
at least two AZs (three is better for the quorum services — see
Zone spread for the quorum services).
3. Dedicated node pools (recommended)
For predictable placement and isolation, give the stateful quorum services their
own node pools and label them with nvcf.nvidia.com/workload. Configure the
selectors under global.nodeSelectors and set enabled: true — the
selectors ship disabled (global.nodeSelectors.enabled: false) and are a no-op
until you turn them on:
These same three classes (controlplane/cassandra/vault) are also the
classes used by global.affinity and global.topologySpreadConstraints
below, so node-pool selection and scheduling-policy tuning stay consistent.
Each dedicated pool must span the availability zones — that is, have capacity
in every AZ you want to spread across (both zones in a 2-AZ cluster, all three
in a 3-AZ cluster). A 3-node Cassandra pool concentrated in one AZ cannot spread
across zones no matter what the stack requests, and with enabled: false the
pods fall back to default scheduling regardless of your labels.
global.nodeSelectors.all only applies to a class (controlplane/cassandra/vault)
whose own selector is unset. The stack’s base.yaml ships controlplane/cassandra/vault
already set to their own dedicated-pool values, so setting all alone has no effect on
them, and nulling out the per-class selectors to force the fallback is not currently safe
(some releases read controlplane directly and require it to be non-null).
To run a single shared pool instead, give each class the same selector explicitly:
Size the shared pool to hold every replica on distinct nodes/AZs.
Enabling HA
Set the mode in your environment file (for example
deploy/stacks/self-managed/environments/<env>.yaml):
That single line activates all of the defaults documented below. Then apply the stack as usual:
An invalid mode fails the render fast with a clear error, so a typo cannot silently disable HA.
What HA changes, by tier
Replica-safe Deployments
Active-active Deployments with no leader election: api, rateLimiter,
natsAuthCalloutService, adminIssuerProxy, llmApiGateway (when the LLM
addon is enabled), and the control-plane services nvctApi, notary, sis,
apiKeys, reval, and ess all follow the same convention. (nvctApi and
sis run scheduled jobs but coordinate them with a distributed lock; ess runs
its scheduled crypto jobs only in a separate worker that is not part of this
stack, so the in-stack ESS API scales safely.)
Note —
invocation-serviceandgrpc-proxyare deferred. These two are stateless too, but their multi-replica scaling is intentionally held at a single replica for now, pending Envoy support in the self-hosted stack. Worker callbacks are host-bound to the specific pod that accepted the request (per-pod pod-IP / DNS addressing), which is safe in a single cluster; the Envoy dependency is for the cross-cluster case. Until then they keep hostname anti-affinity and zone spread (no-ops at one replica) and get no HA PDB (aminAvailable: 1PDB on a singleton would block node drains).
Under HA each of these gets:
- 2 replicas.
- Hostname pod anti-affinity — under
enforcedthe two replicas can never share a node; underpreferredthe scheduler separates them when it can but may co-locate under capacity pressure (sopreferreddoes not by itself guarantee node-loss survival). - Zone topology spread (
topology.kubernetes.io/zone,maxSkew: 1) —enforcedrequires different AZs (a replica stays Pending otherwise);preferredspreads when it can. Requires nodes labelled with the zone. - A PodDisruptionBudget (
minAvailable: 1) so voluntary disruptions (drains, upgrades) never take the last replica. - A surge rolling-update strategy (
maxSurge: 1,maxUnavailable: 0) so a new pod is Ready before an old one is removed.
Affinity and topology spread are shared scheduling policy, so they can be tuned once for every release of a class instead of per-component. See Tuning scheduling policy below.
Quorum services (data durability)
Cassandra, NATS, and OpenBao run as 3-replica quorum StatefulSets with:
- Hostname pod anti-affinity to keep the 3 peers on distinct nodes.
Following the mode, this is preferred (soft) under
preferred— the scheduler may co-locate under capacity pressure — and required (hard) underenforced, the same convention as replica-safe Deployments. OpenBao’s upstream chart ships a hard anti-affinity that the stack disables for single-node installs and re-enables (soft/hard by mode) under HA. - Zone topology spread across
topology.kubernetes.io/zone, same soft/hard-by-mode convention. See the next section for what it takes for this to actually protect against an AZ loss. - PodDisruptionBudgets sized to tolerate exactly one voluntary disruption
(
maxUnavailable: 1), so a drain takes down at most one member even when you run more than three replicas.
Zone spread for the quorum services
Zone topology spread for the quorum peers is part of the same mode convention as everything else — it is not a separate toggle. That said, whether it actually protects you from an AZ loss depends on infrastructure you must provide:
- StorageClass
volumeBindingMode: WaitForFirstConsumer. Each quorum pod has a zonal PersistentVolume, and a zonal disk can only attach to a node in its own AZ. WithWaitForFirstConsumer, the scheduler places the pod first (honoring the spread constraint) and the PV is then created in that pod’s zone. WithImmediatebinding the PV’s zone is chosen up front and the pod is pinned to it, which fights the spread constraint and can leave podsPending. The stack cannot set this for you — it is a property of the StorageClass you supply. - Capacity in at least 3 AZs. A 3-member quorum only survives an AZ loss if no single AZ holds a majority. With only 2 AZs one zone inevitably holds 2 of 3 members, and losing that zone breaks quorum.
If you do not have 3-AZ capacity or a WaitForFirstConsumer StorageClass,
zone spread for the quorum tier will not reliably schedule or will not
protect you from an AZ loss even though it renders. Switching from enforced
to preferred alone does not fix this — preferred’s soft
(ScheduleAnyway) constraint will let pods co-locate in that case rather than
failing loudly, which hides the gap rather than closing it. To disable
generated zone spread for the quorum tier explicitly instead of relying on a
mode change, use the global.topologySpreadConstraints escape hatch:
(NATS shares the controlplane class with the stateless tier; scope an
override to just NATS via a component-shaped value instead if you need to
disable spread for NATS specifically without affecting the rest of
controlplane.)
Cassandra places data by rack, not by zone. Spreading the pods across
zones does not by itself make the data zone-diverse. NetworkTopologyStrategy
puts each replica on a different rack, and the Cassandra image assigns the
rack from the pod ordinal (r1, r2, r3 for ordinals 0, 1, 2, repeating as
ordinal mod 3), not from the node’s zone. With exactly three pods spread one
per zone, the three racks land in three zones, so RF=3 keeps one copy of each
row in every zone. With more than three pods, pods that share a rack can sit in
different zones, so replica placement is not zone-aware; zone-aware rack
assignment is not implemented yet. NATS and OpenBao (Raft) replicate per member
and need only the pod spread.
Once a quorum pod’s PV is created in a zone it is pinned there for the life of that StatefulSet ordinal — steady-state placement stays spread, but a pod whose AZ is lost cannot reschedule elsewhere until the AZ returns (its two peers carry quorum in the meantime).
Beyond placement, HA also raises the data-durability settings:
NATS JetStream replica factor
Streams default to a single replica. When NATS runs more than one server, the
stack derives the JetStream replica factor (RF) from the server count, capped
at 3 as the NATS documentation recommends, so the HA cluster of 3 servers
gets RF=3. It is set on the three services that create streams: nvcf-api (via
NVCF_NATS_REPLICAS), invocation-service (via NATS_PROPERTIES__REPLICAS), and
icms (the sis release, via ICMS_NATS_REPLICAS, for CreateNvcaFunctionTaskStream and
TerminateNvcaStream). To override it for nvcf-api or invocation-service, set
that variable in api.env or invocation.env. JetStream streams use Raft quorum: RF=3 tolerates the loss of
one replica. RF=2 is not sufficient — a 2-member Raft group loses quorum
the moment either replica is unavailable, so it provides no resilience benefit
over RF=1.
RF applies when a stream is created. Streams already created at RF=1 are not
rewritten by changing this value — after enabling HA, recreate or edit those
streams (for example nats stream edit) to raise their replica factor. The
exception is icms, which raises its two streams to the configured RF when it
starts and on its periodic stream validation.
Cassandra replication and consistency
The application keyspaces use NetworkTopologyStrategy in the ncp
datacenter. Their replication factor is set once, when a keyspace is first
created, from the Cassandra replica count, so a fresh HA install creates
them at RF=3. An existing install keeps the replication factor it was created
with (RF=1 for a single-node install) until you raise it; see
Upgrading an existing install.
The control-plane services default to LOCAL_QUORUM consistency, and some
nvcf-api read paths use LOCAL_ONE. With RF=3:
- Single datacenter:
LOCAL_QUORUMtolerates the loss of one replica for reads and writes.LOCAL_ONEreads stay available too, but can briefly return data that has not reached every replica yet. - Multi-AZ: with exactly three pods spread one per zone, the rack mapping
above puts one replica in each zone, so
LOCAL_QUORUMkeeps reads and writes available through the loss of one zone. This needs three zones (see Zone spread for the quorum services) and does not hold above three pods.
Tuning scheduling policy (global.affinity / global.topologySpreadConstraints)
Affinity and zone topology spread are shared scheduling policy, so you can tune them once per workload class instead of overriding every component individually. Both resolve in the same order:
The available classes match global.nodeSelectors: controlplane,
cassandra, vault. These values are optional — absence means “use the
mode-derived convention” — and an explicit override replaces the generated
policy for that class entirely, it does not merge with it:
An explicit {} (for affinity) or [] (for topologySpreadConstraints)
suppresses the generated policy for that class — this is different from
omitting the key, which falls back to all and then to the convention.
Component-shaped values (for example api.affinity, cassandra.affinity) in
your environment file remain the final, single-release escape hatch and take
precedence over both global.affinity and the generated convention.
Validation
After applying HA, confirm replicas are spread as expected.
Check that quorum peers landed on distinct nodes and AZs:
Confirm the replica-safe Deployments scaled and spread:
Confirm the JetStream RF took effect (streams report Replicas: 3):
Confirm Cassandra keyspace replication:
The application keyspaces should report 'ncp': '3'. On an install that
predates HA they keep their original value until you follow step 5 of
Upgrading an existing install.
If any pod is stuck Pending under enforced, it usually means a node pool
lacks capacity in a second node/AZ. Add capacity, or drop to preferred to
let it schedule while you rebalance — but see
Zone spread for the quorum services
for why that alone does not restore the AZ-loss guarantee if the underlying
capacity gap remains.
Recovery objectives and failure behavior
With HA enabled and capacity in at least two AZs:
- Single node loss: Multi-replica replica-safe Deployments keep
serving from their surviving replica; the scheduler recreates the lost pod on
another node (and the PDB prevents drains from removing the last one).
invocation-serviceandgrpc-proxy(single replica until Envoy) are briefly unavailable while the scheduler restarts the pod on another node. Quorum services (Cassandra RF=3/LOCAL_QUORUM, NATS RF=3, OpenBao 3-node Raft) retain quorum with 2 of 3 members and continue serving reads and writes. For Cassandra this assumes the keyspaces are at RF=3, which an upgraded install reaches only after step 5 of the upgrade. - Single AZ loss: Only if zone spread actually took effect — see
Zone spread for the quorum services
for its prerequisites (3-AZ capacity,
WaitForFirstConsumerstorage). With those met, the control plane stays available on the surviving AZ(s) and recovery time is dominated by pod reschedule/restart time rather than any manual failover. Without them, an AZ loss can take a majority of a quorum service’s members and pause writes even with HA enabled. - Two simultaneous quorum-member losses: A 3-member quorum service loses quorum and pauses writes until a member returns. This is why three AZs (or at least three nodes across two AZs, with the third member able to reschedule) is the durable target.
HA reduces recovery to automatic rescheduling within surviving failure domains; it does not replace backups. Continue to back up Cassandra and OpenBao per the Control Plane Operations runbooks.
Upgrading an existing install
HA ships in a specific self-managed stack release. On a running install, adopt it in this order:
-
Upgrade to the release that includes HA. Check out the stack release that carries these changes and use its
deploy/stacks/self-managed/directory. Yourenvironments/<environment-name>.yamlandsecrets/carry over unchanged. -
Make the new artifacts available. If you pull directly from NGC (
nvcr.io/helm.ngc.nvidia.com), no action is needed. If you mirror to a private registry, re-mirror the stack’s charts and images before syncing — otherwise pods fail withImagePullBackOff. See Image Mirroring and the Artifact Manifest. -
Choose the mode. HA is on by default (
preferred), so an environment file that does not sethighAvailability.modeactivates HA on upgrade. Confirm the cluster prerequisites first —preferredneeds ≥3 schedulable nodes;enforcedalso needs ≥3 AZs with capacity, AZ labels, andWaitForFirstConsumerstorage. If the cluster is not ready (or it is a single-node install), sethighAvailability.mode: noneto keep the current single-replica behavior. -
Apply.
HELMFILE_ENV=<environment-name> helmfile sync. Replica-safe Deployments roll withmaxUnavailable: 0/maxSurge: 1(needs headroom for one surge pod); quorum services roll one member at a time and keep quorum. -
Raise Cassandra keyspace replication. The sync scales Cassandra to three members, but the application keyspaces keep the replication factor they were created with, so a single Cassandra node loss can still make data unavailable. Once all three Cassandra pods are Ready, raise each application keyspace to RF=3, then run a full repair on every member so the new replicas receive the existing data. Skip any keyspace your install does not have (the query in Validation lists them):
Run the repairs one member at a time. Until they finish, reads at
LOCAL_ONEcan miss rows that exist only on the original replica. -
Verify. Confirm the Validation checks — replica counts (2 or 3), PodDisruptionBudgets, zone spread, and keyspace replication.
Because
preferredis the default, upgrading with an environment file that does not pinhighAvailability.modewill scale services up (2/3 replicas, PDBs, placement) on the firstsync. Setmode: nonebeforehand if you are not ready for HA.
Related
- Control Plane Operations — service reference, key rotation, and upgrade runbooks.
- Infrastructure Sizing — node pool sizing guidance.
- Helmfile Installation — how environment values
and
global.yaml.gotmplare applied.