> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nvcf/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nvcf/_mcp/server.

# High Availability

This guide explains how to run the self-hosted NVCF control plane in a
high-availability (HA) topology so that it survives the loss of a single node
or availability zone (AZ). It covers the cluster prerequisites you must provide,
how to turn HA on, what each service does under HA, and how to validate the
result.

HA is on by default (`highAvailability.mode: preferred`). Single-node installs
(local, CI, small proofs of concept) should turn it off by setting
`highAvailability.mode: none`.

## Overview

HA is controlled by a single switch in your Helmfile environment values,
`highAvailability.mode`. There is no per-component HA configuration beneath
it: sizing and placement are derived uniformly from the mode for every
in-scope release, in `deploy/stacks/self-managed/global.yaml.gotmpl`.

| `highAvailability.mode` | Behavior |
| --- | --- |
| `none` | Keep existing single-replica chart/env values. Nothing in this guide applies. Set this for local, CI, and single-node installs. |
| `preferred` (default) | HA sizing (multiple replicas, quorum services at 3). Node/AZ spread is a **preference**: if a second node or AZ has no capacity, pods still schedule (co-located) rather than staying `Pending`. |
| `enforced` | Same HA sizing, but node/AZ spread is **required**: a replica that cannot land on a distinct node/AZ stays `Pending` instead of packing onto an occupied one. |

`preferred` is the recommended starting point: it gives you the full HA
topology while degrading gracefully on a constrained cluster. Move to
`enforced` once you have confirmed your node pools have capacity in every AZ
and you want hard placement guarantees.

For the replicated services, `preferred` guarantees continuity after a single
**pod** failure but provides only best-effort node/zone separation — under
capacity pressure, replicas can still end up co-located. `enforced` guarantees
continuity after a single pod **or node** failure when the prerequisites below
are met. `invocation-service` and `grpc-proxy` are still single-replica in both
modes, so they are briefly unavailable while their Pod is rescheduled. Neither mode alone
claims arbitrary availability-zone or site-loss tolerance; see
[Zone spread for the quorum services](#zone-spread-for-the-quorum-services)
and [Recovery objectives](#recovery-objectives-and-failure-behavior) for what
it takes to survive an AZ loss specifically.

## Cluster prerequisites

HA depends on infrastructure the **operator** provides. The stack cannot create
nodes or AZs for you — it only schedules against what you label.

### 1. Node count

Both HA modes require **at least 3 schedulable nodes** in the pool(s) that host
control-plane and quorum workloads for the *intended* spread to actually
schedule — the quorum services (Cassandra, NATS, OpenBao) run 3 replicas.
Under `preferred`, fewer nodes degrade to co-located (but still Ready) pods
rather than blocking the install; under `enforced`, insufficient nodes leave
replicas `Pending`.

### 2. Availability-zone labels

To spread replicas across failure domains, the scheduler uses the standard
Kubernetes well-known label on your nodes:

```text
topology.kubernetes.io/zone=<az>
```

You (the operator) must ensure this label is present on every node. Managed
Kubernetes services (EKS, AKS, GKE) apply it automatically. On bare-metal or
custom clusters, set it yourself, for example:

```bash
kubectl label node <node-name> topology.kubernetes.io/zone=az-1
```

If the label is absent, zone topology spread has nothing to spread across. In
`preferred` this silently degrades to node-level spread only; in
`enforced` zone-constrained pods can stay `Pending`. Aim for capacity in
**at least two AZs** (three is better for the quorum services — see
[Zone spread for the quorum services](#zone-spread-for-the-quorum-services)).

### 3. Dedicated node pools (recommended)

For predictable placement and isolation, give the stateful quorum services their
own node pools and label them with `nvcf.nvidia.com/workload`. Configure the
selectors under `global.nodeSelectors` and **set `enabled: true`** — the
selectors ship disabled (`global.nodeSelectors.enabled: false`) and are a no-op
until you turn them on:

```yaml
global:
  nodeSelectors:
    enabled: true                 # required; selectors are ignored when false
    controlplane:
      key: nvcf.nvidia.com/workload
      value: control-plane
    cassandra:
      key: nvcf.nvidia.com/workload
      value: cassandra
    vault:
      key: nvcf.nvidia.com/workload
      value: vault
```

| Pool | Selector value | Hosts |
| --- | --- | --- |
| `controlplane` | `control-plane` | All Tier-1 Deployments (api, invocation, grpcproxy, adminIssuerProxy, rateLimiter, natsAuthCalloutService, llmApiGateway, nats) |
| `cassandra` | `cassandra` | Cassandra StatefulSet |
| `vault` | `vault` | OpenBao StatefulSet |

These same three classes (`controlplane`/`cassandra`/`vault`) are also the
classes used by `global.affinity` and `global.topologySpreadConstraints`
below, so node-pool selection and scheduling-policy tuning stay consistent.

Each dedicated pool must span the availability zones — that is, have **capacity
in every AZ you want to spread across** (both zones in a 2-AZ cluster, all three
in a 3-AZ cluster). A 3-node Cassandra pool concentrated in one AZ cannot spread
across zones no matter what the stack requests, and with `enabled: false` the
pods fall back to default scheduling regardless of your labels.

`global.nodeSelectors.all` only applies to a class (`controlplane`/`cassandra`/`vault`)
whose own selector is unset. The stack's `base.yaml` ships `controlplane`/`cassandra`/`vault`
already set to their own dedicated-pool values, so setting `all` alone has no effect on
them, and nulling out the per-class selectors to force the fallback is not currently safe
(some releases read `controlplane` directly and require it to be non-null).

To run a single shared pool instead, give each class the same selector explicitly:

```yaml
global:
  nodeSelectors:
    enabled: true
    controlplane:
      key: nvcf.nvidia.com/workload
      value: system
    cassandra:
      key: nvcf.nvidia.com/workload
      value: system
    vault:
      key: nvcf.nvidia.com/workload
      value: system
```

Size the shared pool to hold every replica on distinct nodes/AZs.

## Enabling HA

Set the mode in your environment file (for example
`deploy/stacks/self-managed/environments/<env>.yaml`):

```yaml
highAvailability:
  mode: preferred
```

That single line activates all of the defaults documented below. Then apply
the stack as usual:

```bash
helmfile -e <env> apply
```

An invalid mode fails the render fast with a clear error, so a typo cannot
silently disable HA.

## What HA changes, by tier

### Replica-safe Deployments

Active-active Deployments with no leader election: `api`, `rateLimiter`,
`natsAuthCalloutService`, `adminIssuerProxy`, `llmApiGateway` (when the LLM
addon is enabled), and the control-plane services `nvctApi`, `notary`, `sis`,
`apiKeys`, `reval`, and `ess` all follow the same convention. (`nvctApi` and
`sis` run scheduled jobs but coordinate them with a distributed lock; `ess` runs
its scheduled crypto jobs only in a separate worker that is not part of this
stack, so the in-stack ESS API scales safely.)

> **Note — `invocation-service` and `grpc-proxy` are deferred.** These two are
> stateless too, but their multi-replica scaling is intentionally **held at a
> single replica for now**, pending Envoy support in the self-hosted stack.
> Worker callbacks are host-bound to the specific pod that accepted the request
> (per-pod pod-IP / DNS addressing), which is safe in a single cluster; the
> Envoy dependency is for the cross-cluster case. Until then they keep hostname
> anti-affinity and zone spread (no-ops at one replica) and get **no HA PDB**
> (a `minAvailable: 1` PDB on a singleton would block node drains).

Under HA each of these gets:

- **2 replicas.**
- **Hostname pod anti-affinity** — under `enforced` the two replicas can never
  share a node; under `preferred` the scheduler separates them when it can but
  may co-locate under capacity pressure (so `preferred` does not by itself
  guarantee node-loss survival).
- **Zone topology spread** (`topology.kubernetes.io/zone`, `maxSkew: 1`) —
  `enforced` requires different AZs (a replica stays Pending otherwise);
  `preferred` spreads when it can. Requires nodes labelled with the zone.
- **A PodDisruptionBudget** (`minAvailable: 1`) so voluntary disruptions
  (drains, upgrades) never take the last replica.
- **A surge rolling-update strategy** (`maxSurge: 1`, `maxUnavailable: 0`) so a
  new pod is Ready before an old one is removed.

Affinity and topology spread are shared scheduling policy, so they can be
tuned once for every release of a class instead of per-component. See
[Tuning scheduling policy](#tuning-scheduling-policy-globalaffinity--globaltopologyspreadconstraints)
below.

### Quorum services (data durability)

`Cassandra`, `NATS`, and `OpenBao` run as **3-replica quorum StatefulSets** with:

- **Hostname pod anti-affinity** to keep the 3 peers on distinct nodes.
  Following the mode, this is preferred (soft) under `preferred` — the scheduler
  may co-locate under capacity pressure — and required (hard) under `enforced`,
  the same convention as replica-safe Deployments.
  OpenBao's upstream chart ships a hard anti-affinity that the stack disables
  for single-node installs and re-enables (soft/hard by mode) under HA.
- **Zone topology spread** across `topology.kubernetes.io/zone`, same
  soft/hard-by-mode convention. See the next section for what it takes for
  this to actually protect against an AZ loss.
- **PodDisruptionBudgets** sized to tolerate exactly one voluntary disruption
  (`maxUnavailable: 1`), so a drain takes down at most one member even when
  you run more than three replicas.

#### Zone spread for the quorum services

Zone topology spread for the quorum peers is part of the same mode
convention as everything else — it is **not** a separate toggle. That said,
whether it actually protects you from an AZ loss depends on infrastructure
you must provide:

- **StorageClass `volumeBindingMode: WaitForFirstConsumer`.** Each quorum pod
  has a zonal PersistentVolume, and a zonal disk can only attach to a node in
  its own AZ. With `WaitForFirstConsumer`, the scheduler places the pod first
  (honoring the spread constraint) and the PV is then created in that pod's
  zone. With `Immediate` binding the PV's zone is chosen up front and the pod
  is pinned to it, which fights the spread constraint and can leave pods
  `Pending`. The stack cannot set this for you — it is a property of the
  StorageClass you supply.
- **Capacity in at least 3 AZs.** A 3-member quorum only survives an AZ loss
  if no single AZ holds a majority. With only 2 AZs one zone inevitably holds
  2 of 3 members, and losing that zone breaks quorum.

If you do not have 3-AZ capacity or a `WaitForFirstConsumer` StorageClass,
zone spread for the quorum tier will not reliably schedule or will not
protect you from an AZ loss even though it renders. Switching from `enforced`
to `preferred` alone does **not** fix this — `preferred`'s soft
(`ScheduleAnyway`) constraint will let pods co-locate in that case rather than
failing loudly, which hides the gap rather than closing it. To disable
generated zone spread for the quorum tier explicitly instead of relying on a
mode change, use the `global.topologySpreadConstraints` escape hatch:

```yaml
global:
  topologySpreadConstraints:
    cassandra: []
    vault: []
```

(NATS shares the `controlplane` class with the stateless tier; scope an
override to just NATS via a component-shaped value instead if you need to
disable spread for NATS specifically without affecting the rest of
`controlplane`.)

**Cassandra places data by rack, not by zone.** Spreading the *pods* across
zones does not by itself make the *data* zone-diverse. `NetworkTopologyStrategy`
puts each replica on a different **rack**, and the Cassandra image assigns the
rack from the pod ordinal (`r1`, `r2`, `r3` for ordinals 0, 1, 2, repeating as
ordinal mod 3), not from the node's zone. With exactly three pods spread one
per zone, the three racks land in three zones, so RF=3 keeps one copy of each
row in every zone. With more than three pods, pods that share a rack can sit in
different zones, so replica placement is not zone-aware; zone-aware rack
assignment is not implemented yet. NATS and OpenBao (Raft) replicate per member
and need only the pod spread.

Once a quorum pod's PV is created in a zone it is pinned there for the life of
that StatefulSet ordinal — steady-state placement stays spread, but a pod whose
AZ is lost cannot reschedule elsewhere until the AZ returns (its two peers carry
quorum in the meantime).

Beyond placement, HA also raises the data-durability settings:

#### NATS JetStream replica factor

Streams default to a single replica. When NATS runs more than one server, the
stack derives the JetStream replica factor (RF) from the server count, capped
at **3** as the NATS documentation recommends, so the HA cluster of 3 servers
gets RF=3. It is set on the three services that create streams: `nvcf-api` (via
`NVCF_NATS_REPLICAS`), `invocation-service` (via `NATS_PROPERTIES__REPLICAS`), and
`icms` (the `sis` release, via `ICMS_NATS_REPLICAS`, for `CreateNvcaFunctionTaskStream` and
`TerminateNvcaStream`). To override it for nvcf-api or invocation-service, set
that variable in `api.env` or `invocation.env`. JetStream streams use Raft quorum: RF=3 tolerates the loss of
one replica. **RF=2 is not sufficient** — a 2-member Raft group loses quorum
the moment either replica is unavailable, so it provides no resilience benefit
over RF=1.

RF applies when a stream is **created**. Streams already created at RF=1 are not
rewritten by changing this value — after enabling HA, recreate or edit those
streams (for example `nats stream edit`) to raise their replica factor. The
exception is `icms`, which raises its two streams to the configured RF when it
starts and on its periodic stream validation.

#### Cassandra replication and consistency

The application keyspaces use `NetworkTopologyStrategy` in the `ncp`
datacenter. Their replication factor is set once, when a keyspace is **first
created**, from the Cassandra replica count, so a **fresh** HA install creates
them at RF=3. An existing install keeps the replication factor it was created
with (RF=1 for a single-node install) until you raise it; see
[Upgrading an existing install](#upgrading-an-existing-install).

The control-plane services default to `LOCAL_QUORUM` consistency, and some
`nvcf-api` read paths use `LOCAL_ONE`. With RF=3:

- **Single datacenter:** `LOCAL_QUORUM` tolerates the loss of one replica for
  reads and writes. `LOCAL_ONE` reads stay available too, but can briefly
  return data that has not reached every replica yet.
- **Multi-AZ:** with exactly three pods spread one per zone, the rack mapping
  above puts one replica in each zone, so `LOCAL_QUORUM` keeps reads and
  writes available through the loss of one zone. This needs three zones (see
  [Zone spread for the quorum services](#zone-spread-for-the-quorum-services))
  and does not hold above three pods.

## Tuning scheduling policy (`global.affinity` / `global.topologySpreadConstraints`)

Affinity and zone topology spread are shared scheduling policy, so you can
tune them once per workload class instead of overriding every component
individually. Both resolve in the same order:

```text
class-specific global value  ->  global.<key>.all  ->  the highAvailability.mode-derived convention
```

The available classes match `global.nodeSelectors`: `controlplane`,
`cassandra`, `vault`. These values are optional — absence means "use the
mode-derived convention" — and an explicit override replaces the generated
policy for that class entirely, it does not merge with it:

```yaml
highAvailability:
  mode: preferred

global:
  affinity:
    cassandra:
      podAntiAffinity:
        requiredDuringSchedulingIgnoredDuringExecution:
          - labelSelector:
              matchLabels:
                app.kubernetes.io/instance: cassandra
            topologyKey: kubernetes.io/hostname
  topologySpreadConstraints:
    all: []   # disable generated zone spread everywhere except explicit class overrides
```

An explicit `{}` (for `affinity`) or `[]` (for `topologySpreadConstraints`)
suppresses the generated policy for that class — this is different from
omitting the key, which falls back to `all` and then to the convention.

Component-shaped values (for example `api.affinity`, `cassandra.affinity`) in
your environment file remain the final, single-release escape hatch and take
precedence over both `global.affinity` and the generated convention.

## Validation

After applying HA, confirm replicas are spread as expected.

Check that quorum peers landed on distinct nodes and AZs:

```bash
# Nodes and their AZ labels
kubectl get nodes -L topology.kubernetes.io/zone

# Cassandra / NATS / OpenBao pods with their nodes
kubectl -n cassandra-system get pods -o wide
kubectl -n nats-system get pods -o wide
kubectl -n vault-system get pods -o wide
```

Confirm the replica-safe Deployments scaled and spread:

```bash
kubectl -n nvcf get deploy nvcf-api -o wide
kubectl -n nvcf get pods -o wide -l 'app.kubernetes.io/instance=api,!app.kubernetes.io/component'
kubectl -n api-keys get deploy admin-token-issuer-proxy -o wide
# invocation-service and grpc-proxy stay at 1 replica for now (deferred until Envoy)
kubectl -n nvcf get deploy invocation-service grpc-proxy-deployment -o wide
# ess-api scales to 2 (its scheduled crypto jobs run only in a separate worker, not in-stack)
kubectl -n ess get deploy ess-api-deployment -o wide
```

Confirm the JetStream RF took effect (streams report `Replicas: 3`):

```bash
kubectl -n nats-system exec -it nats-0 -- nats stream ls
kubectl -n nats-system exec -it nats-0 -- nats stream info <stream-name>
```

Confirm Cassandra keyspace replication:

```bash
read -rs -p "Cassandra administrator password: " CASSANDRA_PASSWORD; echo
kubectl -n cassandra-system exec cassandra-0 -- \
  cqlsh -u cassandra -p "$CASSANDRA_PASSWORD" \
  -e "SELECT keyspace_name, replication FROM system_schema.keyspaces;"
```

The application keyspaces should report `'ncp': '3'`. On an install that
predates HA they keep their original value until you follow step 5 of
[Upgrading an existing install](#upgrading-an-existing-install).

If any pod is stuck `Pending` under `enforced`, it usually means a node pool
lacks capacity in a second node/AZ. Add capacity, or drop to `preferred` to
let it schedule while you rebalance — but see
[Zone spread for the quorum services](#zone-spread-for-the-quorum-services)
for why that alone does not restore the AZ-loss guarantee if the underlying
capacity gap remains.

## Recovery objectives and failure behavior

With HA enabled and capacity in at least two AZs:

- **Single node loss:** Multi-replica replica-safe Deployments keep
  serving from their surviving replica; the scheduler recreates the lost pod on
  another node (and the PDB prevents drains from removing the last one).
  `invocation-service` and `grpc-proxy` (single replica until Envoy) are briefly
  unavailable while the scheduler restarts the pod on another node. Quorum
  services (Cassandra RF=3/`LOCAL_QUORUM`, NATS RF=3, OpenBao 3-node Raft)
  retain quorum with 2 of 3 members and continue serving reads and writes.
  For Cassandra this assumes the keyspaces are at RF=3, which an upgraded
  install reaches only after
  [step 5 of the upgrade](#upgrading-an-existing-install).
- **Single AZ loss:** Only if zone spread actually took effect — see
  [Zone spread for the quorum services](#zone-spread-for-the-quorum-services)
  for its prerequisites (3-AZ capacity, `WaitForFirstConsumer` storage). With
  those met, the control plane stays available on the surviving AZ(s) and
  recovery time is dominated by pod reschedule/restart time rather than any
  manual failover. Without them, an AZ loss can take a majority of a quorum
  service's members and pause writes even with HA enabled.
- **Two simultaneous quorum-member losses:** A 3-member quorum service loses
  quorum and pauses writes until a member returns. This is why three AZs (or at
  least three nodes across two AZs, with the third member able to reschedule) is
  the durable target.

HA reduces recovery to automatic rescheduling within surviving failure domains;
it does not replace backups. Continue to back up Cassandra and OpenBao per the
[Control Plane Operations](/nvcf/dev/self-managed/control-plane-operations) runbooks.

## Upgrading an existing install

HA ships in a specific self-managed stack release. On a running install, adopt it
in this order:

1. **Upgrade to the release that includes HA.** Check out the stack release that
   carries these changes and use its `deploy/stacks/self-managed/` directory. Your
   `environments/<environment-name>.yaml` and `secrets/` carry over unchanged.
2. **Make the new artifacts available.** If you pull directly from NGC
   (`nvcr.io` / `helm.ngc.nvidia.com`), no action is needed. If you mirror to a
   private registry, re-mirror the stack's charts and images **before** syncing —
   otherwise pods fail with `ImagePullBackOff`. See
   [Image Mirroring](/nvcf/dev/manifest/image-mirroring) and the [Artifact Manifest](/nvcf/dev/manifest/artifacts).
3. **Choose the mode.** HA is on by default (`preferred`), so an environment
   file that does not set `highAvailability.mode` **activates HA on upgrade**.
   Confirm the [cluster prerequisites](#cluster-prerequisites) first — `preferred`
   needs ≥3 schedulable nodes; `enforced` also needs ≥3 AZs with capacity, AZ
   labels, and `WaitForFirstConsumer` storage. If the cluster is not ready (or it
   is a single-node install), set `highAvailability.mode: none` to keep the
   current single-replica behavior.
4. **Apply.** `HELMFILE_ENV=<environment-name> helmfile sync`. Replica-safe
   Deployments roll with `maxUnavailable: 0` / `maxSurge: 1` (needs headroom for
   one surge pod); quorum services roll one member at a time and keep quorum.
5. **Raise Cassandra keyspace replication.** The sync scales Cassandra to three
   members, but the application keyspaces keep the replication factor they were
   created with, so a single Cassandra node loss can still make data
   unavailable. Once all three Cassandra pods are Ready, raise each application
   keyspace to RF=3, then run a full repair on every member so the new replicas
   receive the existing data. Skip any keyspace your install does not have
   (the query in [Validation](#validation) lists them):

   ```bash
   read -rs -p "Cassandra administrator password: " CASSANDRA_PASSWORD; echo
   for ks in nvcf_api nvct_api ess_api api_keys_api sis_api event_ledger nvcf_autoscaler; do
     kubectl -n cassandra-system exec cassandra-0 -- \
       cqlsh -u cassandra -p "$CASSANDRA_PASSWORD" \
       -e "ALTER KEYSPACE $ks WITH replication = {'class': 'NetworkTopologyStrategy', 'ncp': '3'};"
   done
   for pod in cassandra-0 cassandra-1 cassandra-2; do
     kubectl -n cassandra-system exec "$pod" -- nodetool repair --full
   done
   ```

   Run the repairs one member at a time. Until they finish, reads at
   `LOCAL_ONE` can miss rows that exist only on the original replica.
6. **Verify.** Confirm the [Validation](#validation) checks — replica counts
   (2 or 3), PodDisruptionBudgets, zone spread, and keyspace replication.

> Because `preferred` is the default, upgrading with an environment file that
> does not pin `highAvailability.mode` will scale services up (2/3 replicas,
> PDBs, placement) on the first `sync`. Set `mode: none` beforehand if you are
> not ready for HA.

## Related

- [Control Plane Operations](/nvcf/dev/self-managed/control-plane-operations) — service reference,
  key rotation, and upgrade runbooks.
- [Infrastructure Sizing](/nvcf/dev/overview/infrastructure-sizing) — node pool sizing
  guidance.
- [Helmfile Installation](/nvcf/dev/self-managed/helmfile-installation) — how environment values
  and `global.yaml.gotmpl` are applied.