0.6.0 to 0.6.1 Upgrade

View as Markdown

The 0.6.1 release replaces the archived Bitnami Cassandra runtime with the official Apache Cassandra 5.0.8 image and an in-house Cassandra Helm chart. This removes the CVE-affected Bitnami runtime and its bundled components.

Fresh installs need no special handling. Existing clusters that already hold data need a migration, because three things change at once: the runtime image, the Helm chart shape, and the on-disk data layout. A normal in-place Helm upgrade does not carry an existing Bitnami-based cluster onto the new stack:

  • The Cassandra StatefulSet changes shape (name, selector, volume claim template). Kubernetes freezes most StatefulSet fields, so the StatefulSet cannot be updated in place; it has to be recreated.
  • The Bitnami image stored data nested one directory deeper than the official image expects (/bitnami/cassandra/data/... versus /var/lib/cassandra/...), so the new image does not find the old data by default.
  • The old and new node configuration must agree on the settings the existing data was written with.

Because of this, an operator with an existing data-bearing cluster chooses a migration method. This guide describes the methods and their trade-offs; none is mandated over the others.

These steps assume 0.6.0 is already deployed and healthy, that you have downloaded the 0.6.1 stack, and that you run make from the extracted nvcf-self-managed-stack/ directory with your environment set.

$export HELMFILE_ENV=<your-environment> # for example eks-cds-qa

Prepare the 0.6.1 configuration

Return to the extracted 0.6.1 stack before running any 0.6.1 sync command. Reconcile the site-specific settings from the 0.6.0 deployment into the 0.6.1 configuration files:

  • environments/$HELMFILE_ENV.yaml
  • secrets/$HELMFILE_ENV-secrets.yaml

Start from the files or templates included with 0.6.1 and carry forward the required values from 0.6.0. Do not replace the 0.6.1 files wholesale because available settings can change between releases. Preserve the existing registry, storage, endpoint, credential, and secret values unless this procedure explicitly instructs you to change them.

Confirm that both files exist before continuing:

$for file in \
> "environments/${HELMFILE_ENV}.yaml" \
> "secrets/${HELMFILE_ENV}-secrets.yaml"; do
$ [ -f "$file" ] || { echo "missing required file: $file"; exit 1; }
$done

Fresh installs

A new 0.6.1 install has no existing Cassandra data and needs no migration. Install the stack normally; Cassandra comes up on the official image with the in-house chart:

$make install HELMFILE_ENV="$HELMFILE_ENV"

The rest of this guide applies only when migrating an existing data-bearing cluster.

Check Cassandra before migrating

Confirm the current Cassandra deployment is healthy and record its state:

$kubectl -n cassandra-system get pods,pvc
$kubectl -n cassandra-system exec cassandra-0 -c cassandra -- nodetool status

All Cassandra pods should be Running and report UN (Up/Normal) in nodetool status, and the data PersistentVolumeClaims should be Bound. Record the keyspaces and a row count for a known table so you can verify data after the migration:

$kubectl -n cassandra-system exec cassandra-0 -c cassandra -- \
> cqlsh -u <user> -p <password> -e "DESCRIBE KEYSPACES;"

Take a backup before migrating. On every node, run nodetool flush and then nodetool snapshot, and copy the snapshots off-node. Do not delete Cassandra PersistentVolumeClaims during any migration method.

Choose a migration method

Three methods are supported. They are listed in no particular order; choose the one that fits your deployment and your tolerance for downtime and operational steps.

MethodDowntimeData handlingRequires
Datacenter expansionEffectively noneNew nodes stream the data; old volumes untouchedCapacity to run both datacenters at once, datacenter-aware clients
In-place volume adoptionBrief restart per nodeThe new pods reuse the existing volumesA manual StatefulSet recreate; a subPath or a one-time layout relocation
Backup and restoreMaintenance windowSnapshot the old cluster, load into a fresh oneCapacity for a parallel cluster during the copy

Datacenter expansion keeps the cluster online by adding the new stack as a second Cassandra datacenter and streaming data to it, then cutting clients over. It has the most moving parts and needs enough capacity to run both datacenters during the migration.

In-place volume adoption reuses the existing data volumes. It is the smallest change in footprint, but it requires deleting and recreating the Cassandra StatefulSet (its volumes are retained) and pointing the new pods at the existing Bitnami on-disk layout.

Backup and restore is the most conservative for data integrity: the existing cluster is left untouched until a fresh 0.6.1 cluster is verified. It needs a maintenance window and capacity to run both clusters during the copy.

Method: datacenter expansion

Add the 0.6.1 stack as a second Cassandra datacenter in the same cluster, replicate the data to it, move clients over, then remove the old datacenter. Because the new nodes bootstrap empty and receive their data by streaming, this method avoids the volume-layout and configuration-compatibility concerns of an in-place adoption.

Validate this procedure in a non-production environment before running it against production. Perform the keyspace replication and decommission steps deliberately; an incorrect replication change or an early decommission can lose data.

Two settings on the new datacenter’s nodes decide whether the join and rebuild succeed:

  • storage_compatibility_mode must match the existing cluster. The existing Cassandra fleet runs NONE (full Cassandra 5.0 native format), and the 0.6.1 chart defaults the new nodes to NONE for this reason. If the new nodes run a different mode (the Apache Cassandra 5.0 default is CASSANDRA_4), gossip and schema exchange still succeed and the datacenters appear healthy, but the nodetool rebuild in step 4 fails: the streaming connection is rejected as version-incompatible.
  • Token allocation. When the existing cluster already holds a large schema, a node that joins with schema-aware token allocation (allocate_tokens_for_local_replication_factor) must reach schema readiness within a fixed window during bootstrap, or it aborts with a “Could not achieve schema readiness” error and restarts. If the new nodes fail to finish joining with that error, have them allocate tokens without the schema-aware option (random tokens are acceptable for a new, empty datacenter) and retry.
  1. Stand up the new datacenter joined to the existing cluster. Deploy the 0.6.1 Cassandra nodes with the same cluster name as the existing cluster, a distinct datacenter name, seeds that reach the existing datacenter, and auto_bootstrap: false so the new nodes join without streaming during bootstrap. Do not run the schema-initialization or migration jobs against the new datacenter: the schema already exists in the cluster and the new nodes receive it through gossip. Disable them in the new datacenter’s values with cassandra.hooks.initializeCluster.enabled: false and cassandra.hooks.migrations.enabled: false. This matters because the migrations include destructive ALTER statements and the new datacenter is not yet in every keyspace’s replication (that happens in step 3): running the migration job there can re-apply those statements from an earlier tracked version and race the existing schema.

  2. Confirm both datacenters are visible and healthy from a node in either datacenter:

    $kubectl -n cassandra-system exec cassandra-0 -c cassandra -- nodetool status

    Every node should report UN. The new datacenter nodes appear with no or minimal load until the rebuild in step 4.

  3. Extend replication for every application keyspace, and for system_auth, system_distributed, and system_traces, to include the new datacenter. For each keyspace:

    1ALTER KEYSPACE <keyspace> WITH replication = {
    2 'class': 'NetworkTopologyStrategy',
    3 '<old-datacenter>': <rf>,
    4 '<new-datacenter>': <rf>
    5};
  4. On each node in the new datacenter, stream the existing data across:

    $kubectl -n cassandra-system exec <new-dc-pod> -c cassandra -- \
    > nodetool rebuild -- <old-datacenter>
  5. Move clients to the new datacenter. Point application traffic at the new datacenter using LOCAL_QUORUM against the new datacenter name, and confirm reads and writes succeed there.

  6. Remove the old datacenter from replication, then decommission it. Alter every keyspace again to drop the old datacenter from the replication map, then decommission the old-datacenter nodes and remove the old stack.

Method: in-place volume adoption

Reuse the existing data volumes with the new stack. The StatefulSet must be recreated because its shape is immutable, and the new pods must be pointed at the Bitnami on-disk layout.

The StatefulSet recreate step deletes the StatefulSet object but keeps its pods and PersistentVolumeClaims. Do not delete the PersistentVolumeClaims. Confirm the StatefulSet’s persistentVolumeClaimRetentionPolicy.whenDeleted is not Delete before deleting it.

The chart supports adopting the existing volume with persistence.subPath: data, which surfaces the Bitnami nested layout at the paths the official image expects. A one-time relocation script is also available to convert the volume to the official layout so no subPath is needed afterward; use one or the other.

This method is validated for the standard NVCF Cassandra configuration. It also depends on the new nodes reading the data the old fleet wrote, which requires the 0.6.1 cassandra.yaml to match the settings the old data was written with. The chart carries the settings the default NVCF deployment needs. If your Bitnami deployment ran a customized cassandra.yaml, verify those settings carry over to 0.6.1 before you commit to this method: a missed setting can prevent the node from starting on the existing data. When in doubt, use datacenter expansion or backup and restore, which do not depend on on-disk configuration compatibility.

  1. Set the migration values. In the Cassandra values:

    • persistence.subPath: data, so the new pods adopt the existing volume and the official image finds the data at the paths it expects.
    • podSecurityContext.fsGroup: 999. The Bitnami image wrote the data as UID 1001; the official image runs as UID 999 and cannot read the existing files without this. Setting the fsGroup re-groups the volume to the official image’s user at mount time so it can read and write the adopted data.
    • the same cluster.name as the existing cluster. Cassandra refuses to start if the configured cluster name does not match the name stored in the adopted data.
    • the new image and the remaining auth and config values for 0.6.1.
  2. Recreate the Cassandra StatefulSet so the new-shape StatefulSet is created and re-adopts the existing volumes:

    $kubectl -n cassandra-system delete statefulset cassandra --cascade=orphan
    $kubectl -n cassandra-system delete pod cassandra-0
    $make install HELMFILE_ENV="$HELMFILE_ENV" HELMFILE_SELECTOR=name=cassandra

    For a multi-node cluster, delete and recreate one pod at a time and wait for each to become ready before the next.

  3. Confirm the new pod comes up on the existing data:

    $kubectl -n cassandra-system exec cassandra-0 -c cassandra -- nodetool status

    The node should report UN, and the row count for the known table you recorded earlier should match.

Method: backup and restore

Leave the existing cluster in place and load its data into a fresh 0.6.1 cluster.

  1. Snapshot the existing cluster: on every node, run nodetool flush then nodetool snapshot, and copy the snapshots off-node.

  2. Deploy a fresh 0.6.1 Cassandra cluster on new volumes and let the init and migration jobs create the schema.

  3. Restore the data into the new cluster with sstableloader, or place the snapshot sstables and run nodetool refresh per table.

  4. Verify row counts and application connectivity against the new cluster, then decommission the old cluster.

Verify the migration

After any method, confirm the Cassandra cluster is healthy on the 0.6.1 image:

$kubectl -n cassandra-system get pods
$kubectl -n cassandra-system exec cassandra-0 -c cassandra -- nodetool status

All Cassandra pods should be Running and report UN. Confirm the known table row count matches what you recorded before the migration, and confirm the control plane is healthy:

$kubectl get pods -A | grep -vE 'Running|Completed|READY'

This command should print nothing. Any output is a pod that is not yet Running or Completed.