How Rancher Saved Us During an OS Upgrade in an Air-Gapped Environment

Disclaimer: I am not an advocate of Rancher. Since it is open source, I hope sharing this won’t create any issues.

The Setup

We had a running RKE1 clusters in an air-gapped environment (I know it’s EOL, but it’s the customer’s choice). We were in a pre-production environment with two clusters, each an HA setup with 3 masters and 15 workers. Both clusters were, luckily, imported into Rancher.

What Happened

The task was an OS upgrade from RHEL 7.9 to 8.10, and a Docker upgrade from version 19 to 25. After the os upgrade and reboot, that machines simply wouldn’t come up and the both clusters were broken.

The Problem

At that point, we had two options, we could have used the cluster.yaml to rebuild the cluster, but we didn’t have it. No cluster.yaml, no RKE state file. Things were looking difficult.

How Rancher Saved Us

We rebuilt the machines with RHEL 8.10 and Docker installed. Then:

  • Deleting the node: In Rancher, navigate to your cluster, go to the Nodes section, and simply delete the affected node from the UI.
  • Re-adding the node: Head over to the Node Registration section in Rancher. It generates a Docker command that you simply run on the new machine something like:
docker run -d --privileged --restart=unless-stopped \
    -v /etc/kubernetes:/etc/kubernetes \
    -v /var/run:/var/run \
    rancher/rancher-agent:v2.x.x \
    --server https://<rancher-url> \
    --token <token> \
    --worker

That command deploys the Rancher agent container on the node, and it registers itself back into the cluster automatically. Pretty straightforward.

The Takeaway

If you are using rancher, always import your clusters into it, even if things seem fine. You never know when it’ll save you.

Now, we can easily move forward with prod upgrade :)