Rancher upgrade
Plan and sequence COMMUNITY-edition Rancher upgrades across air-gapped multi-cluster fleets — a management/"hosting" Rancher cluster plus the downstream RKE2/K3s clusters it provisions. Covers the community release model (2.11→2.14, community-vs-Prime cadence, EOL), the Kontainer Driver Metadata (KDM) matrix deciding which downstream k8s minors each Rancher version can manage (and the stranding risk when a host-Rancher bump outruns its sub-clusters), cross-cluster upgrade ordering, the embedded-CAPI→Rancher-Turtles migration, Fleet coupling, cert-manager/Helm/backup prerequisites, backup- restore-operator + etcd rollback, and the air-gapped upgrade procedure (which images/charts/KDM to mirror). Assumes the Helm-on-Kubernetes install. Community editions only; Prime content excluded. Companion to k8s-components-checker, which owns the management-cluster k8s compatibility verdict; this skill owns the upgrade methodology and downstream coordination.From its SKILL.md
npx -y skills add air-gapped/skills --skill rancher-upgradeAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 3 stars3 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
SKILL.md
12.2 KB, ~2.7k tokens by cl100k_base, as published. Nobody here has run it
rancher-upgrade
Plan, sequence, and de-risk a community-edition Rancher upgrade across a real fleet: a
management ("hosting") cluster and the downstream RKE2/K3s clusters it provisions. The hard part
of a Rancher upgrade is almost never the helm upgrade itself — it's the cross-cluster
coordination: a management-cluster Rancher bump silently changes which downstream Kubernetes
versions the fleet can run, and getting the order wrong strands sub-clusters or forces a rebuild.
Community editions only. SUSE Rancher Prime backports, Prime-only patches, and Prime support matrices are out of scope — flag them when they appear, never build a plan on them. House Rule #1.
Companion to k8s-components-checker — who owns what
This skill is the upgrade-methodology companion to k8s-components-checker (same repo), mirroring
the argo-cd-apps ↔ compat/argo-cd.md split. The boundary is load-bearing — respect it to avoid
drift:
| Question | Skill |
|---|---|
| "Is Rancher version X compatible with the management cluster's k8s minor?" | k8s-components-checker — references/compat/rancher.md is the single source of truth for the mgmt-cluster k8s window per Rancher minor. Cite it; never restate those numbers here. |
| "Is component Y (Cilium, Rook, …) OK on k8s 1.NN on one cluster?" | k8s-components-checker (per-cluster verdict) |
| "How do I plan/sequence a Rancher version upgrade, and what happens to my downstream clusters?" | this skill |
| "Which downstream RKE2/K3s k8s minors can Rancher 2.X manage?" (KDM) | this skill |
| Harvester host→guest upgrade coordination | harvester-upgrade (same repo) — it owns the Harvester ladder. Load it alongside this skill when Harvester sits under the fleet: each Harvester hop is gated on upgrading the external Rancher first, so a Rancher plan here is the leading half of that pairing. Get the Harvester↔Rancher pairing from harvester-upgrade; never restate it here. |
The mental model — two coupled axes
A Rancher fleet upgrade moves on two axes that are decoupled but ordered:
- Management-cluster axis — the k8s minor the Rancher server itself runs on, and the Rancher
version. (k8s window → cite
compat/rancher.md.) - Downstream axis — which RKE2/K3s k8s minors Rancher can provision and manage, governed by Kontainer Driver Metadata (KDM), NOT by the management cluster's own k8s version.
The rule that ties them: introducing a downstream k8s minor requires upgrading Rancher first;
downstream patch upgrades do not. So the management Rancher always moves before any downstream
minor bump — and a trailing downstream cluster can fall out of support when the host Rancher rolls
forward. This is the single most common way a fleet upgrade goes wrong. See
references/kdm-downstream-matrix.md.
Workflow
1. Establish the change set and the fleet shape
- Install type first — this skill assumes the Helm-on-Kubernetes install. The same command that
gets the version answers this: an empty result means there is no Helm-managed
rancherrelease, and the likeliest reason is a single-node Docker install (docker run rancher/rancher, still a documented method). That upgrades by stopping the container, backing up its data volume, and starting a new container from the newer image against the same volume — nohelm upgrade, no cert-manager chart, no BRO, no etcd snapshot, no KDM coordination. Effectively nothing below applies. Say so and stop rather than emitting a Helm plan; point at the Rancher docs' single-node-with-Docker upgrade page. - Versions: current Rancher minor+patch, target Rancher minor. Get this from the cluster
(
helm list -A -n cattle-system | grep rancher), not from memory. - Fleet shape: does this Rancher actually provision downstream clusters?
kubectl get clusters.cluster.x-k8s.io -A(populated = real fleet; empty = standalone mgmt cluster — the CAPI/Turtles migration is then a near-non-event).kubectl get clusters.provisioning.cattle.io -A— any row whose name ≠localis a downstream cluster. - Air-gap? If the registry is internal-mirror-only, the air-gap procedure (what to mirror)
applies —
references/air-gap-procedure.md. (Direct-pull clusters skip mirroring entirely.) - Kubeconfig path: if the operator's kubeconfig is Rancher-proxied (
servercontains/k8s/clusters/), switch to a Rancher-independent admin kubeconfig BEFORE upgrading —helm upgradeof the rancher release restarts the proxy pods and severs the API mid-apply.references/prereqs-and-ordering.md§ Pre-flight.
2. Compute the upgrade path (no minor skipping)
The only supported path between minors is latest-COMMUNITY-patch-of-current-minor → latest- COMMUNITY-patch-of-next-minor, one minor at a time (2.11→2.12→2.13→2.14; never 2.11→2.14). Intra-minor patch jumps are fine.
⛔ The rung is the community ceiling, not the newest tag. Rancher ships both editions to one
GitHub feed, so for every minor except the current one the newest tag is Prime-only and
uninstallable here — sort -V | tail -1 silently yields a target the operator cannot use. Derive
each rung edition-aware (House Rule #1 + #3): references/lifecycle.md § Latest patch per minor.
3. For each minor step, run the pre-flight → upgrade → post-flight runbook
The per-minor breaking changes and the ordered runbook live in references/per-minor-runbook.md.
Prerequisites common to every step (cert-manager window, Helm floor, mandatory backup, RKE1 sweep,
API aggregation layer) and the cross-cluster ordering + rollback floor are in
references/prereqs-and-ordering.md.
4. Coordinate the downstream clusters
Before bumping the host Rancher, check every downstream cluster's k8s minor against the target
Rancher's KDM window (references/kdm-downstream-matrix.md). Lift any trailing downstream into
the new window first, or it loses manageability after the host upgrade. After the host upgrade
(which ships a new KDM branch), the newly-unlocked downstream minors become selectable — bump
downstreams then.
5. Produce the plan
Emit an ordered plan: per-minor-step pre-flight gates, the upgrade command (air-gap variant if applicable), post-flight verification, and the downstream-coordination steps interleaved at the right points. Cite the reference + grounded version for every specific claim. Recognizable shape:
<cluster/fleet> — Rancher <current> → <target> (path: <minor steps, no skips>)
prerequisites - cert-manager <window> · Helm <floor> · RKE1 sweep · BRO backup + etcd snapshot
downstream - <cluster>: k8s <ver> — <in target's KDM window | lift to ≥X BEFORE host upgrade>
source: references/kdm-downstream-matrix.md
per step (×N) 1. pre-flight gates 2. upgrade mgmt Rancher 3. post-flight 4. bump downstream
blockers - <one-way rollback / known regression / Prime-gated path> source: references/<file>.md
Every cited version is cluster-reported or freshly gh-grounded (House Rule #3), never from memory;
target versions follow look-ahead (House Rule #4).
House rules (encode into every plan)
- Community editions only. Flag Prime-gated versions/patches/features; never plan on them.
The reliable community-vs-Prime signal is the GitHub release-notes first line — see
references/lifecycle.md. - Cite, don't restate, the mgmt-cluster k8s window. That lives in
k8s-components-checker/references/compat/rancher.md. Pointing at it keeps the two skills from drifting. - Never invent versions; ground or abstain (House Rule #8 in k8s-components-checker). k8s windows and KDM
mechanics are methodology the skill states. Specific release/patch numbers — "latest 2.13
patch", "Turtles version on 2.14", "BRO chart for 2.12" — are volatile and the #1 fabrication
risk. State a specific version only if it is cluster-reported, freshly grounded via
gh, or explicitly markedUNVERIFIED. Grounding method (anti-confirmation): anchor ongh api repos/rancher/rancher/releases/latest, enumerate-and-derive the real latest patch, never ask "does vX.Y.Z exist?" (existence/list queries get rubber-stamped).ghmust be run with valid auth from the operator's workstation — anonymous is 60 requests/hour and exhausts almost immediately on an enumeration sweep (gh auth statusfirst). Repo map + protocol:references/lifecycle.md§ Grounding. - Look-ahead version targeting. When recommending any target version, pick the one that covers the operator's next planned hop too, not the bare immediate minimum — a version sitting at its own support ceiling forces an avoidable second upgrade. (Same rule as k8s-components-checker House Rule #9.)
- Management Rancher upgrades BEFORE downstream k8s minor bumps, always. And no minor skipping on the Rancher axis. Violating either is the classic fleet-stranding / rebuild path.
- Back up before every step. An RKE2 etcd snapshot of the management cluster is the real
rollback floor — the 2.14 CAPI-v1beta2 boundary is one-way without it. Add a backup-restore-operator
backup if BRO is installed; many community installs have none, so don't block on it (substitute
targeted resource exports — Rancher/Fleet objects + the auth AuthConfig & secrets). See
references/prereqs-and-ordering.md§ Backup & rollback.
Decision guide
| Task | Read |
|---|---|
Community vs Prime, cadence, EOL, latest-patch grounding, gh repo map | references/lifecycle.md |
| Which downstream k8s a Rancher minor can manage; stranding; air-gap KDM mirror | references/kdm-downstream-matrix.md |
| cert-manager / Helm / backup prereqs, cross-cluster ordering, etcd rollback | references/prereqs-and-ordering.md |
| embedded-CAPI→Turtles migration, CAPRKE2/CAAPF, Fleet per-minor + Helm v4 | references/capi-turtles-fleet.md |
Air-gapped upgrade: what to mirror, helm upgrade flags, downstream RKE2 SUC | references/air-gap-procedure.md |
| Per-minor (2.11→2.14) breaking changes + ordered pre/upgrade/post runbook | references/per-minor-runbook.md |
Each reference carries a header stating what was grounded and when — and what was NOT re-derived on the latest pass. Read that header before citing anything from the file: a claim under a "not re-derived" disclaimer must be re-grounded at use time regardless of the file's headline date. Otherwise re-ground per House Rule #3 only if the stamp is stale (≳1–2 weeks) — same-day re-grounding is wasted churn the operator will (rightly) push back on; don't present it as an unconditional step.
What ships with it: 8 files
69.4 KB alongside SKILL.md
references/
- air-gap-procedure.md6.8 KB
- capi-turtles-fleet.md4.6 KB
- improvement-backlog.md14.1 KB
- kdm-downstream-matrix.md5.5 KB
- lifecycle.md13.5 KB
- per-minor-runbook.md9.4 KB
- prereqs-and-ordering.md9.3 KB
- sources.md6.1 KB