Skip to content

Upgrades and rotation

A cluster changes shape in three ways: the control plane moves to a newer Kubernetes version, the workers follow, and a node pool moves onto a different instance plan. All three are rolling operations: capacity is replaced in waves, so the cluster keeps serving while it changes.

  • A running cluster. See Clusters.
  • The target Kubernetes version is registered and Active, with a valid upgrade path from your current version. Your provider curates the list; see Admin setup.
  • Workloads have sane PodDisruptionBudgets. Every rolling operation drains nodes, and a strict budget is the usual reason an operation stalls.

In the user panel go to Services > Clusters, open the cluster, then use the Upgrades tab. This tab is user-panel only; the admin panel’s cluster page has no Upgrades tab. Progress for every operation lands on the Tasks tab. Plan rotation lives on the Node pools tab.

The tab shows two cards: Upgrade Workers first, Upgrade Control Plane below it.

Workers can only be upgraded up to the control plane’s current version; upgrade the control plane first if you are moving a minor version ahead. In the Upgrade Workers card:

Field What to enter
Target version The version to move workers to. Same forward-only, one-minor-step rules as the control plane.
Max surge How many extra workers may exist at once during the upgrade (1 to the current worker count). Higher is faster and costs more temporarily.
Drain grace (sec) Shutdown time per pod during drains (0-3600).

Click Upgrade. The upgrade runs in waves: surge workers are provisioned on the target version and confirmed ready, then the same number of old workers are cordoned, drained and destroyed. The pool never drops below its target capacity.

In the Upgrade Control Plane card:

  1. Pick the Target version. Only valid upgrade targets from the current version are offered.
  2. Set Drain grace (sec): how long each pod gets to shut down cleanly during the drain (0-3600).
  3. Click Upgrade CP.

The upgrade runs in two phases. First, new control plane nodes on the target version are provisioned one at a time; each one is confirmed ready, verified as an etcd member and added to the API load balancer before the next starts. Then the old control plane nodes are removed one at a time. The API endpoint stays reachable through the load balancer throughout. A concurrent worker scale or worker upgrade is blocked while a control plane upgrade runs.

Changing a pool’s Instance Plan does not touch existing workers; new workers use the new plan and the pool shows a drift badge counting workers still on the old one. To move them, click the badge (Rotate now):

  • Rotation is a surge replacement, one wave at a time: new workers on the new plan are provisioned and confirmed ready, then old workers are cordoned, drained and destroyed.
  • The pool never runs below its target capacity during the rotation.
  • Downsizing to a smaller plan asks for an explicit second confirmation.
  • An interrupted rotation resumes on its own where it left off.

Every upgrade and rotation is a task on the Tasks tab with live logs. When an operation stalls:

  • Read the task log first; it names the node and phase. A pod with a strict PodDisruptionBudget is the usual blocker. Relax the budget or scale the workload temporarily, then retry.
  • A task that makes no progress for 30 minutes is automatically marked failed, with the reason written to its log, so the operation can be retried instead of hanging forever.
  • A plan rotation that never actually started is re-dispatched automatically, once.

Admins can also use Cancel Task on the cluster’s Destructive tab to stop an active task that is no longer making progress. See Admin setup.

  • The target version is not offered. Upgrade paths are curated: the target must be Active, a forward step of at most one minor version, and for workers no newer than the control plane. Ask your provider to register the missing version or path.
  • An upgrade stalls on one node. A pod with a strict disruption budget blocks the drain. Relax it or raise the drain grace period, then retry.
  • Rotation seems stuck. Check the Tasks tab; if the task was marked failed by the inactivity watchdog, start the rotation again. It resumes from where it stopped.