Kubernetes troubleshooting
Every cluster operation is a task with logs, so start almost anywhere the same way: open the cluster’s Tasks tab and read the failing task’s log. Pool-specific symptoms are on Node pools; autoscaler symptoms are on Autoscaling.
The cluster is stuck in Provisioning or Starting
Section titled “The cluster is stuck in Provisioning or Starting”Open the Tasks tab; the failing task’s log names the phase. Usual causes:
- The VPC has no healthy NAT gateway. Workers cannot fetch packages or container images without outbound access, so bootstrap never finishes. Check the VPC’s NAT gateway, then retry.
- The version’s control plane or worker image was unregistered after the cluster was created. Your provider must re-register images for the version; see Admin setup.
- A task hung and was auto-failed. A task that makes no progress for 30 minutes is marked failed with the reason in its log, so you can retry the operation. See Upgrades and rotation.
Admins can use Reset State or Cancel Task on the cluster’s Destructive tab for a cluster whose state flag no longer matches reality.
kubectl gets connection refused or times out
Section titled “kubectl gets connection refused or times out”- Confirm the cluster is
Running. - For a public endpoint, the load balancer scope of the cluster’s security group must allow your source IP on TCP 443 (the API port on the load balancer). See Access and security.
- For a private-only endpoint,
kubectlmust run from inside the VPC, for example over a VPN gateway. - If the error is x509 rather than a refusal, your downloaded client certificate expired (90-day validity). Download a fresh kubeconfig.
Pods stuck in Pending
Section titled “Pods stuck in Pending”- Check node pressure on the Nodes tab. If autoscaling is on, the pool may have hit its Max Size; if off, raise Desired Size or add workers from the Node pools tab.
- The pod matches no pool. Its
nodeSelectoror tolerations do not match any pool’s labels and taints, or no pool’s plan fits the pod’s resource requests.kubectl describe pod <name>shows the scheduler’s reason. - The autoscaler is not installed. Panel bounds alone do not scale anything; apply the manifest from the Autoscaler tab. See Autoscaling.
A Service of type=LoadBalancer never gets an EXTERNAL-IP
Section titled “A Service of type=LoadBalancer never gets an EXTERNAL-IP”Run kubectl describe svc <name> and read the failure event at the bottom; its message carries the exact cause.
The most common causes, in the error text you will see:
annotation 'service.beta.kubernetes.io/managed-loadbalancer-plan' is required: every Service needs the plan annotation naming the load balancer plan to deploy. Add it and re-apply the Service.unknown_lb_plan: the plan name in the annotation does not exist or is not enabled. Your provider lists the valid names.lb_plan_not_in_region: the named plan is not offered in the cluster’s region. Pick a plan attached to that region.
The full annotation set is documented on Load balancer Services. Provisioning normally completes within 30 to 90 seconds of the annotations being valid.
An upgrade or rotation stalls on one node
Section titled “An upgrade or rotation stalls on one node”A pod with a strict PodDisruptionBudget, or a safe-to-evict: false annotation, blocks the drain. Relax the budget or scale the workload temporarily, raise the drain grace period, then retry. The rolling operations resume where they stopped; the mechanics are on Upgrades and rotation.
Nodes come up NotReady after provisioning
Section titled “Nodes come up NotReady after provisioning”The CNI failed at cluster create time. This is image-side: confirm with your provider that the CNI containers are pre-pulled in the registered worker image, and that the Pod CIDR the cluster was created with matches the CNI’s expectation (Cilium accepts any; Flannel defaults to 10.244.0.0/16). See Image baking.

