ISSUES FACED DURING REBUILDING A KUBERNETES CLUSTER
We had a Kubernetes POC cluster with version 1.18. This cluster got corrupted during experimentation by our team. Instead of starting with clean slate, which is comparatively easy, we tried to rebuild cluster with existing kubelet, etcd, kubeadm. This post is basically my own reference/notes of the issues faced and their fixes.
Issue 1: kubeadm join failing with below error. This happens only while join control plane nodes not worker nodes.
failure loading certificate for CA: couldn’t load the certificate file /etc/kubernetes/pki/ca.crt: open /etc/kubernetes/pki/ca.crt: no such file or directory
Solution: Copied following files from node 1 to other control plane nodes
/etc/kubernetes/pki/ca.crt
/etc/kubernetes/pki/ca.key
/etc/kubernetes/pki/sa.key
/etc/kubernetes/pki/sa.pub
/etc/kubernetes/pki/front-proxy-ca.crt
/etc/kubernetes/pki/front-proxy-ca.key
Issue 2: Flannel pods are crashlooping. Found following error continuously in kube-proxy logs.node.go:125]
Failed to retrieve node info: Unauthorized
Solution: deleting kube-proxy-token-xxxxx secret resolved issue.
Issue 3: Core DNS pods are in running state but were not in READY state. With kubectl describe found the following event:
Readiness probe failed: HTTP probe failed with statuscode: 503
pod logs have following messages:
pkg/mod/k8s.io/client-go@v0.17.2/tools/cache/reflector.go:105: Failed to list *v1.Endpoints: Unauthorized
pkg/mod/k8s.io/client-go@v0.17.2/tools/cache/reflector.go:105: Failed to list *v1.Namespace: Unauthorized
pkg/mod/k8s.io/client-go@v0.17.2/tools/cache/reflector.go:105: Failed to list *v1.Service: Unauthorized
pkg/mod/k8s.io/client-go@v0.17.2/tools/cache/reflector.go:105: Failed to list *v1.Endpoints: Unauthorized
pkg/mod/k8s.io/client-go@v0.17.2/tools/cache/reflector.go:105: Failed to list *v1.Namespace: Unauthorized
pkg/mod/k8s.io/client-go@v0.17.2/tools/cache/reflector.go:105: Failed to list *v1.Service: Unauthorized
Solution: kubectl delete secret <podname> -n kube-system
Issue 4: networkPlugin cni failed to set up pod network: Found below error in logs:
failed to delegate add: failed to set bridge addr: “cni0” already has an IP address different from
Solution: Delete network interfaces and restart kubelet which will create the network interfaces again.
Issue 5: pods are getting evicted with reason: node(s) had taint {node.kubernetes.io/disk-pressure
Solution: /var/log directory utilization is more than 80%. Removed old logs.
Issue 6: Unable to connect to the server: x509: certificate signed by unknown authority (possibly because of “crypto/rsa: verification error” while trying to verify candidate authority certificate “kubernetes”
Resolution: Copy config file to $HOME/.kube and provide necessary permissions.
mkdir -p $HOME/.kube
sudo cp -i /etc/kubernetes/admin.conf $HOME/.kube/config
sudo chown $(id -u):$(id -g) $HOME/.kube/config
Comments
Post a Comment