ISSUES FACED DURING REBUILDING A KUBERNETES CLUSTER

We had a Kubernetes POC cluster with version 1.18. This cluster got corrupted during experimentation by our team. Instead of starting with clean slate, which is comparatively easy, we tried to rebuild cluster with existing kubelet, etcd, kubeadm. This post is basically my own reference/notes of the issues faced and their fixes.

Issue 1: kubeadm join failing with below error. This happens only while join control plane nodes not worker nodes.

failure loading certificate for CA: couldn’t load the certificate file /etc/kubernetes/pki/ca.crt: open /etc/kubernetes/pki/ca.crt: no such file or directory

Solution: Copied following files from node 1 to other control plane nodes
/etc/kubernetes/pki/ca.crt
/etc/kubernetes/pki/ca.key
/etc/kubernetes/pki/sa.key
/etc/kubernetes/pki/sa.pub
/etc/kubernetes/pki/front-proxy-ca.crt
/etc/kubernetes/pki/front-proxy-ca.key

Issue 2: Flannel pods are crashlooping. Found following error continuously in kube-proxy logs.node.go:125]

Failed to retrieve node info: Unauthorized

Solution: deleting kube-proxy-token-xxxxx secret resolved issue.

Issue 3: Core DNS pods are in running state but were not in READY state. With kubectl describe found the following event:

Readiness probe failed: HTTP probe failed with statuscode: 503

pod logs have following messages:

pkg/mod/k8s.io/client-go@v0.17.2/tools/cache/reflector.go:105: Failed to list *v1.Endpoints: Unauthorized
pkg/mod/k8s.io/client-go@v0.17.2/tools/cache/reflector.go:105: Failed to list *v1.Namespace: Unauthorized
pkg/mod/k8s.io/client-go@v0.17.2/tools/cache/reflector.go:105: Failed to list *v1.Service: Unauthorized
pkg/mod/k8s.io/client-go@v0.17.2/tools/cache/reflector.go:105: Failed to list *v1.Endpoints: Unauthorized
pkg/mod/k8s.io/client-go@v0.17.2/tools/cache/reflector.go:105: Failed to list *v1.Namespace: Unauthorized
pkg/mod/k8s.io/client-go@v0.17.2/tools/cache/reflector.go:105: Failed to list *v1.Service: Unauthorized

Solution: kubectl delete secret <podname> -n kube-system

Issue 4: networkPlugin cni failed to set up pod network: Found below error in logs:

failed to delegate add: failed to set bridge addr: “cni0” already has an IP address different from

Solution: Delete network interfaces and restart kubelet which will create the network interfaces again.

Issue 5: pods are getting evicted with reason: node(s) had taint {node.kubernetes.io/disk-pressure

Solution: /var/log directory utilization is more than 80%. Removed old logs.

Issue 6: Unable to connect to the server: x509: certificate signed by unknown authority (possibly because of “crypto/rsa: verification error” while trying to verify candidate authority certificate “kubernetes”

Resolution: Copy config file to $HOME/.kube and provide necessary permissions.
mkdir -p $HOME/.kube
sudo cp -i /etc/kubernetes/admin.conf $HOME/.kube/config
sudo chown $(id -u):$(id -g) $HOME/.kube/config

Comments

Popular posts from this blog

HOW WE REDUCED SOA OSB PROVISIONING FROM 4 DAYS TO 4 HOURS

NOT ABLE TO START RABBITMQ CLUSTER: CANNOT DECLARE A QUEUE ‘~S’ ON NODE ‘~S’: ~255P

SOA SUITE 12.2.1.4 INSTALLATION: GOT EXCEPTION WHEN AUTO CONFIGURING THE SCHEMA COMPONENT(S) WITH DATA OBTAINED FROM SHADOW TABLE