Troubleshoot cluster components
Reference diagram:
%%{init: {'themeVariables': {'edgeLabelBackground':'transparent'}}}%%
flowchart TB
%% Control Plane
subgraph ControlPlane [Control Plane Node]
direction TB
API[kube-apiserver]
ETCD[(etcd cluster)]
SCHED[kube-scheduler]
CM[kube-controller-manager]
CCM[cloud-controller-manager]
%% Internal Control Plane Relationships
ETCD -->|gRPC| API
SCHED -->|HTTPS| API
CM -->|HTTPS| API
CCM -->|HTTPS| API
end
User([User / kubectl])
User -->|HTTPS| API
%% Worker Node 1
subgraph Worker1 [Worker Node N]
direction TB
Klet1[kubelet]
Kproxy1[kube-proxy]
CR1{Container Runtime}
subgraph Pods1 [Pods]
direction LR
Pod1A((Pod A))
Pod1B((Pod B))
end
Klet1 <-->|GRPC| CR1
CR1 --> Pod1A
CR1 --> Pod1B
end
%% Cluster Communication
Klet1 <-->|HTTPS| API
Kproxy1 -->|HTTPS| API
class ControlPlane controlplane;
class Worker1,Worker2 worker;
class Pod1A,Pod1B,Pod2A pod;
class API core;
Control Plane Node Components
ETCD
Usually, most etcd implementations also include etcdctl, which can aid in monitoring the state of the cluster. If you’re unsure where to find it, execute the following:
find / -name etcdctl
Leveraging this tool to check the etcd cluster status:
etcdctl --write-out=table --endpoints=$ENDPOINTS endpoint status
+------------------------+------------------+---------+---------+-----------+------------+-----------+------------+--------------------+--------+
| ENDPOINT | ID | VERSION | DB SIZE | IS LEADER | IS LEARNER | RAFT TERM | RAFT INDEX | RAFT APPLIED INDEX | ERRORS |
+------------------------+------------------+---------+---------+-----------+------------+-----------+------------+--------------------+--------+
| https://127.0.0.1:2379 | 4e30a295f2c3c1a4 | 3.6.5 | 8.1 MB | true | false | 3 | 7903 | 7903 | |
+------------------------+------------------+---------+---------+-----------+------------+-----------+------------+--------------------+--------+
The cluster this was executed on has only one control plane node, hence only one result from the script. You will normally receive a response for each etcd member in the cluster.
Alternatively, leverage kubectl get componentstatuses:
kubectl get componentstatuses #ComponentStatus is deprecated in v1.19+
NAME STATUS MESSAGE ERROR
scheduler Healthy ok
controller-manager Healthy ok
etcd-1 Healthy {"health":"true"}
etcd-0 Healthy {"health":"true"}
Etcd may also be running as a Pod:
kubectl logs etcd -n kube-system
Kube-apiserver
This is dependent on the environment for which the Kubernetes platform has been installed on. For systemd based systems:
journalctl -u kube-apiserver
Or
cat /var/log/kube-apiserver.log
Or for instances where Kube-API server is running as a static pod:
kubectl logs kube-apiserver-k8s-master-03 -n kube-system
Kube-Scheduler
For systemd-based systems
journalctl -u kube-scheduler
Or
cat /var/log/kube-scheduler.log
Or for instances where Kube-Scheduler is running as a static pod:
kubectl logs kube-scheduler-k8s-master-03 -n kube-system
Kube-Controller-Manager
For systemd-based systems
journalctl -u kube-controller-manager
Or
cat /var/log/kube-controller-manager.log
Or for instances where Kube-controller manager is running as a static pod:
kubectl logs kube-controller-manager-k8s-master-03 -n kube-system
Worker Node(s)
CNI
Obviously this is dependent on the CNI in use for the cluster you’re working on. However, using Flannel as an example:
journalctl -u flanneld
If running as a pod, however:
Kubectl logs --namespace kube-system <POD-ID> -c kube-flannel
kubectl logs --namespace kube-system weave-net-pwjkj -c weave
Kube-Proxy
For systemd-based systems
journalctl -u kube-proxy
Or
cat /var/log/kube-proxy.log
Or for instances where Kube-proxy manager is running as a static pod:
kubectl logs kube-proxy -n kube-system
Kubelet
journalctl -u kubelet
Or
cat /var/log/kubelet.log
Container Runtime
Similarly to the CNI, this depends on which container runtime has been deployed, but using containerd as an example:
For systemd-based systems:
journalctl -u containerd
Or
cat /var/log/containerd.log
Exam Tip
A lot of modern distributions, including kubeadm package control plane components as static Pods. Therefore, we can use regular kubectl commands to extract logs.
Exam Tip
A lot of modern distributions, including kubeadm deploy the Kubelet as a systemd service. Therefore, we can use journalctl to view its logs. This is a crucial communication path between worker nodes and the API server.