Skip to content

Troubleshoot clusters and nodes

Reference diagram:

%%{init: {'themeVariables': {'edgeLabelBackground':'transparent'}}}%%
flowchart TB
    %% Control Plane
    subgraph ControlPlane [Control Plane Node]
        direction TB
        API[kube-apiserver]
        ETCD[(etcd cluster)]
        SCHED[kube-scheduler]
        CM[kube-controller-manager]
        CCM[cloud-controller-manager]

        %% Internal Control Plane Relationships
        ETCD -->|gRPC| API
        SCHED -->|HTTPS| API
        CM -->|HTTPS| API
        CCM -->|HTTPS| API
    end


        User([User / kubectl])
        User -->|HTTPS| API

    %% Worker Node 1
    subgraph Worker1 [Worker Node N]
        direction TB
        Klet1[kubelet]
        Kproxy1[kube-proxy]
        CR1{Container Runtime}

        subgraph Pods1 [Pods]
            direction LR
            Pod1A((Pod A))
            Pod1B((Pod B))
        end

        Klet1 <-->|GRPC| CR1
        CR1 --> Pod1A
        CR1 --> Pod1B
    end

    %% Cluster Communication
    Klet1 <-->|HTTPS| API
    Kproxy1 -->|HTTPS| API

    class ControlPlane controlplane;
    class Worker1,Worker2 worker;
    class Pod1A,Pod1B,Pod2A pod;
    class API core;

How can we determine the overall health of a cluster? simply running something like kubectl get nodes is a good way to test several, key components:

david@fedora:~/cka$ kubectl get no
NAME         STATUS   ROLES           AGE   VERSION
srv-rk1-01   Ready    control-plane   39d   v1.36.4+k3s1
srv-rk1-02   Ready    <none>          39d   v1.36.4+k3s1
srv-rk1-03   Ready    <none>          39d   v1.36.4+k3s1
srv-rk1-04   Ready    <none>          39d   v1.36.4+k3s1

If we can get an output from kubectl this tells us that:

  • The Kubernetes API server is running and responding to requests
  • ETCD is working

At a node level. we're looking at the status role and version values. If a node status is anything other than Ready we know the kubelet is reporting an issue.

kubectl describe node <name> - Enables us to inspect node conditions, for example, if it's exhibiting MemoryPressure, DiskPressure and allocatable resources

kubectl get events --sort-by-'.lastTimestamp' - Presents a list of cluster wide events, sorted in chronological order to spot systemic issues such as scheduling failures.

journalctl -u kubelet - Run this directly on a node to inspect the kubelet service logs if a node is not reporting a Ready status.

Exam Flashcards

Exam Tip

The Kubelet is the component that reports back to the Kubernetes API server. If a node is not ready. Checking its logs is often the first clue

Exam Tip

Ensure you're using the correct kubeconfig file that references the IP/FQDN of the Kubernetes API server you're trying to access.

Exam Tip

Try to discern the topology of the cluster you're working with - how many nodes their are, their roles, if roles are consolidated, etc.