Linux Foundation
Linux Foundation KCNA: Observability for Kubernetes Workloads
Kubernetes makes infrastructure more dynamic, but that dynamism can make failure harder to understand. Pods restart, endpoints move, nodes come and go, controllers continually reconcile state, and application requests cross several layers before they reach a backend. Observability provides the evidence needed to explain that behavior rather than guessing from a single dashboard. For learners building cloud-native infrastructure, the useful model is to treat metrics, logs, traces, and Kubernetes events as complementary signals. The current KCNA scope places observability inside cloud-native architecture, but the production skill is broader: operators need…
Linux Foundation KCNA: Kubernetes Service Discovery in Practice
Kubernetes workloads are ephemeral: Pods are created, replaced, scaled, and rescheduled as controllers maintain desired state. Service discovery gives clients a stable way to reach a logical application even though the individual backend Pod addresses change. In most clusters, that stability comes from Service objects, EndpointSlices, and cluster DNS working together. Within cloud-native infrastructure, service discovery connects application design to cluster networking. The current KCNA competencies include networking, troubleshooting, containerization, and application delivery, all of which depend on understanding how a name becomes traffic to a healthy backend. Troubleshooting should…
Linux Foundation KCNA: Kubernetes Pod Networking
Kubernetes assumes that each Pod receives its own cluster-reachable IP address and that Pods can communicate across nodes according to the cluster network model unless policy intentionally restricts the traffic. Kubernetes defines the model and APIs, while a network implementation—commonly through CNI—provides the data plane that makes those addresses and routes real. Within cloud-native infrastructure, Pod networking is the layer that connects scheduling to application communication. The current KCNA competencies include networking under container orchestration, and Kubernetes documentation separates Pod networking from Service proxying, NetworkPolicy, and external ingress. Troubleshooting improves…
Linux Foundation KCNA: Container Runtime Fundamentals
Kubernetes does not run application processes by itself. On each node, kubelet relies on a container runtime to pull images, create and start containers, manage their lifecycle, and report status. The Container Runtime Interface gives Kubernetes a standard way to communicate with different runtimes without baking one runtime implementation into kubelet. Within cloud-native infrastructure, the runtime layer explains many problems that otherwise look like mysterious Pod failures. The current KCNA competencies include containerization under Kubernetes fundamentals, while current Kubernetes documentation describes CRI as the main protocol between kubelet and the…
CNCF CKA: Kubernetes Troubleshooting Starts With the Control Plane Story
Kubernetes troubleshooting becomes much easier when the cluster stops looking like a collection of commands and starts looking like a control system. A workload declaration enters through the API, controllers reconcile desired state, the scheduler selects a node, the kubelet turns the Pod specification into a running workload, networking makes services reachable, and storage provides state where needed. When something breaks, the fastest path is usually to find the point where that story stopped progressing. This mental model matters more than memorizing a long sequence of kubectl commands. A…
CNCF CKA: Pods Are Disposable, State Is Not
Kubernetes is built around the idea that Pods are replaceable. A controller can create a new Pod when one fails, a scheduler can place the replacement on another node, and a rolling update can deliberately destroy old Pods while new ones take over. That model is powerful for stateless applications because compute instances become temporary implementation details rather than permanent pets. Data does not follow the same lifecycle automatically. A database, message store, artifact repository, or other stateful workload needs information to survive Pod replacement and often node replacement…
CNCF CKA: Services, Ingress, and Gateway APIs for Cluster Traffic
Kubernetes networking becomes much easier to reason about when traffic is treated as a path rather than as a collection of resource types. A client does not reach a Pod because a Service, Ingress, or Gateway object merely exists. Traffic has to move through a sequence of decisions: how the client finds an entry point, which routing object accepts the request, which Service represents the backend, which EndpointSlices identify usable Pods, and which data-plane implementation actually forwards packets. That sequence matters because Pods are deliberately replaceable. Their IP addresses…
CNCF CKA: RBAC in Kubernetes: Who Can Do What, Where, and Why
Kubernetes RBAC is easiest to understand as an authorization decision about an API request. An identity asks the API server to perform a verb on a resource in a scope. RBAC rules and bindings determine whether that request is allowed. The important words are not the object names by themselves; they are who, can do what, to which resources, and where. That framing separates several ideas that are often blurred together. Authentication establishes who the caller is. Authorization determines what that identity may do. Admission can then evaluate or…
CNCF CKA: Persistent Volumes From Provisioning to Mounting
Kubernetes storage becomes confusing when provisioning, binding, attaching, and mounting are treated as one event. They are separate stages with different owners and different failure modes. A PersistentVolumeClaim can be perfectly valid while no suitable volume exists. A claim can be bound while the storage cannot attach to the chosen node. A volume can attach while a filesystem or mount operation still fails. The durable mental model starts by separating the API abstraction from the storage system underneath it. A PersistentVolume represents storage available to the cluster. A PersistentVolumeClaim…
CNCF CKA: Scheduling Problems
A Pending Pod is not evidence that the Kubernetes scheduler is confused. It usually means the scheduler is doing exactly what the Pod specification and cluster state require, but no node satisfies all of the constraints at the same time. The fastest way to troubleshoot scheduling is therefore to reconstruct the filter that the scheduler is applying. Resource requests are one part of that filter. Taints can repel Pods. Tolerations can make a taint acceptable. Node selectors and node affinity can restrict eligible nodes. Pod affinity and anti-affinity can…
CNCF CKA: etcd Is Small, Critical, and Worth Understanding
etcd is easy to ignore because most Kubernetes administrators interact with the API server rather than with the backing store directly. Yet the control plane depends on etcd for authoritative cluster state. Deployments, Secrets, ConfigMaps, Nodes, custom resources, leases, and the rest of the API model ultimately depend on a storage system that must remain consistent and available. That does not mean every administrator should treat etcd as a database to tune casually. The opposite is safer. etcd deserves respect because unnecessary intervention can turn a recoverable control-plane issue…
CNCF CKA: Cluster Upgrades Are an Operations Exercise, Not a Version Bump
A Kubernetes upgrade changes more than a version string. It changes a distributed control plane while workloads are running, nodes are serving traffic, add-ons depend on Kubernetes APIs, and automation assumes particular behaviors. The safest upgrade plan therefore looks more like a controlled operations exercise than a software installer. The high-level order is simple: understand the current cluster, prepare recovery, upgrade the control plane, move through worker nodes, and verify the system at each stage. The complexity comes from the dependencies around that order. API removals can break manifests….
CNCF CKA: Network Policies: Security Depends on the CNI You Actually Run
Kubernetes NetworkPolicy is a powerful example of declarative intent that depends on an implementation. You can create a perfectly valid NetworkPolicy object and see it stored by the API server even when the cluster network does not enforce that policy. The security outcome therefore depends not only on the YAML but on the CNI or network implementation that turns the policy into packet filtering. This matters because Kubernetes networking is generally open between Pods unless something deliberately restricts it. A NetworkPolicy selects Pods and defines allowed ingress, egress, or…
CNCF CKA: Resource Requests and Limits Shape Cluster Stability
Kubernetes resource settings are not merely performance tuning. Requests influence where Pods can be scheduled, limits influence how the node enforces consumption, and the combination affects quality-of-service classification, eviction behavior, capacity planning, and the amount of useful work a cluster can safely host. Poor values can make a cluster look full when it is idle or make it look efficient until memory pressure causes a cascade of failures. The important distinction is that a request and a limit answer different questions. A request tells the scheduler how much of…
CNCF CKA: Reading Kubernetes Events Before Reaching for the Logs
Kubernetes events are often the shortest path from a vague symptom to the component that first noticed it. A Pod is Pending, a volume will not mount, an image cannot be pulled, or a node turns NotReady; in each case, the control plane and node agents may already have emitted a concise explanation before an administrator opens any application log. That does not make events a replacement for logs. Events are best-effort, have limited retention, and summarize state changes rather than preserving every detail. Their value is triage. They…