Practice Exams:

Cluster Upgrades Are an Operations Exercise, Not a Version Bump

 

A Kubernetes upgrade changes more than a version string. It changes a distributed control plane while workloads are running, nodes are serving traffic, add-ons depend on Kubernetes APIs, and automation assumes particular behaviors. The safest upgrade plan therefore looks more like a controlled operations exercise than a software installer.

The high-level order is simple: understand the current cluster, prepare recovery, upgrade the control plane, move through worker nodes, and verify the system at each stage. The complexity comes from the dependencies around that order. API removals can break manifests. A CNI or CSI driver can be incompatible with the target release. Pod disruption budgets can block node drains. An application may tolerate one node being offline but not two.

Current Kubernetes guidance recommends remaining on supported minor releases and moving to current patch levels promptly. That security advice is useful, but speed should not replace discipline. A fast upgrade that ignores workload availability, extension compatibility, or rollback evidence can create more risk than the old version it was meant to remove.

Begin with an inventory of what the cluster actually runs

Before choosing commands, record the cluster version, control-plane topology, node versions, container runtime, CNI, CSI drivers, ingress or Gateway implementation, metrics components, operators, admission webhooks, and other extensions that participate in normal operation. A cluster is an ecosystem, not just kube-apiserver and kubelet.

This inventory identifies the compatibility work that the Kubernetes release notes cannot do for you. A third-party controller might use an API removed in the target version. A storage driver may require its own upgrade first. A managed load-balancer controller may support only a particular Kubernetes range. Those constraints belong in the plan before maintenance begins.

People working through CKA cluster administration should make this dependency map a habit. Knowing the target version is useful; knowing which components must remain compatible while getting there is the operational skill.

Recovery preparation should be finished before the first component changes

For self-managed control planes, back up the state that would be difficult or impossible to reconstruct quickly. etcd deserves particular attention because it contains the Kubernetes API state. kubeadm upgrade workflows also create temporary backups of local etcd member data and static Pod manifests, but an operator should understand what those backups cover rather than assuming every failure can be rolled back automatically.

Recovery preparation also includes configuration, certificates, infrastructure-as-code definitions, and a record of the pre-upgrade state. If the cluster depends on external systems, know which ones would need reconciliation after a rollback or restore. A snapshot without the information required to rebuild access, networking, or storage is an incomplete recovery plan.

The same principle appears in disaster-recovery planning: recovery capability is proven by a tested procedure and known dependencies, not by the mere existence of a backup file.

Version skew and upgrade order are safety constraints

Kubernetes components are designed with supported version-skew relationships. That is why upgrade documentation specifies an order rather than telling administrators to update every binary simultaneously. The control plane is upgraded before worker nodes, and kubelet versions are managed within supported skew relative to the API server.

In a highly available control plane, upgrading one control-plane node at a time preserves quorum and API availability when the architecture is healthy. Worker nodes can then be moved through the process in controlled groups. The exact mechanics depend on how the cluster was installed, but the principle remains: change the smallest failure domain that still lets the system serve its purpose.

CLI tools such as kubectl also have supported version relationships. Treating client, server, kubelet, and add-on versions as one undifferentiated “Kubernetes version” hides the compatibility boundaries that matter during the transition.

Cordon and drain protect workloads differently

Cordoning a node marks it unschedulable for normal new Pods. It does not evict the Pods already running there. Draining goes further by using the eviction process to remove eligible workloads so they can run elsewhere before node maintenance. A safe upgrade uses these behaviors intentionally rather than treating node replacement as a surprise to the scheduler.

Pod disruption budgets, local data, DaemonSets, static Pods, and unmanaged Pods can all affect a drain. A blocked drain is not necessarily an obstacle to bypass; it may be evidence that the application cannot currently tolerate the disruption you planned. Forcing the operation can convert a maintenance task into an outage.

High-availability concepts from resilient cloud architecture apply directly: the maintenance plan must respect the number and placement of replicas required to keep the service available while infrastructure is intentionally removed.

API compatibility should be checked before the binary upgrade

Kubernetes evolves its APIs over time. A resource that worked on an older release can become deprecated and later removed. If manifests, Helm charts, operators, or admission integrations still depend on a removed API version, the failure may appear only when the new API server stops serving that endpoint.

That is why deprecation review belongs before the maintenance window. Search live resources and source repositories for APIs affected by the target release, update controllers that create those resources, and verify that replacement versions preserve the intended semantics. Fixing a deprecated manifest is easier while the old cluster still runs normally.

CustomResourceDefinitions need the same attention because an operator’s compatibility is not determined solely by built-in Kubernetes objects. Extension authors define their own versions, conversion behavior, and upgrade procedures. The cluster upgrade plan should include those application-level control planes as well.

Networking and storage add-ons can determine whether upgraded nodes are usable

A node is not operational merely because kubelet registers successfully. Pods still need a working CNI path, service networking, DNS, and any storage drivers required by their claims. If the CNI daemon, CSI node plugin, or service proxy fails on the new version, the node may be technically Ready while real workloads cannot function correctly.

Upgrade compatibility matrices for those components should be reviewed in advance. In some environments the add-on must be upgraded before Kubernetes; in others it follows. The important point is to treat add-ons as part of the platform, not as incidental Pods that will somehow adapt to any control-plane change.

The wider cloud-native ecosystem is built from these interacting projects. A Kubernetes version change can expose assumptions in networking, storage, observability, policy, and deployment tooling all at once.

Maintenance windows should also define explicit stop conditions. If API latency rises sharply, a control-plane component enters a restart loop, storage attachments fail on the upgraded node, or a representative application loses redundancy, the correct action may be to pause rather than continue because the calendar says the upgrade should finish. A good runbook contains criteria for proceeding, pausing, and rolling back.

The workload layer deserves pre-upgrade testing as well. Pod disruption budgets, anti-affinity rules, single-replica services, and local persistent data can all make a seemingly ordinary node drain unsafe. Finding those constraints before the window lets application owners increase replicas, relocate state, or approve a specific exception deliberately instead of discovering the issue while a node is already half upgraded.

Verification after each stage is more valuable than one final health check

After a control-plane node is upgraded, verify API responsiveness, control-plane component health, etcd health where applicable, and the ability of controllers to reconcile normal changes. After a worker node is upgraded, verify node conditions, CNI and CSI components, DNS, Service connectivity, volume attachment, and representative workloads.

Also watch what changes over time. A node that becomes Ready for thirty seconds and then develops MemoryPressure is not healthy. A controller that restarts repeatedly may still serve enough requests to make a shallow smoke test pass. Events, logs, rollout status, and application-level checks should agree that the platform is stable before the next batch moves.

This staged verification is a core DevOps operations habit: small changes with fast feedback create a safer path than a large coordinated change followed by a difficult search for which step caused the failure.

A canary approach can reduce uncertainty on worker upgrades. Move one representative node through the new version, schedule ordinary workloads onto it, and verify networking, storage, DNS, monitoring, and security agents before expanding the rollout. The canary should resemble the rest of the fleet closely enough to expose compatibility problems. If a special test node lacks the same CNI mode, storage attachments, kernel settings, or workload mix, a successful upgrade there provides false confidence rather than useful evidence.

After the rollout, update the baseline documentation rather than leaving the cluster in a half-documented transitional state. Record the final component versions, remove temporary compatibility settings, clear obsolete backups according to retention policy, and close any exceptions created for the maintenance. The next upgrade becomes safer when the previous one ends with an accurate inventory instead of a collection of undocumented one-off fixes.

Practice upgrades should teach decision points, not a memorized command sequence

A useful lab starts with a multi-node kubeadm cluster and a simple application that can demonstrate availability. Record versions and health, take the relevant backups, upgrade one control-plane node, verify the system, then proceed through remaining control-plane and worker nodes. Introduce one incompatibility or blocked drain so the exercise includes a real decision rather than a perfect happy path.

The CKA exam expects hands-on administration, but the commands are only the surface. The stronger skill is understanding why the upgrade order exists, what evidence permits the next step, and what signals should stop the change before the failure domain grows.

Across CNCF certifications and production environments, a successful upgrade is not “the new version installed.” It is a controlled transition in which the cluster, its extensions, and its workloads remain understandable and recoverable from the beginning of the change to the end.

Related Posts

• Why Network Segmentation Still Stops Real Attacks

• Least Privilege as an Architecture Principle

• Availability Sets, Zones, and Scale Sets Solve Different Problems

• Entra Groups, Roles, and Access Reviews in Everyday Administration

• Spanning Tree Still Matters in a World of Faster Switches

• Network Automation Starts With Structured Data, Not Python

• Agents Need Boundaries More Than They Need More Tools

• Data Governance for RAG Pipelines That Touch Sensitive Information

• Campus Fabric Changes Segmentation

• SD-WAN Policy Turns Intent Into Path Selection