Practice Exams:

Backups Are Not a Disaster-Recovery Plan

 

Backups are essential, but a backup is only one input to disaster recovery. A snapshot can preserve bytes while the service that depends on those bytes remains unrecoverable because infrastructure, identity, networking, DNS, secrets, configuration, dependencies, or operational knowledge are missing. Treating “we have backups” as proof of recoverability confuses data protection with service restoration.

This distinction sits inside the resilience decisions covered by SAA-C03. An architect has to design for defined recovery objectives, not simply enable a backup feature. The business requirement is usually expressed as how much data loss is acceptable and how long the service can be unavailable; the architecture must then prove that the entire workload can meet those limits.

A strong disaster-recovery plan therefore begins with RTO and RPO, maps every dependency required for recovery, assigns ownership for the failover process, and tests the path under realistic conditions. Backups matter, but recovery is a system property.

RPO describes recoverable data; RTO describes recoverable service

Recovery Point Objective is the maximum acceptable amount of data loss measured in time. If a database is backed up every hour, the theoretical backup interval might suggest an RPO near one hour, but actual recoverability also depends on whether the backup completed, whether it is readable, and whether later logs or continuous replication are required to reach the desired point. Recovery Time Objective is the maximum acceptable time to restore service after disruption.

Those two objectives lead to different architectures. A workload that can tolerate a day of downtime might use backup-and-restore into a recovery Region. A service that must return in minutes may require pre-provisioned infrastructure, replicated data, tested routing changes, and a warm or active secondary environment. The recovery strategy is therefore a response to business consequences rather than a generic “best practice.”

The distinction between AWS high availability and fault tolerance matters here as well: high availability reduces disruption from expected component failures, while disaster recovery prepares for a broader loss of capability. All three can use redundancy, but they start from different failure assumptions.

A restore that cannot rebuild dependencies is not a recovery plan

Imagine restoring an RDS snapshot successfully while the original VPC no longer exists, security groups are missing, application subnets use different routes, the KMS key is unavailable, secrets were not recreated, and DNS still points at the failed environment. The data may be intact and the service may still be hours or days away from functioning. This is why recovery planning must inventory dependencies, not just storage systems.

Infrastructure as code can reduce this gap by making network, compute, IAM, load-balancing, and application configuration reproducible. But the code itself has dependencies: state files, repositories, artifact registries, credentials, deployment permissions, and pipeline runners must remain accessible during the same disaster. A plan that assumes its deployment system survives without testing that assumption contains a hidden single point of failure.

Operationally mature teams maintain a recovery bill of materials: data sources, keys, secrets, images, network constructs, certificates, DNS records, queues, external integrations, quotas, and required people or approval paths. The goal is not documentation for its own sake; it is to make the restoration sequence executable under pressure.

Recovery strategies trade cost for time and certainty

AWS commonly frames disaster-recovery strategies along a spectrum. Backup and restore minimizes continuously running recovery infrastructure but usually produces the longest RTO. Pilot-light patterns keep critical core components ready while much of the environment is recreated during recovery. Warm standby maintains a scaled-down but functional environment. Multi-site active-active keeps more than one site serving production traffic and can reduce failover time further, but it increases cost and application complexity.

The labels are useful only if the details match the application. A “warm standby” database that is current but an application tier that takes hours to deploy may not satisfy the business RTO. An active-active front end whose data layer cannot accept consistent writes in both Regions may still require a carefully controlled failover. Recovery time is determined by the slowest critical step, not by the most impressive component in the architecture diagram.

At portfolio and organization scale, SAP-C02 extends the same reasoning across organizational complexity, migration, network design, cost, and operating processes, where recovery choices can no longer be made application by application in isolation.

Backups need isolation from the incident they are supposed to survive

A backup that can be deleted by the same compromised identity as production may not be useful during a destructive security event. Recovery design should therefore consider account separation, vault controls, retention policies, immutability where appropriate, encryption-key availability, and the permissions required to restore. The threat model includes mistakes and malicious actions, not only hardware failure.

Cross-Region or cross-account copies can reduce common-mode risk, but location alone does not guarantee protection. If an automation bug, retention policy, or privileged credential can affect both copies, the architecture may still have correlated failure. Recovery assets should be protected by deliberately different controls where the risk justifies it.

Testing must also validate that encrypted backups can actually be decrypted in the recovery path. Losing or disabling the required KMS key can turn a perfectly retained backup into unusable ciphertext. Key policy, grants, cross-account access, and recovery ownership should be part of the recovery exercise rather than discovered during an emergency.

Failover requires operational decisions that backups cannot make

Someone must decide when an outage becomes a disaster, who authorizes failover, whether the process is automated, how split-brain is prevented, and how customers are routed to the recovery environment. These are control-plane questions as much as infrastructure questions. A technically recoverable system can still miss its RTO if the organization spends hours deciding whether it is safe to switch traffic.

Runbooks should define evidence and decision points: what health signals indicate the primary environment is not recoverable quickly, what data-replication lag is acceptable, what features can operate in degraded mode, and who communicates status to dependent teams. If the recovery process requires manual steps, those steps should be rehearsed often enough that they are not novel during the incident.

The operational side is explicit in SOA-C03: monitoring, business continuity, recovery, automation, and incident response determine whether an architecture is recoverable in practice rather than merely on paper.

Restore testing should measure the whole path

A successful test is not “the snapshot restored.” It is “the business service returned within the target, with an acceptable recovery point, and users could complete critical transactions.” That means exercising data restoration, infrastructure deployment, application startup, certificate and secret retrieval, DNS or routing changes, dependency validation, observability, and rollback or return-to-primary procedures.

Tests also reveal capacity assumptions. The recovery Region may have lower service quotas, unavailable instance types, insufficient IP space, or missing reserved capacity. External partners may whitelist only primary addresses. A queue may contain messages that need replay, while downstream systems are not ready to accept them. These details are invisible until the recovery plan is executed end to end.

For AWS Certified CloudOps Engineer – Associate work, recovery testing is an operations discipline. Teams should observe the process, record timing, identify manual bottlenecks, verify alerting and access paths, and use each exercise to reduce uncertainty in the next one.

Recovery planning changes architecture before a disaster occurs

Once teams model recovery honestly, they often redesign production. They may reduce hidden state, automate deployments, remove static credentials, replicate critical configuration, simplify DNS, isolate backups, or choose managed services whose recovery characteristics are easier to operate. Disaster recovery is therefore not a document added after architecture; it can expose whether the architecture is reproducible and understandable.

The AWS Certified Solutions Architect – Associate baseline reinforces the architectural point: resilience begins with explicit failure assumptions. A service that can only be rebuilt by one engineer following undocumented steps has a recovery risk even if every database is backed up.

Backups answer “can we recover data?” Disaster recovery asks a larger question: “can we restore the service, with the right data, within the time the business can tolerate?” Treating those as separate questions is the first step toward a recovery strategy that can survive contact with a real incident.

Application state outside the database can break an otherwise successful recovery

Database-centric recovery plans often miss state stored elsewhere. Object storage may contain generated files or uploaded documents. Queues may contain unprocessed work. Search indexes may lag the system of record. Caches may be disposable but expensive to warm. Certificates, secrets, feature flags, scheduler state, and external SaaS configuration may all influence whether the restored service behaves correctly. Teams should decide which state must be backed up, which can be reconstructed, and which must be replayed from another source of truth.

This matters especially after partial failures. If a queue continues accepting work while the primary database is unavailable, the recovery process must decide whether to replay those messages, discard them, or reconcile them against restored data. If object storage is current but the relational database is restored to an earlier point, users may see files whose metadata no longer exists. Recovery planning should therefore define consistency expectations across components rather than treating each backup independently.

A useful exercise is to walk one critical business transaction from entry to completion and identify every durable or semi-durable artifact it creates. Then ask how each artifact is recovered to a mutually consistent point. That analysis often reveals that the hardest recovery problem is not the largest database; it is the set of loosely coordinated state changes around it.

Recovery exercises should also record the time spent waiting on people, approvals, quota increases, external vendors, or security exceptions. Those delays are part of real RTO even though they do not appear in a backup console. If a restore technically completes in twenty minutes but production access waits three hours for a certificate, firewall change, or executive approval, the service did not meet a twenty-minute recovery objective.

Related Posts

• How Attack Paths Form Across Enterprise Systems

• Start With Risk When Choosing Security Controls

• Azure RBAC: Separate Scope From Role

• Why Azure VNets Fail: Address Spaces, Routes, and DNS

• Azure Backup and Site Recovery Protect Against Different Failures

• NSGs, ASGs, and Azure Firewall: Put the Control in the Right Place

• Subnetting Gets Easier When You Stop Memorizing Tables

• DHCP and DNS: Two Services That Make Everything Else Look Broken

• REST APIs for Network Engineers Who Grew Up on the CLI

• Containers or Lambda? Operations Usually Decides