Microsoft AZ-104: Designing Recovery with Azure Backup
Azure Backup design starts with recovery requirements, not with creating a vault. The workload team needs to define what must be recoverable, how much data loss is acceptable, how quickly restores must complete, which failures the backup protects against, who can authorize destructive changes, and how recovery continues if the primary region or administrator account is compromised. Azure Backup then becomes one part of a broader recovery architecture alongside application-native replication, snapshots, Site Recovery, and business runbooks.
Microsoft’s current Azure Backup guidance emphasizes vault redundancy, soft delete, immutable vaults, multi-user authorization, Cross Region Restore, RBAC, locks, encryption, and restore testing. These controls solve different failure classes. Redundancy protects backup data availability; soft delete helps recover from deletion; immutability blocks destructive changes; MUA adds authorization separation for critical operations; restore drills prove the business can actually use the recovery points.
Backup design therefore belongs inside Azure Architecture in Practice.
Start with RPO and RTO
Recovery point objective defines how much data the business can lose; recovery time objective defines how long the service can remain unavailable.
Backup and Site Recovery protect against different failures, so teams should not assume a daily backup meets an application that needs minute-level failover.
Use replication for availability and rapid failover where required, and backup for recoverable historical copies, corruption, deletion, ransomware, and longer-term retention.
Choose the correct vault type
Azure Backup uses Recovery Services vaults and Backup vaults for different supported workloads.
Select the vault according to the data source and current Azure Backup support rather than standardizing one vault type for every workload.
Platform teams should publish supported vault patterns and avoid creating one enterprise vault so large that ownership, RBAC, and failure blast radius become difficult to manage.
Choose redundancy from recovery scope
Locally redundant, zone-redundant, and geo-redundant options protect backup data against different infrastructure failures where supported.
Microsoft currently recommends ZRS when backup availability must survive a zone failure while keeping data in-region, and GRS when cross-region recovery requirements justify replication to the paired region.
Failure-domain design should align vault redundancy with the failures the workload expects Backup to survive.
Use Cross Region Restore where regional recovery matters
Cross Region Restore can let eligible GRS vaults restore supported workloads in the paired region without waiting for Microsoft to declare a disaster.
This can support drills, audit requirements, and recovery from a regional outage.
CRR has workload-specific RPO and support characteristics, so teams should test the exact restore scenario rather than treat “GRS” as proof that recovery will meet the application target.
Protect recovery points from deletion
Soft delete keeps deleted backup items or vaults recoverable for a configurable period according to current service behavior.
Immutable vaults can block operations that would remove recovery points and can be locked so the immutability decision is irreversible.
These controls are valuable against operator error and destructive attacks, but they should be configured before an incident and included in normal change governance.
Use multi-user authorization for critical operations
Multi-user authorization can require Resource Guard approval for selected sensitive Backup operations.
This creates separation between the backup administrator and the principal authorized to approve critical changes.
Azure RBAC should keep backup administration, security oversight, and resource ownership separate where the business requires stronger protection against compromised privileged accounts.
Use policies aligned to workload retention
Backup policies define frequency and retention for supported resources.
Use application and compliance requirements rather than one universal schedule across every workload.
Long retention improves historical recovery but increases cost and data lifecycle obligations, while short retention can leave the business unable to recover a slow-moving corruption discovered weeks later.
Test restores, not just backup jobs
A successful backup job proves data was captured, not that the service can be restored correctly.
Run periodic restore drills into isolated resources or a recovery environment, validate application state, permissions, network access, secrets, and dependencies, and measure actual recovery time.
Document the commands and approvals required so the recovery process does not depend on one administrator’s memory.
Integrate backup with incident response
Ransomware or destructive-administration incidents can involve both production and backup control planes.
Monitor changes to vault settings, immutability, soft delete, policies, and privileged roles as security-relevant events.
For AZ-104 and AZ-305, durable recovery design is workload requirement → backup policy → vault redundancy → deletion protection → authorization separation → restore drill → incident integration.
Backup architecture should also document dependencies required during recovery. DNS, identity, Key Vault, networking, private endpoints, subscriptions, and quotas can all prevent a technically successful restore from becoming a usable application. A recovery drill should follow the entire service flow rather than stop when the VM or database object appears.
Encryption design matters for recovery continuity. Customer-managed keys can strengthen control, but the key vault and permissions must remain available during restore. Protecting backup data with a key that the disaster scenario also destroys creates a hidden recovery dependency.
Vault locks and Azure resource locks should be applied with operational understanding. Locks can reduce accidental deletion but can also affect legitimate maintenance. Document the authorized procedure for changing protected recovery infrastructure and test it before emergency work is required.
Backup cost should be monitored by retention, protected instance, storage tier, and restore-testing strategy. Recovery controls are intentional cost, but duplicated policies and forgotten protected items can create waste after workloads are retired.
The recovery program is mature when the business can answer which restore point would be used, who can approve it, how long recovery takes, which region it runs in, what data could be lost, and when the process was last proven successfully.
Backup architecture should distinguish operational restore from disaster recovery. A developer who deletes one database row needs a different recovery path from a region outage or compromised administrator. Keep recovery scenarios explicit—single file, application item, entire VM, database point-in-time, subscription-scale recovery, or cross-region restore—and make sure the selected Azure Backup feature actually supports the required granularity.
Recovery Services vault and Backup vault placement should follow ownership and failure-domain strategy. A vault can protect many resources, but centralizing too much can make RBAC, policy, reporting, and blast radius difficult to manage. Separate critical business units or environments where independent administration or recovery authority is required, while avoiding unnecessary proliferation that makes monitoring fragmented.
Soft delete settings should be treated as part of the threat model. If ransomware or an attacker with backup permissions deletes recovery points, the retained soft-deleted state can provide time to respond. Operators need to know the retention period, how to undelete protected items, and which actions remain possible while data is in soft-delete state.
Immutable vaults can strengthen destructive-operation resistance, especially after the immutable setting is locked. Because a locked immutable configuration is intentionally difficult or impossible to reverse, test backup and restore workflows before locking the control. Security controls should protect recovery without creating an operational dead end the team has never rehearsed.
Multi-user authorization is especially useful for organizations that want backup administrators to operate routine jobs without being able to weaken critical recovery protections alone. Resource Guard should have separate ownership and permissions from the protected vault. Periodically test the approval path so the second control is available during a real incident.
Cross-subscription restore or alternative recovery subscriptions can reduce blast radius when the source subscription is compromised or unavailable, where supported by the workload. The recovery subscription should have network, identity, quota, policy, and security foundations ready enough that restored resources can become usable quickly.
Azure Policy can help standardize backup enrollment, vault settings, and retention expectations across subscriptions. Policy-driven coverage is useful only when exceptions are visible and remediation failures have owners. A compliance dashboard that shows a VM as noncompliant without a team responsible for onboarding it does not create recoverability.
Restore testing should use realistic data validation. For a VM, confirm boot, application startup, DNS, identity, monitoring, and dependency connectivity. For a database, run integrity checks and application queries. For file data, validate permissions and representative content. “Restore job succeeded” is only the first stage of recovery validation.
Recovery runbooks should identify whether the application must remain isolated before reconnecting to production. During a cyber incident, immediately restoring into the normal network can re-expose a recovered system to the same attacker or compromised dependency. Clean-room or isolated recovery patterns may be appropriate for high-risk workloads.
Backup monitoring should detect missed jobs, stale recovery points, policy changes, vault-deletion attempts, disabled protection, immutability changes, and restore failures. Alert severity should reflect the workload’s recovery target so a failed backup on a critical database receives different urgency from one missed snapshot on a disposable dev VM.
Retention should be aligned with legal and business requirements. Long-term retention can support audit and historical recovery, but keeping sensitive data beyond its purpose creates privacy and storage cost. Backup lifecycle belongs in data-governance review, especially when production deletion requirements do not immediately remove retained recovery points.
Finally, Azure Backup should be integrated with architecture decision records. Document why the chosen redundancy, policy, retention, soft-delete, immutability, and MUA settings meet the workload’s risk. When the workload moves region, changes database technology, or updates RTO/RPO, review the backup architecture as part of the same change.