Troubleshoot an Azure VM Before You Redeploy It
A broken virtual machine creates pressure to do something dramatic. When RDP or SSH stops working, an application disappears, or the Azure portal reports a failed state, redeploying the VM can look like the fastest reset button. Sometimes it is the right move. But redeploying too early can erase useful evidence, introduce avoidable downtime, and leave the original fault untouched if the real problem lives inside the guest operating system, the network path, or the application itself.
The better habit for an Azure administrator is to narrow the failure domain first. Microsoft’s current AZ-104 blueprint still expects administrators to manage compute, networking, identity, security, and monitoring together. A VM incident usually crosses several of those layers. The task is not to memorize every possible repair command; it is to prove which layer failed before choosing a recovery action.
That means starting with symptoms and state, then moving from the Azure control plane toward the guest. If the evidence eventually points to the physical host or to a stuck platform state, reapply or redeploy can be appropriate. If the evidence points elsewhere, a host move only adds disruption.
Start by classifying what “broken” actually means
A VM that will not boot is different from a VM that boots but cannot be reached, and both are different from a healthy VM whose application is failing. Treating all three as the same “VM problem” is how troubleshooting becomes random.
First look at the resource’s provisioning and power state in Azure. If the resource itself shows a failed provisioning state, that points toward a control-plane or recent-operation problem. If Azure reports the VM as running, ask whether boot diagnostics show a usable operating system. If the guest appears healthy, move outward to network reachability and then upward to the application.
This sequence preserves cause and effect. A stopped web service on an otherwise healthy VM should not trigger a redeploy. A kernel or boot-loader error should not begin with firewall changes. And a route or NSG problem will not be fixed by rebuilding the guest.
That layered thinking is central to the work behind the Azure Administrator role: resource state, compute state, network state, and guest state are related, but they are not interchangeable.
Use the activity log to understand what changed
A useful incident question is “what happened immediately before the symptom?” Azure Activity Log records control-plane operations such as starts, stops, redeployments, disk changes, network-interface updates, extension operations, and many configuration changes. A failed operation often contains an error message that is more actionable than the top-level portal status.
If the VM stopped responding after a resize, extension deployment, policy change, or network-interface edit, that history narrows the search. A generic “failed” label matters less than the operation that placed the resource in that state.
Change context also protects administrators from treating coincidence as cause. A VM may become unreachable after an unrelated deployment while the real issue is an expired application certificate or a full OS disk. The activity log is one source of evidence, not an automatic verdict, but it gives the investigation a timeline instead of a guess.
Boot diagnostics can tell you whether the guest ever became healthy
When a VM does not appear reachable, boot diagnostics are one of the fastest ways to distinguish platform availability from guest-operating-system failure. Azure can expose console output and a screenshot from the hypervisor. Those artifacts can show a login screen, filesystem errors, boot loops, recovery prompts, or an operating system that never completed startup.
A login screen is important evidence: the VM is probably booting, so the investigation should shift toward network access, credentials, services, or host firewall rules. A boot error points the other direction, toward disk and operating-system repair.
For severe Windows boot or disk failures, Azure’s VM repair workflow can copy the OS disk, attach the copy to a repair VM, allow offline remediation, and then restore the repaired disk. That is far more targeted than assuming a host move will repair damaged boot configuration or corrupted files.
Linux incidents follow the same reasoning even though the guest tools differ. Console output, serial access, filesystem state, startup services, and cloud-init or agent logs can reveal a guest problem that redeployment cannot solve.
Prove the network path before blaming the VM
A running VM can be completely healthy while every user experiences it as down. Network security groups, user-defined routes, Azure Firewall, load balancers, public IP changes, DNS, peering, VPN paths, and the guest firewall can all interrupt access.
Start with the exact flow that is failing: source, destination, protocol, port, and the address the client actually resolved. Then inspect effective security rules and effective routes on the VM’s network interface. Network Watcher tools can help determine the next hop and whether Azure sees a path that matches the design.
When the incident is really about virtual-network behavior rather than compute, the deeper AZ-700 networking scope becomes relevant. That does not mean every administrator needs to turn a VM incident into an architecture project. It means the troubleshooting branch should follow the evidence.
A common mistake is to “fix” access by opening a broader NSG rule or attaching a public IP. That may restore connectivity while creating a security problem and hiding the original fault. Diagnose the intended path before creating a new one.
Check guest agents and extensions as their own failure domain
Azure VM extensions depend on the guest agent and on the VM being able to process the desired state supplied by the platform. An extension can fail while the operating system and application continue to work. Conversely, an extension problem can be the only visible symptom of a guest agent that is unhealthy.
Look at extension status and, when possible, extension logs inside the VM. Microsoft documents VM Reapply as a way to trigger the platform to reapply the VM’s state and deliver a new goal state. Reapply is different from redeploy: it does not intentionally move the VM to a new host.
That distinction matters. If the resource is stuck because the latest model or provisioning state did not apply cleanly, reapply is a more proportionate action than moving the workload. If the extension itself is misconfigured, fix the extension. If the guest agent is damaged, repair the agent or guest dependency.
This is why broad Azure administration is operational work rather than a sequence of portal buttons: the same symptom can originate from resource configuration, the guest agent, the network, or the application.
Reapply is a state repair, not a general reboot
When an Azure VM is stuck in a failed state after an operation, Microsoft recommends reapplying the VM state before redeploying in applicable scenarios. Reapply tells the Azure platform to reapply the resource model. It is useful when the desired configuration and the platform’s current state are out of sync.
Administrators should still treat reapply as a change. It can cause short downtime, and Microsoft notes that rare cases can trigger a pending update that requires a restart. Use it because the evidence points to provisioning state, not because it sounds safer than diagnosis.
If reapply clears a failed provisioning state and the guest remains unreachable, that is new evidence: the control-plane state improved but the incident remains. Continue the investigation instead of immediately escalating to redeploy.
Redeploy solves a narrower problem than many people assume
Redeploy shuts down the VM, moves it to a new Azure host node, and powers it back on. That is valuable when the current host is suspected, when platform-level connectivity remains abnormal after other checks, or when Microsoft troubleshooting guidance specifically recommends a host move.
It does not reinstall the operating system. It does not repair an application configuration. It does not fix a bad route, an incorrect DNS record, a guest firewall rule, or a full filesystem. The associated Azure resources and configuration are retained, which is why a configuration problem often follows the VM to the new host.
Redeploy also has consequences. Microsoft warns that data on the temporary disk and ephemeral disk can be lost, and dynamic IP addresses associated with the virtual network interface can change. The operation therefore belongs after evidence collection, not before it.
A useful decision test is simple: if moving the exact same VM model and disks to another physical node would plausibly remove the suspected cause, redeploy is relevant. If not, keep troubleshooting the actual layer that failed.
Preserve evidence before recovery changes the environment
Incidents become harder to understand after several well-intentioned fixes. A restart clears transient state. A redeploy moves the host. Disk repair changes files. Firewall changes alter the path. If every action is performed before evidence is recorded, the team may restore service without learning what happened.
Capture the provisioning state, boot diagnostics, relevant activity-log events, effective routes, security rules, extension states, disk health indicators, and application symptoms before disruptive changes. For recurring systems, that evidence is often more valuable than the individual fix because it improves monitoring and runbooks.
Teams that treat reliability as part of Azure optimization eventually move from heroic recovery toward measurable operating signals: they know what healthy boot, network reachability, disk usage, application response, and platform state look like before an incident.
The best VM troubleshooting ends with a smaller failure domain
A strong investigation does not need to identify every possible cause at once. It needs to remove whole categories of causes with evidence. If boot diagnostics show a healthy login screen, stop treating the incident as a boot failure. If the resolved IP and effective route are correct, move past basic routing. If reapply restores provisioning but not application access, separate platform state from guest behavior.
For candidates preparing around Microsoft certifications, that reasoning is more durable than memorizing a list of recovery buttons. Azure changes, portal labels move, and individual commands evolve, but the operational sequence remains useful: define the symptom, identify the layer, collect evidence, apply the least disruptive action that matches the evidence, and verify the result.
Redeployment is an important Azure recovery tool. It is simply more effective when it is the conclusion of troubleshooting rather than the beginning.