Azure Site Recovery Disaster Recovery Implementation
Implementation of Azure Site Recovery — Microsoft's disaster-recovery-as-a-service — for on-premises VMware, Hyper-V, or physical servers replicating to Azure, or for Azure VMs replicating between regions. IT Partner deploys the Recovery Services vault and replication components, configures replication policies and network mapping, brings every in-scope server to a healthy replicated state, builds an ordered recovery plan, executes one test failover in an isolated network — production untouched — and hands over a documented failover runbook. $150 per server plus a $2,500 tenant fee as the working estimate; the final quote is fixed, in writing, after scoping. Azure consumption charges are Microsoft's, billed separately to your subscription.
What this engagement is
Backup answers one question: can we get the data back? Disaster recovery answers a harder one: can the business keep running while we do? If your servers are protected by backup alone — and for most of the companies we meet, they are — a burned-out host, a flooded server room, or a dead hypervisor means rebuild first, restore second: hours to days of downtime even when every backup is perfect. Azure Site Recovery closes that gap. It keeps a continuously updated replica of each protected server in Azure, and when the primary fails, the replica boots as an Azure VM in minutes — no standby datacenter to buy, and no compute charges until the day you actually fail over. This service takes you from zero to a proven DR capability. We design the topology for your platform — on-premises VMware or physical servers through the replication appliance, Hyper-V through the Site Recovery provider on your hosts, or Azure VMs replicating region-to-region — then deploy the Recovery Services vault, configure replication policies (recovery-point retention and app-consistent snapshot frequency), and map your networks to Azure: which virtual network and subnets servers fail over into, how IP addressing and DNS behave, and a separate isolated network for testing. Servers are grouped into recovery plans with an agreed boot order, because a domain controller that comes up after the application that depends on it is not a recovery. The recovery-time and recovery-point targets we design to are set with you and stated in the design as configured targets with measured test results — never as guarantees, because nobody who is being honest guarantees a disaster. The engagement ends the right way: with one executed test failover into the isolated network, your application owners validating that what booted actually works, the measured results written down, and a failover runbook handed to your team. What happens after that is deliberately out of scope — ongoing DR drills, replication-health monitoring, and alerting are operational services (see Managed Backup and Backup-Restore and Azure Resource Monitoring and Maintenance), and the written plan that should wrap around this technology is its own discipline (see Business Continuity and Disaster Recovery Plan Development). One thing we will insist on: Site Recovery complements backup, it does not replace it — replication faithfully copies bad changes as well as good ones, so Azure Backup stays in the picture.
Success criteria
What you receive
How the work unfolds
Server inventory, platforms, dependencies, data churn, and bandwidth are assessed; recovery-time and recovery-point targets are agreed with you; network mapping is designed. Output: the design document and the fixed written quote confirming the estimate.
Recovery Services vault, replication appliance or providers and agents, connectivity, and least-privilege access are deployed and verified.
Replication policies are applied and initial replication runs — scheduled around your bandwidth so production traffic is not starved — until every in-scope server reaches a healthy, current replicated state.
Failover networks, subnets, IP and DNS behavior, and the isolated test network are configured; servers are grouped into recovery plans with the agreed boot order.
A test failover is executed into the isolated network. Your application owners sign in and validate; we document what worked, what needed adjustment, and the measured time to a running environment. Production is never touched.
The failover runbook is finalized with the test findings folded in, walked through with your team, and handed over together with our recommended drill cadence.
Prerequisites
Who does what
IT Partner
- Design the DR topology and document the configured RTO/RPO targets with their rationale.
- Deploy the vault, replication infrastructure, policies, and network mapping.
- Bring every in-scope server to healthy replication and keep you informed of initial-replication progress.
- Build the recovery plans and execute the test failover, documenting measured results honestly — including anything that did not work on the first try.
- Deliver the failover runbook and the Azure consumption outline.
- Hand over cleanly, with our access removed or converted to a support arrangement you choose.
Your team
- Provide the Azure subscription, platform access, and appliance capacity.
- Approve the RTO/RPO targets and the network design — these are business decisions; we model them, you own them.
- Approve change windows for agent and provider installation on production hosts.
- Make application owners available to validate during the test failover — a boot screen is not a validation; a signed-in user is.
- Own Microsoft's consumption billing for the replication footprint and any failover compute.
- Own the DR drill schedule after handover, or engage us separately to run it.
What's not included
Limitations & technical notes
Frequently asked questions
We already use Azure Backup. Why would we need Site Recovery too?
Because they answer different questions. Backup preserves point-in-time copies you restore from — typically hours to days to a running server, since restoring means provisioning something to restore onto. Site Recovery keeps a continuously updated replica that boots in Azure in minutes. Backup protects you from data loss, deletion, and ransomware history; replication protects you from downtime. Most businesses that can quantify the cost of a day offline end up wanting both — and we will tell you plainly if, for your workloads, backup alone is actually enough.
Does the test failover disrupt production?
No — that is the point of doing it properly. The test failover boots replica VMs in an isolated Azure network with no connection to production; source servers keep running and replicating throughout. Your application owners sign in to the failed-over copies and confirm the applications genuinely work, we document the measured recovery time, and then the test environment is cleaned up. A DR implementation that has never been failed over is a hypothesis, not a capability.
Are the RTO and RPO guaranteed?
No, and be wary of anyone who says yes. RTO and RPO are targets we configure and then measure: the replication policy determines how much data you could lose (typically minutes for healthy replication), and the test failover measures how long a recovery actually takes with your recovery plan. Both numbers go in the design document and the test report as configured targets with measured evidence. Microsoft publishes its own SLA for the Site Recovery service itself; a real disaster's variables belong to nobody.
What does Site Recovery cost to run after the project?
Microsoft bills it to your Azure subscription, separately from our fee. As of this writing, Microsoft's published pricing charges per protected instance per month — and makes each instance's first 31 days free, which conveniently covers the implementation and test period. On top of the instance charge come replica storage, storage transactions, and egress; compute charges apply only while failed-over VMs actually run, which is what makes DR-to-Azure so much cheaper than a standby datacenter. We give you a written consumption outline, and you verify current rates with Microsoft at contracting — it is their meter.
Which platforms can you protect?
On-premises VMware VMs and physical Windows or Linux servers (via the replication appliance), Hyper-V VMs (via the Site Recovery provider on your hosts), and Azure VMs replicating to another region. Exact operating-system versions, disk configurations, and churn rates are confirmed against Microsoft's current support matrix during design — we commit to what the matrix supports on the day we design, not to folklore.
What counts as a 'server' in the $150-per-server price?
Each protected instance in scope — one VM or one physical server enabled for replication. Ten servers is $150 × 10 plus the $2,500 tenant fee, which covers the vault, replication infrastructure, network mapping, recovery plans, the test failover, and the runbook. It is an estimate until discovery is done; then you get one fixed written quote, and you pay after you approve delivery.
Will replication slow down our production servers or saturate our internet connection?
Designed properly, no — and bandwidth design is part of the engagement. Initial replication is the heavy lift, so we schedule and, where the platform supports it, throttle it around business hours. Ongoing delta replication is usually modest, but 'usually' is not a plan: we assess your data churn and link capacity during design and tell you before enabling anything whether your connection can carry the estate honestly.
Does Site Recovery protect us against ransomware?
Only partially, and we will not oversell it. Replication is faithful — an encrypted server replicates its encrypted state within minutes. Retained recovery points let you fail over to a point shortly before the damage, which can genuinely help, but retention on replication is measured in hours to days, not the weeks or months of history real ransomware recovery often needs. Backup with proper retention is the ransomware control; Site Recovery is the downtime control. Run both.
Who runs DR after you hand over — and how often should we test?
You do, with the runbook: it covers test, planned, and unplanned failovers, roles, and failback. Our standing recommendation is a test failover at least annually — more often after significant infrastructure change — because replication health today does not prove recoverability next year. If you would rather not own that discipline internally, our Managed Backup and Backup-Restore and Azure monitoring services can carry the drills and the replication-health watching as an ongoing arrangement, priced separately.
Can we fail back after a failover?
Yes — failback is part of the design and the runbook, and the mechanics depend on the platform. Azure VMs re-protect and replicate back to the original region. VMware and Hyper-V estates replicate back to on-premises infrastructure once it is healthy. The honest caveat: under Microsoft's current support matrix, a physical server that failed over to Azure fails back to a VMware VM rather than to bare metal — if bare-metal restoration matters to you, we design for that reality up front instead of surprising you during a disaster.
Our servers are already in Azure. Is this still relevant?
Yes — a region is a blast radius too. Azure-to-Azure Site Recovery replicates VMs to a secondary region with the same vault, policy, recovery-plan, and test-failover discipline, and the same engagement structure applies. During design we also look at whether zone redundancy inside your current region covers your actual risk more cheaply — the right answer is sometimes the smaller one, and we will say so.
How long does the implementation take?
The typical engagement is about two weeks: design and infrastructure in the first, replication stabilization, the test failover, and handover in the second. The honest variable is initial replication — large data volumes on modest bandwidth extend the middle of the project, and we tell you the realistic timeline in the same written quote as the price.