First page of Microsoft's 100,000-partner directory, sorted by responsiveness Microsoft Solutions Partner — Security, Modern Work, Infrastructure, App Innovation Microsoft partner since 2006 1,100+ organizations under management
Home/Services/Azure Site Recovery Disaster Recovery Implementation
Implementation

Azure Site Recovery Disaster Recovery Implementation

Implementation of Azure Site Recovery — Microsoft's disaster-recovery-as-a-service — for on-premises VMware, Hyper-V, or physical servers replicating to Azure, or for Azure VMs replicating between regions. IT Partner deploys the Recovery Services vault and replication components, configures replication policies and network mapping, brings every in-scope server to a healthy replicated state, builds an ordered recovery plan, executes one test failover in an isolated network — production untouched — and hands over a documented failover runbook. $150 per server plus a $2,500 tenant fee as the working estimate; the final quote is fixed, in writing, after scoping. Azure consumption charges are Microsoft's, billed separately to your subscription.

Timeline 2 weeksService owner Roman SotnikAzure Site RecoveryMicrosoft Azure

What this engagement is

Backup answers one question: can we get the data back? Disaster recovery answers a harder one: can the business keep running while we do? If your servers are protected by backup alone — and for most of the companies we meet, they are — a burned-out host, a flooded server room, or a dead hypervisor means rebuild first, restore second: hours to days of downtime even when every backup is perfect. Azure Site Recovery closes that gap. It keeps a continuously updated replica of each protected server in Azure, and when the primary fails, the replica boots as an Azure VM in minutes — no standby datacenter to buy, and no compute charges until the day you actually fail over. This service takes you from zero to a proven DR capability. We design the topology for your platform — on-premises VMware or physical servers through the replication appliance, Hyper-V through the Site Recovery provider on your hosts, or Azure VMs replicating region-to-region — then deploy the Recovery Services vault, configure replication policies (recovery-point retention and app-consistent snapshot frequency), and map your networks to Azure: which virtual network and subnets servers fail over into, how IP addressing and DNS behave, and a separate isolated network for testing. Servers are grouped into recovery plans with an agreed boot order, because a domain controller that comes up after the application that depends on it is not a recovery. The recovery-time and recovery-point targets we design to are set with you and stated in the design as configured targets with measured test results — never as guarantees, because nobody who is being honest guarantees a disaster. The engagement ends the right way: with one executed test failover into the isolated network, your application owners validating that what booted actually works, the measured results written down, and a failover runbook handed to your team. What happens after that is deliberately out of scope — ongoing DR drills, replication-health monitoring, and alerting are operational services (see Managed Backup and Backup-Restore and Azure Resource Monitoring and Maintenance), and the written plan that should wrap around this technology is its own discipline (see Business Continuity and Disaster Recovery Plan Development). One thing we will insist on: Site Recovery complements backup, it does not replace it — replication faithfully copies bad changes as well as good ones, so Azure Backup stays in the picture.

Success criteria

01A Recovery Services vault and the replication components for your platform (appliance, provider and agents, or the Azure VM extension) are deployed and healthy.
02Every in-scope server shows healthy replication with recovery points inside the policy you approved.
03Replication policies — recovery-point retention and app-consistent snapshot frequency — match the approved design.
04Network mapping is complete: failover virtual network, subnets, IP addressing approach, DNS behavior, and an isolated test network that cannot touch production.
05Recovery plans exist with the agreed boot order and dependency groups.
06One test failover has been executed in the isolated network, your application owners have validated the results, and the measured time to a running environment is documented.
07The failover runbook is delivered and walked through with the people who would use it at 2 a.m.
08You have a written estimate of the Azure consumption the replication footprint will generate — Microsoft's meter, modeled honestly.

What you receive

DR design document — scope, platform topology, replication approach, configured RTO/RPO targets, network mapping, and the IP/DNS plan.
Configured Recovery Services vault with the platform-appropriate replication infrastructure deployed (replication appliance for VMware/physical, provider and agents for Hyper-V, Site Recovery extension for Azure VMs).
Replication policies configured to the approved retention and snapshot-frequency design.
Network mapping including the isolated test network.
Recovery plans with boot order and dependency groups, plus scripted steps where scoped.
One executed test failover with a written results report: what booted, what your application owners validated, issues found, and the measured recovery time.
Failover runbook — how to run a test, planned, or unplanned failover, who does what, how decisions get made, and the failback path for your platform.
Azure consumption outline for the replication footprint — Site Recovery per-instance charges, storage, and expected failover compute, so Microsoft's bill is not a surprise.
Handover session with your team.

How the work unfolds

Discovery and design

Server inventory, platforms, dependencies, data churn, and bandwidth are assessed; recovery-time and recovery-point targets are agreed with you; network mapping is designed. Output: the design document and the fixed written quote confirming the estimate.

Vault and replication infrastructure

Recovery Services vault, replication appliance or providers and agents, connectivity, and least-privilege access are deployed and verified.

Enable replication

Replication policies are applied and initial replication runs — scheduled around your bandwidth so production traffic is not starved — until every in-scope server reaches a healthy, current replicated state.

Network mapping and recovery plans

Failover networks, subnets, IP and DNS behavior, and the isolated test network are configured; servers are grouped into recovery plans with the agreed boot order.

Test failover

A test failover is executed into the isolated network. Your application owners sign in and validate; we document what worked, what needed adjustment, and the measured time to a running environment. Production is never touched.

Runbook and handover

The failover runbook is finalized with the test findings folded in, walked through with your team, and handed over together with our recommended drill cadence.

Prerequisites

An Azure subscription — Site Recovery licensing, storage, and any failover compute are billed by Microsoft to it, separately from our fee.
Administrative access for your platform: vCenter/ESXi for VMware, the Hyper-V hosts, or local administrator/root on physical servers — plus appropriately scoped Azure permissions for us.
Capacity to host the replication appliance where the platform needs one (a VM or server on the source side for VMware and physical estates).
Internet or ExpressRoute bandwidth adequate for initial and ongoing replication — we assess this during design rather than discovering it during replication.
The OS and workload list for in-scope servers, so supportability is confirmed against Microsoft's current Site Recovery support matrix at design time.
A decision-maker for the recovery-time and recovery-point targets, and application owners available for the test-failover validation window.
If failing back to on-premises hardware matters to you, say so at design — failback mechanics differ by platform and shape the architecture.

Who does what

IT Partner

  • Design the DR topology and document the configured RTO/RPO targets with their rationale.
  • Deploy the vault, replication infrastructure, policies, and network mapping.
  • Bring every in-scope server to healthy replication and keep you informed of initial-replication progress.
  • Build the recovery plans and execute the test failover, documenting measured results honestly — including anything that did not work on the first try.
  • Deliver the failover runbook and the Azure consumption outline.
  • Hand over cleanly, with our access removed or converted to a support arrangement you choose.

Your team

  • Provide the Azure subscription, platform access, and appliance capacity.
  • Approve the RTO/RPO targets and the network design — these are business decisions; we model them, you own them.
  • Approve change windows for agent and provider installation on production hosts.
  • Make application owners available to validate during the test failover — a boot screen is not a validation; a signed-in user is.
  • Own Microsoft's consumption billing for the replication footprint and any failover compute.
  • Own the DR drill schedule after handover, or engage us separately to run it.

What's not included

Ongoing DR operations — scheduled drills, replication-health monitoring, and alert response after handover are operational services; see Managed Backup and Backup-Restore and Azure Resource Monitoring and Maintenance.
Executing a real failover during an actual disaster — the runbook is written so your team can, and we can assist under a separate support engagement, but standing emergency response is not part of this fixed scope.
Backup implementation — Site Recovery does not replace backup. Server backup is its own service: Azure Backup for physical or virtual servers or the combined server-and-client variant.
The written business continuity and disaster recovery plan around the technology — recovery objectives, roles, communications, and a tested plan document are delivered by Business Continuity and Disaster Recovery Plan Development.
Multi-region application re-architecture — active-active designs, SQL Server availability groups spanning regions, and global load balancing are architecture engagements, not replication configuration.
Azure consumption — Microsoft's charges for Site Recovery licensing, storage, storage transactions, egress, and compute while failed-over VMs run are billed by Microsoft to your subscription, entirely separate from our service fee.
Migrating workloads to Azure permanently — if the real goal is leaving the datacenter rather than protecting it, start with the Azure Migrate datacenter discovery and assessment or go straight to Windows Server migration to Azure.

Limitations & technical notes

!RTO and RPO are configured targets, never guarantees. We design to the targets you approve and the test-failover report states what was actually measured; Microsoft publishes its own SLA for the Site Recovery service. Anyone guaranteeing recovery times for a real disaster is selling you adjectives.
!Replication copies change — including bad change. A ransomware-encrypted server replicates its encrypted state within minutes. Retained recovery points give limited rewind, but backup with proper retention remains the ransomware control, which is why we treat Site Recovery as a complement to Azure Backup, never a substitute.
!Initial replication time depends on data volume and bandwidth. Multi-terabyte estates on modest links take days to reach a protected state, and we schedule around your production traffic rather than pretending physics away.
!Supportability is governed by Microsoft's current Site Recovery support matrix — operating-system versions, disk sizes, and per-disk data-churn limits. High-churn workloads such as busy database servers get explicit design attention; we verify against Microsoft's published matrix at design time rather than assuming.
!Failback mechanics differ by platform. Under Microsoft's support matrix as of this writing, a physical server that has failed over to Azure fails back to a VMware virtual machine, not to bare metal — if that matters to you, it shapes the design, and we raise it before work begins.
!Azure consumption is Microsoft's meter and separate from our fee. As of this writing Microsoft's published pricing bills Site Recovery per protected instance per month (with each instance's first 31 days free under Microsoft's own pricing terms), plus storage, storage transactions, egress, and compute whenever failed-over VMs run — including during test failovers. We model it in the consumption outline; verify current rates with Microsoft at contracting.
!The $150 per server + $2,500 tenant fee is an estimate. Unusual topologies, very high churn, multi-site sources, or many recovery plans move the number; after discovery you receive one fixed quote in writing, and that number is the number.

Frequently asked questions

We already use Azure Backup. Why would we need Site Recovery too?

Because they answer different questions. Backup preserves point-in-time copies you restore from — typically hours to days to a running server, since restoring means provisioning something to restore onto. Site Recovery keeps a continuously updated replica that boots in Azure in minutes. Backup protects you from data loss, deletion, and ransomware history; replication protects you from downtime. Most businesses that can quantify the cost of a day offline end up wanting both — and we will tell you plainly if, for your workloads, backup alone is actually enough.

Does the test failover disrupt production?

No — that is the point of doing it properly. The test failover boots replica VMs in an isolated Azure network with no connection to production; source servers keep running and replicating throughout. Your application owners sign in to the failed-over copies and confirm the applications genuinely work, we document the measured recovery time, and then the test environment is cleaned up. A DR implementation that has never been failed over is a hypothesis, not a capability.

Are the RTO and RPO guaranteed?

No, and be wary of anyone who says yes. RTO and RPO are targets we configure and then measure: the replication policy determines how much data you could lose (typically minutes for healthy replication), and the test failover measures how long a recovery actually takes with your recovery plan. Both numbers go in the design document and the test report as configured targets with measured evidence. Microsoft publishes its own SLA for the Site Recovery service itself; a real disaster's variables belong to nobody.

What does Site Recovery cost to run after the project?

Microsoft bills it to your Azure subscription, separately from our fee. As of this writing, Microsoft's published pricing charges per protected instance per month — and makes each instance's first 31 days free, which conveniently covers the implementation and test period. On top of the instance charge come replica storage, storage transactions, and egress; compute charges apply only while failed-over VMs actually run, which is what makes DR-to-Azure so much cheaper than a standby datacenter. We give you a written consumption outline, and you verify current rates with Microsoft at contracting — it is their meter.

Which platforms can you protect?

On-premises VMware VMs and physical Windows or Linux servers (via the replication appliance), Hyper-V VMs (via the Site Recovery provider on your hosts), and Azure VMs replicating to another region. Exact operating-system versions, disk configurations, and churn rates are confirmed against Microsoft's current support matrix during design — we commit to what the matrix supports on the day we design, not to folklore.

What counts as a 'server' in the $150-per-server price?

Each protected instance in scope — one VM or one physical server enabled for replication. Ten servers is $150 × 10 plus the $2,500 tenant fee, which covers the vault, replication infrastructure, network mapping, recovery plans, the test failover, and the runbook. It is an estimate until discovery is done; then you get one fixed written quote, and you pay after you approve delivery.

Will replication slow down our production servers or saturate our internet connection?

Designed properly, no — and bandwidth design is part of the engagement. Initial replication is the heavy lift, so we schedule and, where the platform supports it, throttle it around business hours. Ongoing delta replication is usually modest, but 'usually' is not a plan: we assess your data churn and link capacity during design and tell you before enabling anything whether your connection can carry the estate honestly.

Does Site Recovery protect us against ransomware?

Only partially, and we will not oversell it. Replication is faithful — an encrypted server replicates its encrypted state within minutes. Retained recovery points let you fail over to a point shortly before the damage, which can genuinely help, but retention on replication is measured in hours to days, not the weeks or months of history real ransomware recovery often needs. Backup with proper retention is the ransomware control; Site Recovery is the downtime control. Run both.

Who runs DR after you hand over — and how often should we test?

You do, with the runbook: it covers test, planned, and unplanned failovers, roles, and failback. Our standing recommendation is a test failover at least annually — more often after significant infrastructure change — because replication health today does not prove recoverability next year. If you would rather not own that discipline internally, our Managed Backup and Backup-Restore and Azure monitoring services can carry the drills and the replication-health watching as an ongoing arrangement, priced separately.

Can we fail back after a failover?

Yes — failback is part of the design and the runbook, and the mechanics depend on the platform. Azure VMs re-protect and replicate back to the original region. VMware and Hyper-V estates replicate back to on-premises infrastructure once it is healthy. The honest caveat: under Microsoft's current support matrix, a physical server that failed over to Azure fails back to a VMware VM rather than to bare metal — if bare-metal restoration matters to you, we design for that reality up front instead of surprising you during a disaster.

Our servers are already in Azure. Is this still relevant?

Yes — a region is a blast radius too. Azure-to-Azure Site Recovery replicates VMs to a secondary region with the same vault, policy, recovery-plan, and test-failover discipline, and the same engagement structure applies. During design we also look at whether zone redundancy inside your current region covers your actual risk more cheaply — the right answer is sometimes the smaller one, and we will say so.

How long does the implementation take?

The typical engagement is about two weeks: design and infrastructure in the first, replication stabilization, the test failover, and handover in the second. The honest variable is initial replication — large data volumes on modest bandwidth extend the middle of the project, and we tell you the realistic timeline in the same written quote as the price.

Didn’t find your question?

Ask it here. A real engineer answers by email within one business day — and if it’s a good one, it becomes part of this page so the next person finds it.

Answered by a person, one time, to your inbox. Nothing you type here is published without a human reviewing and anonymizing it first.

Often combined with

$150 per server + $2,500 tenant fee
2 weeks
Scope my DR implementation