Azure Well-Architected Review with Remediation Sprint
This is a fixed-price, three-week engagement that reviews one production Azure workload against all five pillars of Microsoft's Azure Well-Architected Framework — reliability, security, cost optimization, operational excellence and performance efficiency — and then fixes the top of the list instead of only writing it down. Two weeks of review: Microsoft's own Azure Well-Architected Review assessment (about sixty questions, scored per pillar and in aggregate) run properly against the named workload, with Azure Advisor recommendations pulled in for the agreed subscription or resource group, Microsoft Defender for Cloud findings, monitoring, backup and cost evidence, and architecture interviews with the people who actually build and operate it. Then a remediation sprint capped at five engineer-days that implements the highest-value items you approve in writing — availability-zone and backup gaps, missing diagnostic settings and alert rules, right-sizing, network hardening, cost quick wins — inside your own change process. You finish with a scored baseline you own, a prioritized backlog with effort, risk and cost impact on every item, the approved fixes applied and evidenced, and the residual backlog costed and sequenced. $4,950 fixed per workload; multi-workload estates are quoted per estate in writing before we start. No promised score and no promised savings: those numbers come out of your environment, not off this page.
What this engagement is
Microsoft's Azure Well-Architected Framework organizes workload design around five pillars — reliability, security, cost optimization, operational excellence and performance efficiency — and Microsoft publishes a matching self-assessment, the Azure Well-Architected Review, on Microsoft Assessments. It runs to roughly sixty questions drawn from the framework's key recommendations, scores each pillar you answer plus an aggregate for the workload, can pull in the Azure Advisor recommendations for the subscription or resource group you point it at, exports the recommendations for your own tracking, and carries a milestone feature so a later run can be measured against the first as a baseline. It also contains several sub-assessments — workload-type variants sit alongside the core review — and Microsoft publishes a maturity-model track as well, with a maturity model for each of the five pillars. Choosing the wrong one is where a do-it-yourself run usually goes wrong: pick a workload-type variant by accident, or point the assessment at a subscription holding three unrelated systems, and the score that comes back is about nothing in particular. We select deliberately, record which assessment was used and on what date, and save the milestone under your sign-in so the baseline is yours. The review phase is two weeks and it is not a questionnaire exercise. We complete the assessment for one named workload with the people who build and operate it in the room, because the honest answer to 'do you test your recovery path' only ever surfaces in conversation. Around it we assemble the evidence a questionnaire cannot produce: Azure Advisor's current recommendations and category scores for the workload's scope; Microsoft Defender for Cloud's secure score and security recommendations; the resource topology as deployed rather than as diagrammed; identity, secrets and network paths; Azure Monitor coverage — what has diagnostic settings, what has alert rules, what nobody would hear about at three in the morning; the backup and recovery configuration measured against the recovery objectives you would actually be judged on; and the workload's recent Microsoft Cost Management data. Tool output is triaged, not forwarded — a raw Advisor export is a list of things, and half of them will not apply to you. Everything that survives lands in one backlog, mapped to its pillar and scored for value, effort, implementation risk and cost impact. Then the part that separates this from a report: a remediation sprint of up to five engineer-days, inside the same fixed fee, spent implementing the items you approve. You see a change list first — each item with what changes, the blast radius, the rollback and whether it needs a maintenance window — and nothing is touched until you sign it off. The work that fits this shape is the work that keeps showing up in real estates: closing an availability-zone or redundancy gap where the SKU allows it, turning on backup for the resources that never had it and proving one restore, adding diagnostic settings and the handful of alert rules that would have caught the last incident, right-sizing over-provisioned compute and databases, removing public endpoints and tightening network security groups, moving secrets into Key Vault and connections onto managed identities, deleting orphaned disks, public IPs and stale snapshots, and applying tags and budgets so the next cost spike arrives with a name attached. Items that need application code changes, a re-architecture, or a maintenance window your business will not grant stay on the backlog with an estimate against them. An honest residual list is worth more than a heroic one. Two straight answers about the alternatives. The assessment itself is free — you can take the Azure Well-Architected Review yourself today, and you should if that is all you need; we will point you at it on the discovery call. Microsoft also runs its own Azure Expert Assessment, which Microsoft describes as a no-cost review by Azure Certified Experts covering cost management, security and reliability, with customers self-nominating through an online portal; eligibility and terms are Microsoft's and they change, so check with your Microsoft account team, and if you can get it, take it. What this engagement adds is that the review is run for you against one bounded workload rather than an environment, from evidence rather than self-grading, across all five pillars, and that the highest-value fixes are implemented inside the engagement under your change control instead of joining a backlog nobody has time for. Where the findings run past one workload, the follow-on work is separate and named: estate-wide cost governance is the Azure Cost Optimization and FinOps Assessment, a proper subscription and governance foundation is the Azure Landing Zone and Cloud Adoption Framework implementation, continuous multi-subscription posture management is Microsoft Defender for Cloud CSPM, and keeping the resources monitored and patched afterwards is Azure Resource Monitoring and Maintenance.
Success criteria
What you receive
How the work unfolds
The workload is named and bounded in writing, including which shared services are dependencies rather than review targets. The Azure scope — subscription or resource groups — is agreed, read-only access is provisioned through your own identity governance, and the correct Well-Architected Review sub-assessment is opened under your sign-in. Advisor, Defender for Cloud and Cost Management data are pulled for the agreed scope so the evidence base exists before the first interview.
Working sessions with the people who build and operate the workload, structured by pillar: how it fails and what happens next, how identity and secrets work, what the network path really is, what is monitored and who gets woken, what is backed up and whether a restore has ever been proved, where the money goes. In parallel we record the deployed topology, diagnostic and alert coverage, backup configuration and cost profile. Disagreements between what the tools show and what the team believes are recorded, not smoothed over.
The assessment questions are answered from evidence rather than optimism, and the per-pillar and aggregate scores are recorded as the baseline milestone. Advisor and Defender for Cloud recommendations are triaged into real findings, noise and not-applicable, with the reasoning kept. Each surviving finding is written up with the evidence attached and the pillar it belongs to.
Every finding is scored for value, effort, implementation risk and cost impact, and the backlog is sequenced. We propose the sprint candidates — the items with the best value-to-risk ratio that genuinely fit five engineer-days — each with its change plan, blast radius and rollback, and present the list for your decision. You choose what gets touched; choosing fewer items, or none, is a legitimate answer.
Approved items are implemented highest-value-first inside your change process and maintenance windows, one at a time, each verified and evidenced with before-and-after state. Where something can be tested — a restore, a failover, an alert path — we test it rather than assume it. Write access is scoped to the agreed resources and to the sprint window. When the cap is reached we stop; anything left is estimated and returned to the backlog.
The closing package lands: scorecard and baseline, findings by pillar, the evidence pack for every change made, the residual backlog costed and sequenced, and a management summary that names the risks worth funding next. We walk your engineers through the detail and your leadership through the summary, hand over the assessment milestone, and remove our access.
Prerequisites
Who does what
IT Partner
- Run Microsoft's Azure Well-Architected Review against the named workload with the correct sub-assessment selected, and record the scores as a baseline milestone under your own sign-in.
- Assemble and reconcile the supporting evidence — Advisor, Defender for Cloud, Azure Monitor coverage, backup configuration, deployed topology and Cost Management data — and triage tool output into real findings rather than forwarding a raw export.
- Interview the people who build and operate the workload across all five pillars, and record where the evidence and the team's belief disagree.
- Produce a prioritized backlog in which every item states the action, the effort, the implementation risk and the cost impact.
- Propose sprint items with change and rollback plans, implement only what you approve in writing, work highest-value-first, test what can be tested, and stop at the five-engineer-day cap.
- Use least-privilege access throughout, keep write access scoped to the agreed resources and the sprint window, evidence every change with before-and-after state, and remove our access at the end.
- Deliver the closing report, cost and sequence the residual backlog honestly — including items we are not the right people to do — and walk your team through all of it.
Your team
- Name the workload, agree its boundary, and confirm the Azure scope in writing before work starts.
- Provide the read-only access for the review and the scoped write access for the sprint, and revoke both when we ask you to at the end.
- Make the builders and operators available for interviews in the first week, and answer follow-up questions in days rather than weeks.
- State your availability and recovery objectives, your change-approval process and the maintenance windows the sprint can use.
- Decide in writing which backlog items the sprint implements — the choice is yours, including the choice to implement none of them.
- Accept that some findings will point at decisions your own team made, and that the report will say so plainly rather than diplomatically.
What's not included
Limitations & technical notes
Frequently asked questions
What is the Azure Well-Architected Framework, and who decided on these five pillars?
Microsoft did. The Azure Well-Architected Framework is Microsoft's own guidance for designing and operating workloads on Azure, organized into five pillars: reliability, security, cost optimization, operational excellence and performance efficiency. Microsoft publishes the framework, the per-pillar guidance, the maturity models and the assessment that scores against it. We are not selling you our private opinion of what good architecture looks like — we are running Microsoft's, thoroughly, against your workload, and then fixing the top of the resulting list.
The Well-Architected Review is a free self-assessment. Why would we pay for it?
Take it yourself if that is all you need — it is free on Microsoft Assessments and we will point you at it on the discovery call. What the fee buys is the difference between a questionnaire and a review. The assessment is about sixty questions answered by whoever fills in the form, and it is worth precisely as much as the honesty of those answers; teams reliably grade themselves generously on the questions they are least sure about. We answer them from evidence instead — deployed topology, Advisor, Defender for Cloud, monitoring and backup configuration, cost data — with the people who run the workload in the room to be challenged. Then five engineer-days of the findings get implemented. The self-assessment produces a score. This produces a shorter backlog.
Microsoft offers a free Azure Expert Assessment. Why pay you instead?
Take Microsoft's if you can get it. Microsoft describes the Azure Expert Assessment as a no-cost review by Azure Certified Experts covering cost management, security and reliability, with customers self-nominating through an online portal; eligibility and terms are Microsoft's and they change, so confirm the current position with your Microsoft account team. Three honest differences. It is a review with recommendations — the hands-on remediation in your subscription is still yours to schedule, whereas here five engineer-days of it happen inside the engagement. It covers three pillars in Microsoft's own description; this covers all five, including operational excellence and performance efficiency. And it runs to Microsoft's queue and eligibility rather than to your date. If both are available to you, the sensible move is to take Microsoft's first and bring us the output — we build on it rather than repeat it, and more of the three weeks goes to fixing things.
What counts as 'one workload'?
A workload is one system that delivers a business capability and is operated as a unit — the customer portal, the claims pipeline, the data platform behind reporting — including its compute, data, networking, identity and operational tooling. In practice it is usually a resource group or a small set of them. It is not 'our Azure tenant' and not 'everything in the production subscription'. We agree the boundary in writing on day one, including which shared services (a hub network, the Microsoft Entra ID tenant, a central Log Analytics workspace) are dependencies we take as given rather than things we review. If you need several workloads reviewed we quote per estate, and if you are not sure how many you have, that is a good use of the free discovery call.
What actually happens in the five-day remediation sprint?
Approved, low-blast-radius, high-value fixes get made in your environment, under your change process. Typical items: closing an availability-zone or redundancy gap where the SKU supports it, enabling backup on resources that never had it and proving one restore, adding diagnostic settings and the specific alert rules that would have caught your last incident, right-sizing over-provisioned compute and databases, removing public endpoints and tightening network security groups, moving secrets into Key Vault and connections onto managed identities, clearing orphaned disks, public IPs and stale snapshots, and applying tags and budgets. Every item is proposed with its change plan, blast radius and rollback, and nothing is touched until you approve it in writing. What does not fit this shape — code changes, re-architecture, anything needing a window your business will not grant — stays on the backlog with an estimate.
What if the fixes need more than five engineer-days?
Then they do not happen in the sprint, and you know that at the point of approval rather than at the end. Every backlog item carries an effort estimate before you choose, so the list you approve is one that fits, and we work highest-value-first inside it. Anything that does not fit is estimated, sequenced and priced separately — another sprint, a scoped project, or work for your own engineers. The cap is what keeps the fee fixed and the promise honest. An uncapped 'we will fix whatever we find' would have to be either a much larger number or a much smaller amount of fixing.
Will you make changes in production? How is that controlled?
Only changes you have approved in writing, one item at a time, inside your change process and your maintenance windows. Each proposed change comes with what it touches, its blast radius, the rollback and whether it needs a restart or a window. The review itself needs read-only access; write access is granted for the sprint window, scoped to the agreed resources, and revoked at the end. Where a change cannot be made safely in production without a rehearsal, we say so and it stays on the backlog. Nothing in this engagement requires you to accept an unreviewed change.
Will our Advisor score or secure score go up?
Likely, on the items that map to them — but we do not promise a number and no target score appears in the quote. Azure Advisor scores your resources in the same five categories as the framework, calculated largely from the proportion of healthy resources against the applicable ones, so fixing real findings tends to move it. The score is a proxy, though: it responds to resource state, and some of the most valuable findings in a review — a recovery path nobody has tested, a dependency nobody documented, an alert nobody would answer — move no score at all. We would rather fix the thing than farm the metric, and anyone guaranteeing you a score is telling you which of the two they are doing.
Do the fixes increase our Azure bill?
Some of them do, and the backlog tells you which before you approve anything. Zone redundancy, backup, longer log retention, private endpoints, premium tiers and additional Defender for Cloud plans all cost money — that is what buying reliability and security looks like. Others go the other way: right-sizing, idle and orphaned resources, storage tiering. Every item carries its cost direction. And Microsoft's meters are Microsoft's: Azure consumption, Defender for Cloud plans, Log Analytics ingestion and backup storage are billed by Microsoft to your subscription, not by us. Our fee is fixed and does not vary with what the changes cost or save.
How is this different from your Azure cost optimization services?
It answers a different question. The Azure performance and cost optimization assessment is a short, low-cost review of an Azure environment for efficiency that hands you recommendations and changes nothing. The Azure Cost Optimization and FinOps Assessment goes wider and deeper on money across the estate and sets up the practice that keeps it optimized. This engagement asks whether one workload is well built across all five pillars — cost is one fifth of it — and ends with approved fixes applied rather than recommendations delivered. If your only problem is the bill, buy one of the cost services instead: they will serve you better and cost less, and we would rather say that here than after you have paid.
How is this different from Defender for Cloud CSPM?
Microsoft Defender for Cloud CSPM is a continuous, estate-wide security posture program — plans enabled, recommendations triaged on an ongoing basis, compliance standards tracked over time. This is a point-in-time review of one workload across all five pillars, of which security is one, and we read your existing Defender for Cloud data rather than standing the program up. They fit together: a review that finds real security debt in one workload frequently ends with a recommendation to run CSPM properly across the estate, and the report will say so where the evidence points that way, even though it is a different engagement with a different fee.
We do not have a proper landing zone. Should we fix that first?
Often, yes — and the review will tell you honestly which order to work in. If the workload's problems are mostly foundational — no subscription topology, no policy baseline, no consistent identity or network model — a Well-Architected review will keep producing findings that share one root cause, and the right answer is the Azure Landing Zone and Cloud Adoption Framework implementation first. But a workload that grew organically inside an imperfect environment can still be made materially safer and cheaper in three weeks, and that is often what a team needs before it can fund the larger program. Describe the situation on the discovery call and we will tell you which we would do first, including when the answer is 'not this one yet'.
What access do you need, and for how long?
For the review: Reader on the workload's subscription or resource groups, Cost Management Reader for the spend data, and a Defender for Cloud reader role for the security findings — all read-only, all through your own identity governance, and just-in-time elevation through Microsoft Entra ID is welcome. For the sprint only: contributor-level rights scoped to the agreed resources, granted for the sprint window and revoked at the end. We do not need and will not ask for Owner or User Access Administrator on your subscription, and if someone offers them we will decline in writing. Everything we see is treated as confidential.
Is three weeks realistic?
For one bounded workload, yes — provided the interviews happen in week one. Scope, evidence, assessment and backlog take the first two weeks; the sprint runs in the third alongside your change windows; the closing report and walkthrough land at the end. The thing that stretches the timeline is not the analysis, it is approvals — if the change list waits a week for a change advisory board, the sprint compresses. That is why we book the sprint window with you at kickoff, and where your process genuinely cannot approve inside the window we agree a later sprint date in writing rather than rushing changes into production.
What do we keep at the end, and can we take it elsewhere?
All of it, and yes. The assessment milestone is saved under your own Microsoft Assessments sign-in, so you can re-run the review next quarter and measure the movement without us. The findings report, the backlog with its estimates, the change and rollback plans and the before-and-after evidence pack are working files handed over with nothing proprietary holding them. Run the residual backlog with your own engineers or another provider if you want to — that is a legitimate outcome, and the report is written to be used rather than to be re-sold.