Building a DORA Exit Strategy for Azure: What Article 28(8) Requires
A DORA exit strategy for Azure is proven by your dependency graph, not by the document. Portability assessment, exit triggers, timeline, cost, and the risk-accepted conclusion.
Building a DORA Exit Strategy for Azure: What Article 28(8) Requires
Most Azure exit strategies I read were written to be filed, not executed.
You can tell within a page. The document describes a migration in the abstract, in the passive voice, with no service names in it. Nobody who has ever moved a workload wrote it. A supervisor with any cloud experience can smell that too, and once they do, everything else in your third-party file gets read with a raised eyebrow.
Here is my position, and it is the uncomfortable one. For most financial institutions, exit from Azure for a critical function is technically feasible and operationally prohibitive. That conclusion, documented and costed and accepted by the board, is compliant. Article 28(8) requires the plan for any arrangement supporting critical or important functions. It does not require you to believe you will ever run it.
So stop treating it as a migration plan. It is a risk assessment that happens to produce one.
The value lives in the assessment.
What Article 28(8) actually asks for
Article 28(8) wants a documented exit strategy for any ICT third-party arrangement supporting a critical or important function. The point of it: you can leave without disrupting business activities, breaching a regulatory requirement, or degrading service continuity. Article 30(3)(f) then lists the exit strategy among the mandatory contractual provisions, so it lands in the contract too.
The requirement has a specific shape. You must be able to leave the arrangement and either bring the function in-house or move it to another provider. Without breaking business continuity or your regulatory obligations. It has to account for a transition period where the incumbent and the successor both run at once. And it has to read as credible. Credible enough that a supervisor believes you could execute it under pressure, not just that you drafted it.
That last word is where the whole thing lives. Credible.
A supervisor cannot audit your intentions. They can audit your dependency graph. Which is why, for Azure, the work is not writing the document at all: it is working out what it would actually take to lift a workload off Microsoft’s platform, service by service, and then letting the document fall out of the answer.
This is the other side of the third-party register and concentration analysis in DORA Article 28 and Microsoft as a critical third party. The register documents the concentration risk. The exit strategy is your answer to it.
Portability, in three layers
Split the assessment into three layers: data egress, infrastructure-as-code portability, and platform-as-a-service lock-in. They carry different difficulty and different cost. Keep them separate, because they fail separately.
Start with data egress. Can you get your data out, in a usable format, inside the transition window? Azure no longer charges egress on data leaving via the standard cross-region paths the way it once did across the board, and EU Data Act provisions are eroding egress fees further. Which means the money is no longer the interesting part. Volume and format are what actually bite. Petabytes of blob storage move on a timescale measured in weeks, not hours. For the genuinely large sets there is Azure Data Box, which ships the data physically, on a truck, like it is 2009. Record the volume, the format, the mechanism, and a duration you would defend in a room.
Then infrastructure-as-code. And here is where I usually get the pushback, so let me perform it rather than paraphrase it.
“We are on Terraform, so we are portable.”
No. That is syntax, not semantics. If your environment lives in Azure Resource Manager (ARM) templates or Bicep, it is Azure-specific by construction and nobody argues otherwise. Terraform with the AzureRM provider merely looks better. An azurerm_kubernetes_cluster does not become an AWS EKS cluster because you swapped the provider block. The real question is how much of your IaC describes generic infrastructure versus Azure-native services, and the honest answer is nearly always the same: the compute is portable, the managed services are not.
That brings you to platform-as-a-service. The hard layer. Managed services trade portability for operational simplicity, and you do not get both.
| Azure service | Lock-in driver | Portability difficulty |
|---|---|---|
| Azure SQL Database | T-SQL plus Azure-specific features (elastic pools, serverless tiers) | Moderate. Migrates to SQL Server, harder to non-Microsoft engines |
| Azure Cosmos DB | Proprietary multi-model API, request-unit model | High. No drop-in equivalent |
| Azure Service Bus | Proprietary messaging semantics, sessions, dead-lettering | High. Requires re-architecture to an alternative broker |
| Microsoft Entra ID | Identity provider for the entire estate, federation, conditional access | Very high. Touches every application |
| Azure Kubernetes Service | Kubernetes itself is portable, AKS-specific add-ons are not | Low to moderate |
Out of this falls a per-service difficulty rating. That rating is the spine of everything downstream: the cost, the timeline, all of it. Get it wrong and every number after it is decoration.
Two of those rows do most of the work, and it is the same two every time.
Microsoft Entra ID, because identity is not a component of the estate. It is the thing every other component authenticates against. Move it and you are not migrating a service, you are re-plumbing every application, every service principal, every conditional access assumption and every federation you forgot existed. Most exit documents rate it as one line in a table. It is closer to a programme than a line.
And the hyperscaler itself. Not any individual service, but the accumulated assumption that Azure is there. The identity model, the network model, the deployment pipeline, the monitoring, the way your engineers think about a resource group. That dependency does not appear in an inventory because it is not a resource. It is the substrate.
Everything else on that table is an engineering problem with a price. Those two are the ones that turn twelve months into thirty.
Triggers and timeline
A plan with no trigger is a plan nobody executes. So define the triggers first: the specific, pre-agreed conditions that make the entity invoke the exit. Then build the timeline by summing the per-layer transition durations from the portability work. Both have to exist before anyone should call the plan credible.
For an Azure arrangement, the triggers usually cover sustained failure to meet contractual service levels, a material breach of the Data Processing Addendum, a supervisory direction from De Nederlandsche Bank (DNB) or the European Supervisory Authorities, a Union-level designation event under DORA Articles 31 to 44, an unacceptable shift in concentration risk, or a commercial event. A price change that breaks the business case counts, and people forget that one.
Name who decides on each, and on what evidence.
A trigger nobody owns is decoration.
Now add up the layers. Data egress runs to weeks. IaC rebuild on the target platform runs to months for anything non-trivial. PaaS re-architecture (replacing Azure Cosmos DB, re-platforming off Microsoft Entra ID) runs to quarters. Put it together and a realistic full-exit timeline for a critical function on Azure is rarely under twelve months and frequently over twenty-four. The cost estimate mirrors the same structure: egress cost, rebuild engineering cost, parallel-run cost during the transition, and the opportunity cost of the engineering capacity you burn getting there.
That last line is the one nobody wants to write down, and it is usually the biggest number on the page.
Why “operationally prohibitive” is a valid answer
DORA is a risk-management regulation. It does not tell you to avoid concentration. It tells you to understand it and govern it. So a risk that is assessed, quantified, and accepted at board level satisfies the framework. Article 28(8) asks for a documented and viable exit strategy, not a commitment to exit.
Do the portability assessment honestly and most institutions land in the same place. Leaving Azure for a critical function would take eighteen to thirty months, cost millions, and eat the engineering capacity that would otherwise build the business. None of that is a compliance failure. It becomes one only when it is undocumented. When nobody assessed it, costed it, or signed for it.
If you are the architect who gets handed this and told to “write the exit plan by Friday”, that is the sentence to put in front of whoever handed it to you. You are not being asked to leave Microsoft. You are being asked to prove you know what staying costs.
The compliant version has four parts:
- The portability assessment, showing you know exactly where the lock-in lives.
- The timeline and cost, showing you have quantified the exit.
- The triggers, showing you know when you would act despite the cost.
- The board sign-off, showing the residual risk was accepted by people with the authority to accept it.
Read that, and a supervisor does not see an institution pretending it will leave Microsoft. They see one that knows precisely what staying costs and has chosen, eyes open, to stay.
Same evidentiary posture runs through DORA backup and recovery (see DORA Article 12 on Azure), where documented capability beats theoretical perfection.
The part I got tired of doing by hand
I am not immune to any of this, by the way. My own first instinct, for years, was to reach for the document template rather than the estate, because the template is finishable on a Friday and the estate is not.
The bit that finally wore me down was the dependency map. Every exit assessment starts by working out which workloads actually depend on Azure SQL Database, Cosmos DB, Service Bus and Entra ID, and how those dependencies wire together, and every time I was rebuilding that by hand from a diagram that had gone stale two reorgs ago. So Crimson Owl Technologies built PAA to read it out of the running environment instead. It is read-only, it does not migrate anything, and the judgment about what to re-platform and what to accept stays where it belongs: with your architects.
I have done over a thousand risk assessments and health checks as a PFE at Microsoft, and something north of seven hundred startups and ISVs in Azure engineering after that. Rounding, because I stopped keeping count properly. Across all of it, the dependency map was the artefact nobody could produce on request. Not because it was hard. Because nobody owned it, and the diagram that was supposed to be it had stopped being true.
What I would take away from this
- Article 28(8) requires exit strategies for ICT arrangements supporting critical or important functions, and Article 30(3) makes the exit plan a mandatory contract element. Two obligations, one artefact.
- The plan rests on a portability assessment of three layers: data egress, infrastructure-as-code portability, and platform-as-a-service lock-in. They fail separately, so assess them separately.
- The hardest lock-in points are the managed services: Azure SQL Database, Azure Cosmos DB, Azure Service Bus, Microsoft Entra ID.
- Triggers must be defined in advance, with a named owner each: provider failure, contract breach, supervisory direction, unacceptable concentration, commercial event.
- “Exit is feasible but operationally prohibitive, risk accepted at board level” is a valid DORA position, provided it is documented, costed, and signed.
- And the one underneath all of them: the document is downstream of the dependency graph. If you cannot derive it from what is actually deployed, you have written fiction with a version number on it.
Frequently asked questions
Does DORA require us to actually be able to leave Azure? It requires a viable, documented exit strategy, not a demonstrated migration. The strategy has to be credible (assessed, costed, governed), but Article 28(8) accepts that exit from a hyperscaler is a major undertaking. A risk-accepted decision to remain, backed by a real assessment, is compliant. An absent or fictional plan is not.
How detailed must the portability assessment be? Detailed enough to support a credible timeline and cost. So: a per-service lock-in rating for the managed services your critical function depends on, the data volume and egress mechanism, and an honest split of which components are portable and which need re-architecture. “We could move if needed” satisfies nobody.
Who has to approve the exit strategy? The management body. Article 28(1) puts ICT third-party risk under board-level responsibility, and accepting residual concentration risk because exit is prohibitive is a call for the body with authority to make it. Sign-off from an architecture team alone does not cover a critical-function arrangement.
Should the exit strategy assume multi-cloud as the destination? Not necessarily. The destination can be another provider, an in-house build, or a different region. Multi-cloud carries its own concentration and complexity costs, and naming it as the target without assessing those costs weakens the plan. Name a realistic destination and assess the move to that one specifically.