DORA Broke First. Wardley's Org Model Is Next.

DORA Broke First. Wardley’s Org Model Is Next.

Two standards bodies are moving in opposite directions on the same question right now, and almost nobody managing a MedTech engineering org has connected them. On 16 December 2025 the European Commission published a targeted revision to MDR, proposing a default Class I designation for certain software with narrower escalation criteria, explicitly to correct what the Commission itself now calls systematic up-classification of software devices relative to their actual patient risk. The comment period closed 18 March 2026. In parallel, IEC SC62A has been developing Edition 2 of IEC 62304, which collapses the current three-tier Class A/B/C system into two process-rigor levels and, critically, assigns Level II (the higher tier) automatically to any software that implements a risk control measure, regardless of what external mitigations might otherwise justify a lower classification. One regulator is trying to loosen software classification generally. One standard is trying to tighten it for exactly the software that used to escape tightening through classification tricks.

That contradiction is not a footnote. It’s the precise shape of a problem I’ve been circling for a while: the metrics and org models that unregulated SaaS exports to MedTech assume something about cost that regulated software doesn’t grant. DORA assumes every deployment costs roughly the same to ship, so deployment frequency is a virtue. Simon Wardley’s explorers, villagers, and town planners model (EVTP), which he wrote up again in December 2023, assumes every team handoff costs roughly the same to execute as a component matures, so fluid reorganization around evolution stage is a virtue. Both assumptions die at the same fence: a Design History File that doesn’t care how mature or commodity your component is, only how much patient risk it carries.

Two Metrics Systems, One Assumption They Both Get Wrong

DORA’s four keys, established by Forsgren, Humble, and Kim’s research and now the default vocabulary in every engineering org I’ve hired from, are deployment frequency, lead time for changes, change failure rate, and time to restore service. Three of the four implicitly reward speed. That’s correct in a context where a deploy is a git push and a rollback. It stops being correct the moment a deploy requires a release record, a Design History File update, and a risk file review before it can ship, because now the constrained resource isn’t your CI pipeline, it’s the quality function’s review capacity. Optimizing deployment frequency in that environment doesn’t make you faster. It makes your quality reviewers the bottleneck and your DORA dashboard a lie about where the actual cycle time lives.

Wardley’s EVTP model has the same shape of problem, from the organizational-design side rather than the metrics side. His argument, laid out across six numbered steps in the blog post, is that as a component evolves from genesis to custom-built to product to commodity, the right team size and the right methods change with it: three to five people and high tolerance for failure at the genesis stage, something closer to a two-pizza team (his number, twelve people) once a component is industrialized. Crucially, he proposes a mechanism he calls theft: villager teams declare they’re taking over an explorer’s project once it’s ready to productize, and town-planner teams do the same to villagers once a product is ready to commoditize. Ownership moves fluidly, tracking the evolution curve. The whole system depends on that handoff being cheap.

In a regulated org it isn’t cheap. It’s the single most expensive event in the software lifecycle. Transferring ownership of a component that lives inside a Design History File triggers a class-of-change assessment, forces re-verification of every interface that component touches, and can trigger a new risk analysis or a technical file amendment depending on jurisdiction. Wardley’s theft is a Tuesday afternoon reorg. Yours is a documented, auditable, cost-bearing regulatory event, and pretending otherwise is how organizations end up with an EVTP structure on paper and a frozen, ossified component ownership map in practice.

What Wardley Actually Gets Right for Regulated Orgs

Before tearing the rest of it down, credit where it’s due. Wardley’s biggest political fight, described at length in his piece, is convincing an organization to accept a mandatory map-and-challenge gate before any project above a spend threshold gets built. He calls it the Intelligence Function. Teams resist it, executives grumble, and he warns you need roughly a year before it runs smoothly and a second year before it’s trusted enough to start spotting duplication across the org.

Regulated MedTech shops don’t have to fight that battle. Design control already is an Intelligence Function, mandated by ISO 13485 and IEC 62304, with no opt-out and no political capital required to install it. Every project above a triviality threshold already gets challenged before build starts: does the design input match a documented user need, has risk been assessed, does the architecture map to a classified safety level. That’s a structural advantage over the unregulated orgs Wardley built his model from, and it’s worth saying plainly to a board that assumes compliance is pure drag: you already have, for free, the single hardest organizational capability Wardley spent two years trying to install elsewhere.

The risk is what you do with that gate once it exists. Wardley’s other principle, laid out through his appropriate-methods diagrams, is that different evolution stages need different rigor: heavyweight, contract-driven process for commodity components, lightweight and exploratory process for genesis-stage ones. Most regulated shops apply one rigor level to everything that passes through design control, regardless of where the component actually sits on the evolution curve. That’s the exact failure Wardley calls out when he tells you to fire the consultants pushing “let’s Six Sigma everything.” It’s also, not coincidentally, the exact failure the MDR targeted revision is trying to correct at the regulatory level: redirect notified body scrutiny toward novel, higher-risk technology instead of spreading it evenly across devices with long, boring clinical histories. The regulator arrived at Wardley’s appropriate-methods argument independently, through its own political process, three years after he published his.

Where the Model Breaks: Theft vs the Design History File

Here’s where I stop being generous to Wardley. The theft mechanism is the load-bearing wall of EVTP, the thing that makes the whole structure “self-organize” as he puts it, and it assumes reassignment friction near zero. Regulated software makes that friction the most expensive line item you have.

Say a villager team wants to take over a genesis-stage clinical algorithm once it’s proven enough to productize, exactly the move Wardley’s model wants you to make. Under a functioning QMS, that handoff isn’t a conversation and a Slack channel rename. It’s a formal transfer of DHF ownership. Every interface the new team touches gets re-verified. If the component performs a risk control function, the transfer likely also triggers a risk management file update, because your risk file names an owner and an assumption set, and both just changed. Depending on how the classification shakes out, you may be looking at a technical file amendment before you can ship the next version under the new owner at all.

IEC 62304 Edition 2 makes this worse, not better, at exactly the point where you’d want relief. Right now, teams have some room to argue a component down from Class C to Class B by pointing at external risk-acceptance mechanisms, hardware interlocks, or clinical workflow controls that reduce the software’s own criticality on paper. Edition 2 removes that room for anything implementing a risk control measure: it’s Level II, full stop, regardless of what sits around it. Combine that with theft-based reassignment and you get a bad outcome. A genesis-stage algorithm, three people, high tolerance for failure, exactly Wardley’s Explorer profile, gets assigned maximum process rigor from day one if it happens to implement a risk control function, which a lot of clinical decision support does. Wardley’s model says treat it loosely because it’s immature. The regulation says treat it strictly because it’s risky. Those are different axes, and EVTP has no vocabulary for the second one.

What Replaces DORA

Deployment frequency and lead time for changes aren’t useless under design control, they’re measuring the wrong segment of the pipeline. The fix isn’t a new four-metric package, it’s redrawing where the clock starts and stops.

Time from design input to first Intelligence Function decision, not time to production, is the number that reflects your actual constrained resource. If your quality function takes three weeks to render a first opinion on a change classification, that’s your real lead time, and shipping fast downstream of that decision doesn’t change it.

DHF-impact ratio: the percentage of changes classified as requiring no Design History File update versus those that do. A low ratio, most changes touch the DHF, tells you your architecture isn’t segmented by evolution stage the way Wardley’s appropriate-methods principle wants it to be. A component that’s stable and well-characterized shouldn’t force a DHF review every time you touch it. If it does, your map is wrong, not your process.

Risk file staleness: days since a component’s risk file was last reviewed, tracked against days since the component’s code last changed. A widening gap is the regulated-org equivalent of tech debt, and it’s invisible to every DORA dashboard I’ve ever seen deployed in a MedTech shop, because DORA doesn’t know the risk file exists.

Change failure rate survives mostly intact, but time to restore service needs to be redefined. In a regulated org you are not recovered when the service is back up. You’re recovered when the incident record, the CAPA if one was triggered, and the risk file all match what actually happened in production. Measuring only service restoration and calling it MTTR hides the tail of the incident, which is usually longer and involves more people than the technical fix.

The Two-Pizza Team With a Quality Lead

Wardley’s team-sizing guidance, three to five people at genesis, something around a two-pizza team once industrialized, is a reasonable starting point, but it doesn’t answer the question I get asked most by engineering leaders moving into MedTech: does every pod need its own embedded quality person.

No, and trying to do that is how you burn through your entire quality headcount budget on your smallest teams. Borrow the other half of Wardley’s structure instead, the part people skip past because the theft mechanism gets all the attention: guilds, cutting horizontally across every pool regardless of which pod someone currently sits in. A quality and regulatory affairs guild, staffed to the org’s actual risk classification profile rather than to a one-per-pod headcount rule, shows up at Intelligence Function gates the way Wardley’s guild of engineers shows up wherever engineering judgment is needed. A genesis-stage pod working on something Class A gets light-touch guild involvement. A pod maintaining a Level II risk-control component gets someone from the guild embedded closer to full time. The guild model scales with risk, not with org-chart symmetry, which is the only version of “two-pizza team with a quality lead” that survives contact with a real quality headcount budget.

Hiring Senior Engineers From Unregulated SaaS Without Losing Them in Week One

The failure mode I’ve watched most often: a strong senior engineer joins from an unregulated SaaS company, spends their first two weeks reading QMS SOPs and sitting through ISO 13485 training modules, and mentally checks out before they’ve shipped a single line of code. They came here to build things. You handed them a compliance syllabus.

Reverse the order. Put them on Intelligence Function-adjacent work in week one, mapping a component, sitting in on a design control review, seeing directly why a particular change needed a risk file update and another one didn’t. That’s an architecture problem, and senior engineers respect architecture problems. Pair them with someone from the quality guild rather than a training module, because a person who can explain a specific decision beats a slide deck that explains a general rule every time. Save the formal SOP training for once they’ve felt the constraint firsthand and have a real question to bring to it. The order matters more than the content: paperwork first teaches people that regulation is bureaucracy: context first teaches them it’s a second risk axis their existing SaaS instincts never had to model, which is closer to the truth and a lot more likely to keep them past month three.

Why This Is the VP-Level Argument

Shipping velocity metrics ported wholesale from unregulated SaaS are the fastest way to lose credibility with a board that has sat through a notified body audit. Being able to say precisely where a well-known, widely cited org design model breaks under regulatory load, backed by two dated, current regulatory changes rather than a general appeal to “compliance is different here,” is a different kind of signal. It says you understand both worlds well enough to know exactly where they stop agreeing with each other. That’s the argument I’d put in front of a board before I’d put a DORA dashboard in front of one, and after this year’s IEC 62304 and MDR moves, it’s an argument with a paper trail behind it.