Research · Blogs & Independent Thinkers

Back to sweep

Research sweep · deep · 2022 – 2026

Engineering Maturity Models for Regulated Financial Services

Which engineering maturity models and capability frameworks large financial services technology organisations use to drive an engineering hygiene uplift (automated releases, regression test automation, CI/CD, observability) and how they roll out, monitor and govern the programme at scale, September 2022 to September 2026: DORA capabilities and metrics, CMMI and TMMi, ITIL 4 change enablement, SAFe and Team Topologies, the CNCF Platform Engineering Maturity Model, engineering scorecards (Backstage, Cortex, OpsLevel), and regulator expectations from the EU Digital Operational Resilience Act, PRA SS1/21 and the FCA operational resilience regime

  • Claude Fable 5
  • tech
  • financial
  • academic
  • blogs
  • vc
  • frontier

Synthesised 2026-09-21

Narrative

Independent commentary on engineering maturity in 2022 to 2026 spends less time proposing new frameworks and more time correcting the record on the frameworks vendors and consultancies already sell. The clearest example is DORA itself: RedMonk analyst Rachel Stephens, writing on her own site across the 2024 and 2025 reports, documents that the elite/high/medium/low tier structure produced results the DORA team itself no longer trusts, including medium performers beating high performers on change failure rate in 2024, and that the 2025 report replaced the tiers entirely with seven team archetypes. Any financial services programme still setting an internal target of reaching 'elite' status is benchmarking against a model DORA has retired.

On governance, Team Topologies' own blog is unusually direct for a framework vendor, stating that Change Advisory Boards in regulated banking approved over 90 percent of major changes reviewed and that some firms never rejected a single change in a year, which it uses to argue automated peer-review governance is a better fit for both compliance and delivery speed than a manual gate. This is corroborated from the framework-critique side: Agile Manifesto co-author Ron Jeffries and independent Enterprise Agility writing both argue that SAFe, the operating model most commonly paired with this kind of governance overhaul, imposes top-down process that produces partial benefit while capping how far delivery practice can actually improve, and Team Topologies co-author Matthew Skelton has publicly disputed SAFe's marketing claim that Team Topologies is integrated into SAFe rather than merely referenced in two summary slides.

A second consistent thread is skepticism of measurement itself. Hillel Wayne's Computer Things newsletter and Kent Beck's Tidy First? Substack both interrogate what happens once an engineering metric becomes a target, echoing Goodhart's law concerns that predate DORA. This skepticism extends to commercial engineering-intelligence tooling: an independent buyer evaluation on tianpan.co comparing Jellyfish, LinearB, DX and Swarmia found LinearB's headline 'AI' feature is a rule-based YAML engine unable to distinguish AI from human contribution, and that Jellyfish's AI-generated executive summaries have shallow traceability back to source data, direct evidence for treating vendor AI-measurement claims as unverified until independently checked.

On delivery mechanics, Simon Willison's newsletter and practitioner accounts from IT Revolution converge on a specific 2025 to 2026 finding: AI coding tools amplify a team's existing capability pattern rather than fixing weak ones, meaning AI adoption in an already-hygienic division compounds its advantage while AI adoption in a manual-release division risks compounding instability. Independent dojo case studies, most notably Nerd/Noir's account of Target running dojo challenges for more than 70 teams, remain the clearest documented mechanism for a central function to uplift capability without owning the teams, consistent with how law-firm analysis of PRA SS1/21 frames the underlying regulatory requirement as evidence of governed, repeatable service resilience rather than any particular delivery methodology.


Sources

ID Title Outlet Date Significance
b1 DORA Report 2024 – A Look at Throughput and Stability RedMonk (rstephens) 2024-11 Independent analyst Rachel Stephens dissects the 2024 State of DevOps data and flags that medium performers beat high performers on change failure rate that year, undermining the idea that the elite/high/medium/low ladder is a stable target.
b2 DORA 2025: Measuring Software Delivery After AI RedMonk (rstephens) 2025-12 Follow-up piece explains DORA's 2025 replacement of the four-tier performance ladder with seven team archetypes, a methodological reversal directly relevant to any programme still benchmarking against 'elite' status.
b3 bliki: Team Topologies martinfowler.com 2020 Canonical independent framing of stream-aligned, platform, enabling and complicated-subsystem team types that underpins how large regulated organisations now design platform and enablement functions rather than project-based delivery.
b4 Engineering Maturity Model Fish Food for Thought (Mike Fisher, Substack) 2024 A named CTO-practitioner lays out a layered maturity model starting from deployment pipeline hygiene upward, explicitly framing maturity as continuous investment areas rather than discrete gated stages, a distinction the brief asks synthesis to track.
b5 The DX Core 4 Framework Fish Food for Thought (Mike Fisher, Substack) 2025 Independent engineering-leader commentary on the DX Core 4 outcome framework, useful for distinguishing outcome-metric frameworks from capability-level models as the brief requires.
b6 Agentic Engineering Patterns Simon Willison's Newsletter (Substack) 2026-02 Willison documents concrete coding-agent patterns (plan, implement, test, review loops) and argues throughput rises only when feedback loops are present, directly bearing on where AI coding tools help or worsen stability in 2026 programmes.
b7 How I use LLMs to help me write code Simon Willison's Newsletter (Substack) 2025 First-hand practitioner account of AI-assisted coding discipline (small steps, tests, review) that engineering-hygiene programmes are now trying to encode as guardrails for AI-generated changes.
b8 Developer Productivity Metrics: Education Necessary Software Design: Tidy First? (Kent Beck, Substack) 2024 Kent Beck's interview with DX's Abi Noda interrogates what productivity metrics can and cannot show; the piece is a sponsored/affiliated conversation with a metrics vendor and should be read as such alongside its substantive critique.
b9 Goodhart's Law in Software Engineering Computer Things (Hillel Wayne) 2023 Explains why even honestly pursued engineering metrics degrade the underlying goal once they become targets, the core theoretical objection any group-level scorecard or DORA-telemetry programme has to answer.
b10 Beyond the Hype: Building Actually Useful Platforms with the CNCF Maturity Model (and a Healthy Dose of Realism) blog.eisele.net 2025-02 Independent practitioner critique arguing the CNCF model offers a narrative progression rather than a multi-dimensional, measurable model, and warns that chasing higher maturity levels can produce over-engineered, under-used platforms.
b11 Announcing the Platform Engineering Maturity Model CNCF Blog 2023-11 Primary publication of the CNCF Platforms Working Group model referenced throughout the lane's independent commentary, including the working group's own caveat that pursuing maximum maturity can be actively detrimental.
b12 2024 DORA report summary DX Blog (Laura Tacho) 2024-11 Named practitioner (DX's CTO) summary of the 2024 State of DevOps findings; vendor-affiliated and should be treated as a claim pending independent corroboration per the lane's sourcing rules, but useful for cross-referencing RedMonk's independent reading of the same data.
b13 DORA Metrics: What the Research Actually Says zbmowrey.com 2024 Independent technical blog separating what the underlying DORA research actually supports from the benchmark-chasing behaviour the metrics have generated in enterprise adoption.
b14 Overview of DORA - DevOps Research and Assessment Bridge Apps (Substack) 2023 Independent explainer distinguishing DORA the research programme from DORA the four-keys metric set, a distinction synthesis needs given the EU's unrelated Digital Operational Resilience Act shares the same acronym.
b15 Demystifying Team Topologies in SAFe Pretty Agile 2023 Independent agile-coaching blog documents Team Topologies co-author Matthew Skelton publicly disputing SAFe's marketing claim that Team Topologies has been 'part of SAFe since version 5.1', evidence that a two-slide summary in Leading SAFe materially dilutes the source framework.
b16 Team Topologies Book Review, Summary, and Notes candost.blog 2022 Independent engineer's working notes on applying Team Topologies concepts, useful as a practitioner cross-check against the more marketing-oriented SAFe integration claims.
b17 DevOps Dojos at Extreme Scale Nerd/Noir 2023 Independent account of dojo-based coaching models used to uplift engineering practice across many teams simultaneously, the mechanism most directly relevant to a group function running an uplift programme without owning delivery teams.
b18 AI's Mirror Effect: How the 2025 DORA Report Reveals Your Organization's True Capabilities IT Revolution 2025-12 Practitioner-publisher analysis (IT Revolution, home of The DevOps Handbook) arguing AI adoption amplifies a team's existing capability pattern rather than fixing weak practices, directly relevant to whether AI tooling helps or worsens hygiene in lagging divisions.
b19 Measuring the Productivity of Developers Burkhard Stubert (Substack) 2024 Independent engineering-manager account testing productivity measurement approaches against day-to-day team reality, cross-referencing claims made by DORA- and SPACE-style frameworks.
b20 DORA 2025 for AI Teams: Seven Archetypes Every SAFe RTE Should Know Agility at Scale 2026 Independent SAFe-practitioner blog translates DORA's 2025 archetype model into scaled-agile release-train terms, showing how outcome-metric research is being retrofitted into level-based operating-model frameworks, a friction point the brief specifically asks about.
b21 Jellyfish vs LinearB vs DX vs Swarmia: What We Learned Evaluating Engineering Intelligence Platforms tianpan.co 2025 Independent buyer-side evaluation reporting that LinearB's 'AI' PR routing is a rule-based YAML engine that cannot distinguish AI from human contribution, and that Jellyfish's AI-generated executive narratives have shallow traceability to source data, direct evidence for the brief's instruction to treat engineering-intelligence vendor claims skeptically.
b22 UK Operational Resilience Rules: Are You Ready for 31 March 2025? Data Matters (Sidley Austin law firm blog) 2025-01 Independent legal-practitioner analysis of the PRA SS1/21 and FCA operational resilience deadline, setting out the end-to-end service resilience and impact-tolerance obligations that engineering maturity evidence has to map to.
b23 The Agile Expanse ronjeffries.com 2023 Agile Manifesto co-author Ron Jeffries argues SAFe imposes top-down process that delivers some benefit but caps how far an organisation's delivery practice can actually improve, a first-principles objection to using SAFe as the primary uplift vehicle.
b24 From Dark Scrum to Broken SAFe: some real problems of Agile-at-scale, and a way out Enterprise Agility (ea.rna.nl) 2021 Independent named-author account of failure modes when SAFe is imposed at scale in large organisations, useful cross-reference against consultancy case studies claiming clean SAFe rollouts.

We use analytics cookies to understand site usage and improve the service. We do not use marketing cookies.