Research · Tech Industry & Practitioner

Back to sweep

Research sweep · deep · 2022 – 2026

Engineering Maturity Models for Regulated Financial Services

Which engineering maturity models and capability frameworks large financial services technology organisations use to drive an engineering hygiene uplift (automated releases, regression test automation, CI/CD, observability) and how they roll out, monitor and govern the programme at scale, September 2022 to September 2026: DORA capabilities and metrics, CMMI and TMMi, ITIL 4 change enablement, SAFe and Team Topologies, the CNCF Platform Engineering Maturity Model, engineering scorecards (Backstage, Cortex, OpsLevel), and regulator expectations from the EU Digital Operational Resilience Act, PRA SS1/21 and the FCA operational resilience regime

  • Claude Fable 5
  • tech
  • financial
  • academic
  • blogs
  • vc
  • frontier

Synthesised 2026-09-21

Narrative

DORA's Accelerate State of DevOps research remains the dominant outcome-metrics framework cited by financial services technology functions, and its own decade of data shows the field is not settling into a stable consensus. The 2024 report, based on responses from more than 39,000 professionals, found the "elite" cluster stable but the high-performance cluster shrinking from 31% of respondents in 2023 to 22% in 2024, while the low-performer cluster grew from 17% to 25%, and for the first time the medium-performance cluster showed a lower change failure rate than the high-performance cluster, undercutting the assumption that DORA's four metrics move in lockstep. DORA's 2025 State of AI-assisted Software Development report replaced the four-tier elite/high/medium/low ranking with seven team archetypes and introduced a companion "AI Capabilities Model" naming seven foundational capabilities (a clear AI stance, healthy data ecosystems, AI-accessible internal data, strong version control, small batch working, user-centric focus, and quality internal platforms) that determine whether AI adoption helps or harms delivery. Two consecutive DORA reports found AI adoption correlated with reduced software delivery performance even as it lifted individual productivity, and Google Cloud's own account of the 2024 report quantified this as roughly a 1.5% throughput decrease and 7.2% stability reduction per 25% increase in AI adoption.

On governance, DORA's published capability page on "streamlining change approval" is one of the more consequential and most-cited findings for regulated firms: it reports no evidence that formal external review (CAB approval) reduces change failure rates, while such heavyweight processes correlate with larger, less frequent, higher-impact releases. The original Accelerate book (Forsgren, Humble, Kim) went further, finding organisations with formal external change approval were 2.6 times more likely to be low performers, and that external approval had no correlation with change fail rate but was negatively correlated with lead time, deployment frequency and restore time. This finding is now being explicitly reworked into banking governance literature: a Team Topologies "Expert View" piece on reconciling DORA metrics with bank governance argues for replacing CAB approval with automated, peer-review-driven controls and enablement coaching rather than abandoning governance altogether. InfoQ's coverage of the 2018 DevOps Enterprise Summit London documented Capital One, Barclays, Lloyds Banking Group, Standard Bank, ABN Amro, UBS and RBS presenting DevOps transformation experience reports, which remains one of the few multi-bank practitioner accounts of this transition, though it predates the current brief's window and should be read as a durable baseline rather than current evidence.

Test and process maturity models occupy a narrower niche than DORA in financial services. A peer-reviewed international survey of sixty financial institutions (banks, insurers, pension funds) benchmarked against TMMi found consistent motivations for test process improvement and documented benefits, giving TMMi one of the few empirically surveyed (rather than single-case) footholds in the sector, though the TMMi Foundation's own case-study library remains largely vendor/consultancy testimonial in character. On operating models, an academic case study of a large full-service bank's SAFe transformation two years in identified persistent friction around management and organisation, requirements engineering, quality assurance and systems architecture, echoing generic SAFe criticisms rather than resolving them for a regulated context, and independent practitioner critiques continue to argue SAFe creates conflicts with SRE/operations functions and encourages large-batch, top-down planning that cuts against a metrics-first DORA approach. Team Topologies' own case study library includes Alfa Financial Software (asset finance SaaS serving Santander and Mercedes-Benz among clients) and REA Group's financial services division, both citing reduced cognitive load and faster flow after restructuring into stream-aligned teams, though these are vendor-published case studies rather than independently audited outcomes.

Platform engineering and scorecarding tooling shows a maturing but contested evidence base. The CNCF TAG App Delivery Platform Engineering Maturity Model (built by the CNCF Platforms Working Group) frames maturity across investment, adoption, interfaces, operations and measurement, explicitly warning that each additional maturity level demands more funding and staff time and that reaching the top level should not be a goal in itself; the CNCF whitepaper explicitly invokes Martin Fowler's framing that a maturity assessment's value lies in the list of things to improve, not the score itself. The State of Platform Engineering Report Volume 4 (518 practitioners) found 45.5% of organisations run dedicated but reactive platform teams, only 13.1% reach an "optimised" measurement stage, and 29.6% still measure no success indicator at all, alongside evidence that well-designed golden paths drive voluntary adoption above 80% versus roughly 20% for platforms without them. Backstage-adjacent commentary is more sobering for enterprise rollouts specifically: while Spotify itself reports adoption above 90%, multiple independent sources (Cortex, Humanitec, Earthly) report that other organisations' Backstage rollouts frequently stall below 10% adoption, and a Microsoft Learn case study of an unnamed bank's platform engineering programme (120-person central engineering team, "thousands of custom tools") explicitly frames golden paths as a negotiated compromise against one-size-fits-all mandates rather than a solved problem. Martin Fowler's own bliki entry on maturity models remains the clearest published skeptical-but-constructive framing available: models are "wrong but hopefully useful," easily corrupted by certification incentives, and valuable mainly for sequencing what to learn next rather than for the score itself.


Sources

ID Title Outlet Date Significance
p1 Announcing the 2024 DORA report Google Cloud Blog 2024-10 Google Cloud's own summary of the tenth Accelerate State of DevOps Report, the primary outcome-metrics framework cited across the industry; documents the AI-driven throughput/stability trade-off.
p2 DORA Accelerate State of DevOps 2024 Report Google Research 2024 Primary research publication page confirming survey scale (39,000+ professionals) and scope of the 2024 findings on platform engineering and AI.
p3 Highlights from the 2024 DORA State of DevOps Report DX (getdx.com) 2024-10 DX's practitioner analysis (Laura Tacho, DX CTO) of shifting performance clusters and the platform-engineering paradox in the 2024 report.
p4 2024 DORA Report DX Newsletter 2024-10 Detailed breakdown showing the high-performance cluster shrinking from 31% to 22% and low cluster growing from 17% to 25% between 2023 and 2024.
p5 DORA | Capabilities: Streamlining change approval DORA (dora.dev) 2023 DORA's canonical published finding that formal CAB/external review has no correlation with lower change failure rates and slows delivery, directly relevant to CAB-vs-continuous-delivery reconciliation.
p6 When DORA metrics meet governance in banking - Expert View Team Topologies 2026-01 Practitioner account applying DORA's CAB findings specifically to bank governance, proposing automated peer-review-based change control as the reconciliation mechanism.
p7 DORA's research program DORA (dora.dev) 2019-2024 DORA's own year-by-year timeline showing the 2019 finding that heavyweight change approval negatively affects performance and that communities of practice outperform centres of excellence.
p8 The Definitive Introduction to the DORA Research Information Safety 2022-07 Documents Nicole Forsgren's early Capital One case study showing a 20x increase in release frequency after adopting trunk-based development and streamlined change approval.
p9 InfoQ Homepage Continuous Delivery Content on InfoQ InfoQ 2018-2019 Indexes InfoQ's 'State of DevOps in Banking' report from DOES London 2018, summarising DevOps transformation talks from Capital One, Barclays, Lloyds, Standard Bank, ABN Amro, UBS and RBS.
p10 Case Studies - Team Topologies Team Topologies 2024-2025 Team Topologies' own case study library including Alfa Financial Software and REA Group's financial services division, illustrating vendor-published claims of reduced cognitive load and faster flow.
p11 Case Studies: Customer Platform Engineering Implementations Microsoft Learn 2025-10 Microsoft's own account of an unnamed bank's platform engineering rollout (120-person central team, thousands of custom tools) negotiating golden paths against one-size-fits-all mandates.
p12 Platform Engineering Maturity Model CNCF TAG App Delivery 2024 CNCF TAG App Delivery's own maturity model whitepaper, explicitly warning that higher maturity levels require proportionally more funding and citing Martin Fowler on the purpose of maturity assessment.
p13 Platform engineering maturity in 2026: What the data tells us platformengineering.org 2026-02 State of Platform Engineering Report Volume 4 (518 practitioners) quantifying CNCF maturity dimensions: 45.5% reactive dedicated teams, only 13.1% at optimised measurement stage, 29.6% measuring no success indicator.
p14 70% of Platform Initiatives Fail Without a Golden Path: Inside Vol. 4 of the State of Platform Engineering Report bex.co 2026-08 Reports the above-80% voluntary golden-path adoption figure against sub-20% for platforms without a designed paved road, and cites DORA's 2024 finding on platform-driven productivity gains.
p15 Spotify Backstage: Features, Benefits & Challenges in 2025 Cortex 2025-09 Independent vendor comparison documenting that Backstage adoption stalls below 10% in most organisations outside Spotify, despite Spotify's own near-universal internal adoption.
p16 Backstage by Spotify - The Ultimate Guide [2026] Roadie.io 2026 Documents named enterprise Backstage deployments (American Airlines, Expedia Group managing 20,000 microservices) as concrete rollout-scale evidence.
p17 Test Maturity Model integration (TMMi): Test Maturity in the Financial Domain American Journal of Computer Science and Technology 2024-05 Peer-reviewed international survey of sixty financial institutions (banks, insurers, pension funds) benchmarked against TMMi, one of the few multi-organisation empirical studies of test maturity in finance.
p18 Papers & Case Studies - TMMi TMMi Foundation 2023-2024 TMMi Foundation's own case study and survey index, including the 2nd TMMi World-Wide User Survey on costs and benefits, useful for distinguishing certification-body claims from independent evidence.
p19 Challenges of Adopting SAFe in the Banking Industry – A Study Two Years after its Introduction arXiv / Springer (RISE Research Institutes of Sweden) 2021 Academic case study of a large full-service bank's SAFe transformation identifying persistent challenges in management, requirements engineering, quality assurance and architecture, filling a gap in independent banking-sector SAFe evidence.
p20 bliki: Maturity Model martinfowler.com 2014 Martin Fowler's canonical skeptical-but-constructive framing of maturity models as simplifications whose value lies in sequencing improvement rather than in the score, widely cited by CNCF and others.
p21 How DX Core 4 aims to unify developer productivity frameworks LeadDev 2024-12 Explains the lineage and stated purpose of DORA, SPACE and DevEx, and DX's attempt to unify them into Core 4, clarifying which framework organisations use for which purpose.
p22 Comparing popular developer productivity frameworks: DORA, SPACE, and DX Core 4 Swarmia 2025-05 Independent critical comparison noting DX Core 4 lacks a stated theory of improvement and flags potentially harmful metrics like PRs-per-developer if used for individual evaluation.
p23 2025 DORA AI Capabilities Model DORA / Google Cloud 2025-10 Primary DORA document defining the seven AI capabilities (data ecosystems, version control, small batches, user-centric focus, quality internal platforms) that determine whether AI adoption helps or harms delivery performance.
p24 Fully Automated DORA Metrics Measurement for Continuous Improvement ACM (ICSSP '24) 2024-09 Peer-reviewed ACM conference paper (ICSSP '24) presenting an industrial case study of automated, microservice-level DORA metric collection across 37 services, evidence against manual survey-only measurement.
p25 4 things you need to know from the latest Thoughtworks Tech Radar LeadDev 2024-04 LeadDev's coverage of Thoughtworks' April 2024 Technology Radar debate on pull requests versus true trunk-based continuous integration, a live methodological fault line relevant to CI/CD maturity sequencing.

We use analytics cookies to understand site usage and improve the service. We do not use marketing cookies.