Research · Academic & arXiv
Back to sweepResearch sweep · deep · 2022 – 2026
Engineering Maturity Models for Regulated Financial Services
Which engineering maturity models and capability frameworks large financial services technology organisations use to drive an engineering hygiene uplift (automated releases, regression test automation, CI/CD, observability) and how they roll out, monitor and govern the programme at scale, September 2022 to September 2026: DORA capabilities and metrics, CMMI and TMMi, ITIL 4 change enablement, SAFe and Team Topologies, the CNCF Platform Engineering Maturity Model, engineering scorecards (Backstage, Cortex, OpsLevel), and regulator expectations from the EU Digital Operational Resilience Act, PRA SS1/21 and the FCA operational resilience regime
- Claude Fable 5
- tech
- financial
- academic
- blogs
- vc
- frontier
Synthesised 2026-09-21
Narrative
Academic and preprint coverage of engineering maturity in regulated financial services is thin and lags the practitioner literature by several years. The clearest finding across multiple 2024-2026 papers is a near-total absence of peer-reviewed, tier-1 venue research on the tools banks actually use. A 2026 Frontiers multivocal literature review of platform engineering and internal developer portals found only two peer-reviewed tier-1 papers on the topic globally and reported that "arXiv, which typically captures early-stage computer science research, returned zero directly relevant results," while noting that scorecards, the primary IDP governance mechanism, have "no peer-reviewed empirical evidence of effectiveness exists despite widespread commercial adoption". The same review found the only large-sample empirical evidence on platform engineering's effect on delivery outcomes comes from DORA's own 39,000+ respondent survey programme, not independent academic work, and flagged self-reported vendor and survey data (Spotify, Puppet) as Tier C/D evidence.
DORA's own metric set is well documented as evolving rather than fixed: the framework's history moved from four keys in 2014 to five by adding "reliability" and later replacing MTTR with a Failed Deployment Recovery Time measure, a debunking early in the research that change fail rate did not initially correlate with a single "IT performance" construct. Recent arXiv work extends or critiques the four/five-key model directly: a 2026 paper on "Measuring Delivery Consistency in Practice" argues DORA's dashboard cannot distinguish steady-cadence from unpredictable release patterns even when deployment frequency and change failure rate are identical, and proposes a deployment-consistency metric tested on 120 weeks of multi-platform release data. A separate ACM paper on automated DORA metrics computation notes that "research on automated DORA metrics is still scarce, apart from a few industrial proofs of concept," and validates an open-source measurement pipeline against a 37-microservice case study. DORA's most recent (2024-2025) research on generative AI, cited in several 2025-2026 arXiv papers, found that AI adoption raises individual effectiveness and throughput but is associated with a measurable reduction in delivery stability at team level, described as AI acting as "an amplifier, magnifying existing organisational strengths and weaknesses" rather than a uniform improvement.
On capability-level models, TMMi has the strongest financial-sector-specific empirical base: a peer-reviewed study surveyed sixty financial institutions (banking, insurance, pension funds) globally on test maturity and reported that the most common achieved level was TMMi Level 3 ("Defined"), with financial institutions reporting benefits concentrated in software quality and testing productivity and near-universal parallel use of test automation, especially at system level. CMMI evidence in financial services is largely case-study and ROI-framing work rather than large-sample studies; one exploratory case study quantifies CMMI's return-on-investment impact but at a single organisation, and broader CMMI-appraisal literature (SCAMPI-based) remains dominated by government-contracting contexts rather than banks. SAFe and Team Topologies have a moderate academic literature, mostly critical: a 2022 ICSE-SEIP paper on SAFe adoption issues, based on 25 respondents from seventeen companies in eight countries, found SAFe "subject to criticism: it appears to be quite demanding and expensive," and a 2023 arXiv empirical comparison of scaling approaches found SAFe holds roughly 53% market share among agile-scaling frameworks yet is contested on effectiveness grounds. A 2025 ScienceDirect paper analysing a SAFe case study frames the core tension as balancing organisational alignment, team autonomy, and control, which maps directly onto the central problem for a central standards capability that does not own delivery teams.
The evidence on AI-assisted coding tools in enterprise settings, directly relevant to the 2026 end of this lane's date range, is now a fast-growing arXiv sub-literature. Papers such as "The Fast and Spurious: Developer Productivity with GenAI" and "Debt Behind the AI Boom" report that Google and Microsoft disclosed AI now writes over 20% of new code in 2025, and that AI-generated code shows measurable maintainability and defect-injection issues even as individual throughput rises, corroborating rather than contradicting DORA's stability-versus-throughput finding. METR's HCAST and RE-Bench benchmarks, while not focused on financial-services delivery, are the most rigorous available empirical instruments for measuring what autonomous coding agents can actually do unsupervised, relevant background for anyone evaluating AI-assisted delivery risk; METR's own research finds the length of tasks models can complete autonomously has been "consistently exponentially increasing... with a doubling time of around 7 months."
Sources
| ID | Title | Outlet | Date | Significance |
|---|---|---|---|---|
| a1 | On the Need to Monitor Continuous Integration Practices -- An Empirical Study | arXiv | 2024-09 | Empirical study on gaps in monitoring CI practices, relevant to hygiene uplift measurement |
| a2 | Measuring Delivery Consistency in Practice: A DORA Extension from a Multi-Platform Release Setting | arXiv | 2026-05 | Practitioner-empirical extension of DORA metrics showing the four/five-key model cannot distinguish steady from erratic delivery, tested on 120 weeks of release data |
| a3 | A history of DORA's software delivery metrics | DORA (Google Cloud) | 2026-01 | Primary source tracing how DORA's own metric set changed from four keys (2014) to five (including reliability and failed deployment recovery time) |
| a4 | Auditable DevOps Automation via VSM and GQM | arXiv | 2026-01 | Positions DORA and SPACE against DevOps capability maturity models (e.g. DevOps Institute CAMM), useful for distinguishing metric frameworks from level-based models |
| a5 | Code Red: The Business Impact of Code Quality -- A Quantitative Study of 39 Proprietary Production Codebases | arXiv | 2022-03 | Quantitative study of 39 proprietary codebases explicitly framed as a complement to DORA's four key metrics, covering code-level waste DORA does not measure |
| a6 | SPACEX: Exploring metrics with the SPACE model for developer productivity | arXiv | 2025-11 | Recent (2025) empirical exploration of SPACE framework dimensions using repository-mining metrics |
| a7 | The SPACE of Developer Productivity: There's more to it than you think | ACM Queue / Microsoft Research | 2021-02 | Foundational SPACE framework paper by Forsgren et al., ACM Queue 2021, the canonical outcome-metric alternative to DORA's delivery-only focus |
| a8 | The Fast and Spurious: Developer Productivity with GenAI | arXiv | 2025-10 | 2025 paper contrasting DevEx, DORA and SPACE frameworks and testing them against generative-AI-assisted development data |
| a9 | The Productivity-Reliability Paradox: Specification-Driven Governance for AI-Augmented Software Development | arXiv | 2026-05 | Systematic review finding only 15% of AI-productivity studies examine more than three SPACE dimensions, evidencing thin measurement rigour |
| a10 | Secure DevOps (DevSecOps) Maturity Models for Regulated Financial Institutions | Zenodo | 2026-02 | Direct study of DevSecOps maturity models mapped to regulatory and audit readiness in UK financial institutions |
| a11 | Assessing the maturity of software testing services using CMMI-SVC: An industrial case study | arXiv | 2020-05 | Empirical CMMI-SVC appraisal applied to testing-service providers, relevant to test-process capability assessment methodology |
| a12 | Software Testing Process Models Benefits & Drawbacks: a Systematic Literature Review | arXiv | 2019-01 | Systematic review of 17 testing maturity/process models including TMM, useful for comparing test maturity model landscape |
| a13 | Platform engineering and internal developer portals: a multivocal literature review | Frontiers in Computer Science | 2026-04 | Key finding that peer-reviewed evidence for platform engineering and scorecards is almost absent; identifies DORA's 39,000-respondent survey as the only large-sample evidence base |
| a14 | Knowledge Activation: AI Skills as the Institutional Knowledge Primitive for Agentic Software Development | arXiv | 2026-03 | Discusses internal developer platforms and golden paths as the mechanism for enforcing engineering standards at scale, with projections that 80% of engineering orgs will run platform teams by 2026 |
| a15 | Balancing organizational alignment, team autonomy, and control in large-scale agile organizations | Information and Software Technology (ScienceDirect) | 2025-12 | Peer-reviewed case study analysing SAFe's core tension between alignment, autonomy and control, directly relevant to a non-owning a central standards capability |
| a16 | Do Agile Scaling Approaches Make A Difference? An Empirical Comparison of Team Effectiveness Across Popular Scaling Approaches | arXiv | 2023-10 | Empirical comparison finding SAFe holds roughly 53% market share among agile-scaling approaches, with mixed effectiveness evidence |
| a17 | Issues in the adoption of the scaled agile framework | ACM/ICSE-SEIP | 2022 | ICSE-SEIP 2022 study of 25 respondents across 17 companies in 8 countries on SAFe adoption problems, a rare tier-1 peer-reviewed critique |
| a18 | An empirical analysis of success factors in the adaption of the scaled agile framework | arXiv | 2020-12 | Early empirical survey of factors affecting SAFe adoption success, contributing to the evidence base on scaling frameworks |
| a19 | The Impact of Capability Maturity Model Integration on Return on Investment in IT Industry: An Exploratory Case Study | Engineering, Technology & Applied Science Research | 2017 | Single-organisation CMMI ROI case study, illustrating the thin, non-generalisable empirical base for CMMI's financial impact claims |
| a20 | The Future of Generative AI in Software Engineering: A Vision from Industry and Academia in the European GENIUS Project | arXiv | 2025-11 | Summarises 2024-2025 DORA report findings that generative AI adoption improves throughput but increases software delivery instability |
| a21 | Echoes of AI: Investigating the Downstream Effects of AI Assistants on Software Maintainability | arXiv | 2025-07 | Empirical study of AI coding assistants' downstream effects on maintainability, relevant to whether AI adoption worsens rework in delivery pipelines |
| a22 | HCAST: Human-Calibrated Autonomy Software Tasks | arXiv / METR | 2025-03 | METR benchmark paper measuring frontier AI agents' autonomous software task capability, background for assessing AI-assisted delivery risk in regulated environments |
| a23 | Time Horizon 1.1 | METR | 2026-01 | METR's updated methodology and task-suite expansion for measuring autonomous AI task-completion capability, evidencing rapid capability growth relevant to AI-assisted engineering risk |