Research · Academic & arXiv

Back to sweep

Research sweep · deep · 2022 – 2026

Engineering Maturity Models for Regulated Financial Services

Which engineering maturity models and capability frameworks large financial services technology organisations use to drive an engineering hygiene uplift (automated releases, regression test automation, CI/CD, observability) and how they roll out, monitor and govern the programme at scale, September 2022 to September 2026: DORA capabilities and metrics, CMMI and TMMi, ITIL 4 change enablement, SAFe and Team Topologies, the CNCF Platform Engineering Maturity Model, engineering scorecards (Backstage, Cortex, OpsLevel), and regulator expectations from the EU Digital Operational Resilience Act, PRA SS1/21 and the FCA operational resilience regime

  • Claude Fable 5
  • tech
  • financial
  • academic
  • blogs
  • vc
  • frontier

Synthesised 2026-09-21

Narrative

Academic and preprint coverage of engineering maturity in regulated financial services is thin and lags the practitioner literature by several years. The clearest finding across multiple 2024-2026 papers is a near-total absence of peer-reviewed, tier-1 venue research on the tools banks actually use. A 2026 Frontiers multivocal literature review of platform engineering and internal developer portals found only two peer-reviewed tier-1 papers on the topic globally and reported that "arXiv, which typically captures early-stage computer science research, returned zero directly relevant results," while noting that scorecards, the primary IDP governance mechanism, have "no peer-reviewed empirical evidence of effectiveness exists despite widespread commercial adoption". The same review found the only large-sample empirical evidence on platform engineering's effect on delivery outcomes comes from DORA's own 39,000+ respondent survey programme, not independent academic work, and flagged self-reported vendor and survey data (Spotify, Puppet) as Tier C/D evidence.

DORA's own metric set is well documented as evolving rather than fixed: the framework's history moved from four keys in 2014 to five by adding "reliability" and later replacing MTTR with a Failed Deployment Recovery Time measure, a debunking early in the research that change fail rate did not initially correlate with a single "IT performance" construct. Recent arXiv work extends or critiques the four/five-key model directly: a 2026 paper on "Measuring Delivery Consistency in Practice" argues DORA's dashboard cannot distinguish steady-cadence from unpredictable release patterns even when deployment frequency and change failure rate are identical, and proposes a deployment-consistency metric tested on 120 weeks of multi-platform release data. A separate ACM paper on automated DORA metrics computation notes that "research on automated DORA metrics is still scarce, apart from a few industrial proofs of concept," and validates an open-source measurement pipeline against a 37-microservice case study. DORA's most recent (2024-2025) research on generative AI, cited in several 2025-2026 arXiv papers, found that AI adoption raises individual effectiveness and throughput but is associated with a measurable reduction in delivery stability at team level, described as AI acting as "an amplifier, magnifying existing organisational strengths and weaknesses" rather than a uniform improvement.

On capability-level models, TMMi has the strongest financial-sector-specific empirical base: a peer-reviewed study surveyed sixty financial institutions (banking, insurance, pension funds) globally on test maturity and reported that the most common achieved level was TMMi Level 3 ("Defined"), with financial institutions reporting benefits concentrated in software quality and testing productivity and near-universal parallel use of test automation, especially at system level. CMMI evidence in financial services is largely case-study and ROI-framing work rather than large-sample studies; one exploratory case study quantifies CMMI's return-on-investment impact but at a single organisation, and broader CMMI-appraisal literature (SCAMPI-based) remains dominated by government-contracting contexts rather than banks. SAFe and Team Topologies have a moderate academic literature, mostly critical: a 2022 ICSE-SEIP paper on SAFe adoption issues, based on 25 respondents from seventeen companies in eight countries, found SAFe "subject to criticism: it appears to be quite demanding and expensive," and a 2023 arXiv empirical comparison of scaling approaches found SAFe holds roughly 53% market share among agile-scaling frameworks yet is contested on effectiveness grounds. A 2025 ScienceDirect paper analysing a SAFe case study frames the core tension as balancing organisational alignment, team autonomy, and control, which maps directly onto the central problem for a central standards capability that does not own delivery teams.

The evidence on AI-assisted coding tools in enterprise settings, directly relevant to the 2026 end of this lane's date range, is now a fast-growing arXiv sub-literature. Papers such as "The Fast and Spurious: Developer Productivity with GenAI" and "Debt Behind the AI Boom" report that Google and Microsoft disclosed AI now writes over 20% of new code in 2025, and that AI-generated code shows measurable maintainability and defect-injection issues even as individual throughput rises, corroborating rather than contradicting DORA's stability-versus-throughput finding. METR's HCAST and RE-Bench benchmarks, while not focused on financial-services delivery, are the most rigorous available empirical instruments for measuring what autonomous coding agents can actually do unsupervised, relevant background for anyone evaluating AI-assisted delivery risk; METR's own research finds the length of tasks models can complete autonomously has been "consistently exponentially increasing... with a doubling time of around 7 months."


Sources

ID Title Outlet Date Significance
a1 On the Need to Monitor Continuous Integration Practices -- An Empirical Study arXiv 2024-09 Empirical study on gaps in monitoring CI practices, relevant to hygiene uplift measurement
a2 Measuring Delivery Consistency in Practice: A DORA Extension from a Multi-Platform Release Setting arXiv 2026-05 Practitioner-empirical extension of DORA metrics showing the four/five-key model cannot distinguish steady from erratic delivery, tested on 120 weeks of release data
a3 A history of DORA's software delivery metrics DORA (Google Cloud) 2026-01 Primary source tracing how DORA's own metric set changed from four keys (2014) to five (including reliability and failed deployment recovery time)
a4 Auditable DevOps Automation via VSM and GQM arXiv 2026-01 Positions DORA and SPACE against DevOps capability maturity models (e.g. DevOps Institute CAMM), useful for distinguishing metric frameworks from level-based models
a5 Code Red: The Business Impact of Code Quality -- A Quantitative Study of 39 Proprietary Production Codebases arXiv 2022-03 Quantitative study of 39 proprietary codebases explicitly framed as a complement to DORA's four key metrics, covering code-level waste DORA does not measure
a6 SPACEX: Exploring metrics with the SPACE model for developer productivity arXiv 2025-11 Recent (2025) empirical exploration of SPACE framework dimensions using repository-mining metrics
a7 The SPACE of Developer Productivity: There's more to it than you think ACM Queue / Microsoft Research 2021-02 Foundational SPACE framework paper by Forsgren et al., ACM Queue 2021, the canonical outcome-metric alternative to DORA's delivery-only focus
a8 The Fast and Spurious: Developer Productivity with GenAI arXiv 2025-10 2025 paper contrasting DevEx, DORA and SPACE frameworks and testing them against generative-AI-assisted development data
a9 The Productivity-Reliability Paradox: Specification-Driven Governance for AI-Augmented Software Development arXiv 2026-05 Systematic review finding only 15% of AI-productivity studies examine more than three SPACE dimensions, evidencing thin measurement rigour
a10 Secure DevOps (DevSecOps) Maturity Models for Regulated Financial Institutions Zenodo 2026-02 Direct study of DevSecOps maturity models mapped to regulatory and audit readiness in UK financial institutions
a11 Assessing the maturity of software testing services using CMMI-SVC: An industrial case study arXiv 2020-05 Empirical CMMI-SVC appraisal applied to testing-service providers, relevant to test-process capability assessment methodology
a12 Software Testing Process Models Benefits & Drawbacks: a Systematic Literature Review arXiv 2019-01 Systematic review of 17 testing maturity/process models including TMM, useful for comparing test maturity model landscape
a13 Platform engineering and internal developer portals: a multivocal literature review Frontiers in Computer Science 2026-04 Key finding that peer-reviewed evidence for platform engineering and scorecards is almost absent; identifies DORA's 39,000-respondent survey as the only large-sample evidence base
a14 Knowledge Activation: AI Skills as the Institutional Knowledge Primitive for Agentic Software Development arXiv 2026-03 Discusses internal developer platforms and golden paths as the mechanism for enforcing engineering standards at scale, with projections that 80% of engineering orgs will run platform teams by 2026
a15 Balancing organizational alignment, team autonomy, and control in large-scale agile organizations Information and Software Technology (ScienceDirect) 2025-12 Peer-reviewed case study analysing SAFe's core tension between alignment, autonomy and control, directly relevant to a non-owning a central standards capability
a16 Do Agile Scaling Approaches Make A Difference? An Empirical Comparison of Team Effectiveness Across Popular Scaling Approaches arXiv 2023-10 Empirical comparison finding SAFe holds roughly 53% market share among agile-scaling approaches, with mixed effectiveness evidence
a17 Issues in the adoption of the scaled agile framework ACM/ICSE-SEIP 2022 ICSE-SEIP 2022 study of 25 respondents across 17 companies in 8 countries on SAFe adoption problems, a rare tier-1 peer-reviewed critique
a18 An empirical analysis of success factors in the adaption of the scaled agile framework arXiv 2020-12 Early empirical survey of factors affecting SAFe adoption success, contributing to the evidence base on scaling frameworks
a19 The Impact of Capability Maturity Model Integration on Return on Investment in IT Industry: An Exploratory Case Study Engineering, Technology & Applied Science Research 2017 Single-organisation CMMI ROI case study, illustrating the thin, non-generalisable empirical base for CMMI's financial impact claims
a20 The Future of Generative AI in Software Engineering: A Vision from Industry and Academia in the European GENIUS Project arXiv 2025-11 Summarises 2024-2025 DORA report findings that generative AI adoption improves throughput but increases software delivery instability
a21 Echoes of AI: Investigating the Downstream Effects of AI Assistants on Software Maintainability arXiv 2025-07 Empirical study of AI coding assistants' downstream effects on maintainability, relevant to whether AI adoption worsens rework in delivery pipelines
a22 HCAST: Human-Calibrated Autonomy Software Tasks arXiv / METR 2025-03 METR benchmark paper measuring frontier AI agents' autonomous software task capability, background for assessing AI-assisted delivery risk in regulated environments
a23 Time Horizon 1.1 METR 2026-01 METR's updated methodology and task-suite expansion for measuring autonomous AI task-completion capability, evidencing rapid capability growth relevant to AI-assisted engineering risk

We use analytics cookies to understand site usage and improve the service. We do not use marketing cookies.