Research · Academic & arXiv
Back to sweepResearch sweep · deep · 2023 – 2026
On-prem and open-weight models in regulated, high-security enterprises
Adoption of on-premises and open-weight AI models in financial services, defence and healthcare, September 2023–September 2026: the split between hosted API, private cloud and on-prem deployment, routing and gateway tooling (OpenRouter, LiteLLM, vLLM, NVIDIA NIM), sovereign model providers (Mistral, Cohere, Aleph Alpha) and their government agreements including Cohere and Mistral with the UK Government, and how regulated firms actually use models in high-security environments
- Claude Fable 5.1
- financial
- frontier
- academic
- vc
- blogs
- tech
Synthesised 2026-09-26
Narrative
The academic literature on regulated-sector AI deployment is dominated by frameworks, benchmarks and position papers rather than large-sample measured surveys of actual deployment splits. METR's own publications anchor the empirical end: HCAST (arXiv 2503.17354) and the time-horizon paper (arXiv 2503.14499) establish that frontier model autonomy on long tasks has roughly doubled every seven months since 2019, while RE-Bench (arXiv 2411.15114) finds AI agents beat 61 human ML experts on short time budgets but lose on 8- and 32-hour budgets. These are independently measured capability findings, not vendor claims, and they matter for regulated deployers because both the EU AI Act and the White House National Security Memorandum flag AI R&D automation as a capability of specific concern.
On deployment economics, arXiv papers such as the cost-benefit analysis of on-premise deployment (2509.18101) and the on-premises "middle path" position paper (2410.11182) work through the trade-off between data privacy, model confidentiality and running cost, but neither offers a large-N figure for what share of regulated enterprise traffic actually runs on-prem today. Financial-sector-specific work, including the global survey of generative AI in financial institutions (arXiv 2504.21574) and the DACH-focused DORA-compliant multi-agent architecture paper (arXiv 2609.27632), documents design patterns banks are adopting under DORA's third-party ICT risk regime, corroborated on the regulator side by a BIS survey of central bank cybersecurity experts (BIS working paper 145) showing most central banks have adopted or plan to adopt generative AI for cyber defence specifically.
Sovereignty appears in the literature as a strategic and regulatory question rather than a settled procurement metric. The "Sovereign Large Language Models" paper (arXiv 2503.04745) and the ACM-endorsed "Buy versus Build an LLM" government framework (arXiv 2602.13033) both treat UK-style agreements with Cohere and Mistral as instances of a buy-versus-build spectrum rather than proof of sovereignty delivering a measurable outcome. On tooling, RouterArena (arXiv 2510.00202) and LLMRouterBench (arXiv 2601.07206) are the first independent benchmarks of LLM gateways, finding that embedding-dependent routers such as RouteLLM carry materially higher latency than rule-based alternatives, a concrete figure largely absent from vendor documentation for OpenRouter, LiteLLM or NVIDIA NIM.
Use-case evidence in defence and healthcare is more concrete but narrower in scope. Papers on military LLM applications (arXiv 2511.10093, 2407.03453) document specific programmes such as US DoD wargaming trials and Task Force Lima's 180-plus identified use cases, while a psychiatric on-device deployment study (arXiv 2604.18302) and a government document-redaction paper using a local LLM (arXiv 2605.10211) are rare empirical tests of zero-egress and on-prem architecture in practice rather than descriptions of intended design. Compliance benchmarks such as COMPL-AI (arXiv 2410.07959) and AIReg-Bench (arXiv 2510.01474) let EU AI Act compliance claims be checked against test scores, a mechanism that could eventually discipline vendor sovereignty and compliance marketing if adopted by procurement teams.
Sources
| ID | Title | Outlet | Date | Significance |
|---|---|---|---|---|
| a1 | HCAST: Human-Calibrated Autonomy Software Tasks | arXiv (METR) | 2025-03 | Defines the 189-task, human-time-calibrated suite METR uses to score model autonomy, the methodological basis for time-horizon claims cited across regulated-deployment risk assessments. |
| a2 | Measuring AI Ability to Complete Long Software Tasks | arXiv (METR) | 2025-03 | Introduces the 50%-task-completion time-horizon metric and reports that frontier model autonomy on long tasks has roughly doubled every seven months since 2019, independently measured by METR rather than vendor-reported. |
| a3 | RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts | arXiv (METR) | 2024-11 | Compares AI agents against 61 human ML experts on 7 open-ended research-engineering environments, finding agents beat humans on a 2-hour budget but lose on 8- and 32-hour budgets, directly relevant to claims about autonomous capability in high-security AI R&D settings. |
| a4 | Position: On-Premises LLM Deployment Demands a Middle Path: Preserving Privacy Without Sacrificing Model Confidentiality | arXiv | 2024-10 | Foundational position paper on the technical tension in on-prem deployment between protecting a vendor's model weights and protecting the client's data, framing the design trade-off regulated firms face when they self-host rather than call a hosted API. |
| a5 | A Cost-Benefit Analysis of On-Premise Large Language Model Deployment | arXiv | 2025-09 | Provides an explicit cost-benefit framework for when local deployment of open-source models becomes economically viable against ongoing API spend, filling a gap where vendor claims about on-prem savings usually go unquantified. |
| a6 | Responsible Innovation: A Strategic Framework for Financial LLM Integration | arXiv | 2025-04 | Analyses how in-house LLM deployment gives financial institutions full lifecycle control needed for data residency and audit requirements, a structural argument for why banks lean on-prem or private cloud rather than hosted API alone. |
| a7 | Sovereign Large Language Models: Advantages, Strategy and Regulations | arXiv | 2025-03 | Systematic academic treatment of sovereign LLM strategy, setting out the regulatory and geopolitical rationale behind national champions such as Mistral, Cohere and Aleph Alpha rather than treating sovereignty as a marketing label. |
| a8 | Buy versus Build an LLM: A Decision Framework for Governments | arXiv | 2026-02 | ACM-endorsed decision framework for public-sector buy-versus-build choices, giving a structured lens on why the UK Government worked with Cohere and Mistral rather than building domestic models from scratch. |
| a9 | Sovereign AI Without Building Everything: A Readiness Model for Emerging Economies | SSRN | 2026 | SSRN working paper proposing a readiness model for partial sovereignty (compute, data, models) that lets smaller states or regulators assess sovereign AI claims against measurable capability tiers rather than binary sovereign-or-not framing. |
| a10 | AI Adoption and Central Banks in Emerging Markets: Challenges and Strategies | SSRN | 2025 | Working paper on how central banks in emerging markets are approaching generative AI governance and adoption, a comparator for how the Bank of England and ECB are treating sovereignty and deployment mode. |
| a11 | Generative AI in Financial Institution: A Global Survey of Opportunities, Challenges, and Future Directions | arXiv | 2025-04 | Global academic survey of generative AI use inside financial institutions, cataloguing deployment patterns and regulatory friction points across jurisdictions rather than relying on a single vendor's telemetry. |
| a12 | Generative artificial intelligence and cyber security in central banking | Bank for International Settlements | 2024 | Regulator-published BIS survey of central bank cybersecurity experts finding most central banks have adopted or plan to adopt generative AI for cyber defence, a rare regulator-corroborated adoption figure rather than a vendor claim. |
| a13 | RouterArena: An Open Platform for Comprehensive Comparison of LLM Routers | arXiv | 2025-10 | Independent benchmarking platform for LLM routers that measures accuracy, cost and latency trade-offs directly, giving an evidence base for gateway tooling claims that goes beyond router vendors' own marketing pages. |
| a14 | LLMRouterBench: A Massive Benchmark and Unified Framework for LLM Routing | arXiv | 2026-01 | Large-scale router benchmark using OpenRouter serving statistics to approximate end-to-end latency, finding embedding-dependent routers such as RouteLLM carry materially higher latency than rule-based alternatives, a concrete measured figure for gateway design in latency-sensitive regulated workflows. |
| a15 | COMPL-AI Framework: A Technical Interpretation and LLM Benchmarking Suite for the EU Artificial Intelligence Act | arXiv | 2024-10 | Translates the EU AI Act's abstract obligations into a concrete, testable benchmark suite for LLMs, letting compliance claims by model vendors be checked against measured scores rather than accepted at face value. |
| a16 | AIReg-Bench: Benchmarking Language Models That Assess AI Regulation Compliance | arXiv | 2025-10 | Tests whether LLMs themselves can reliably assess AI regulation compliance, relevant to whether regulated firms can automate DORA or AI Act documentation review rather than needing manual legal sign-off. |
| a17 | Compliant AI Infrastructure for Regulated Finance: A Tiered Multi-Agent Framework with DLT Audit Trails for Financial Operations in DACH | arXiv | 2026-09 | Proposes a tiered architecture with distributed-ledger audit trails specifically to meet DORA's third-party ICT risk and dual-control requirements, showing how German, Austrian and Swiss financial firms are engineering around EU rules rather than treating them as a checklist. |
| a18 | Benchmarking Multi-Agent LLM Architectures for Financial Document Processing: A Comparative Study of Orchestration Patterns, Cost-Accuracy Tradeoffs and Production Scaling Strategies | arXiv | 2026-03 | Measures cost-accuracy trade-offs across orchestration patterns for production financial document processing, offering measured hardware and cost figures rarely disclosed in vendor case studies. |
| a19 | On the Military Applications of Large Language Models | arXiv | 2025-11 | Surveys concrete defence use cases and deployment constraints for LLMs, including edge and tactical hardware limits, grounding claims about air-gapped or classified-enclave defence deployment in documented use cases rather than anecdote. |
| a20 | On Large Language Models in National Security Applications | arXiv | 2024-07 | Early systematic academic treatment of LLM use in national security, including US DoD wargaming trials and Task Force Lima's cataloguing of over 180 defence use cases, a foundational reference for defence-sector adoption trajectory. |
| a21 | ARMOR 2025: A Military-Aligned Benchmark for Evaluating Large Language Model Safety Beyond Civilian Contexts | arXiv | 2026-05 | Purpose-built safety benchmark for military-context LLM use, filling a gap where civilian safety evaluations do not test for compliance with rules of engagement or noncombatant-immunity constraints in tactical deployment. |
| a22 | Toward Zero-Egress Psychiatric AI: On-Device LLM Deployment for Privacy-Preserving Mental Health Decision Support | arXiv | 2026-04 | Demonstrates a fully on-device, zero-egress LLM architecture for a clinical mental-health use case, an empirical test of the air-gapped healthcare deployment pattern rather than a description of intended architecture. |
| a23 | To Redact, or not to Redact? A Local LLM Approach to Deliberative Process Privilege Classification | arXiv | 2026-05 | Applies a locally hosted open-weight model to a government document-redaction task, a measured example of on-prem deployment substituting for API calls specifically because of legal privilege and data-sensitivity constraints. |
| a24 | Experience Deploying Containerized GenAI Services at an HPC Center | arXiv | 2025-09 | Reports first-hand operational experience running NVIDIA NIM inference microservices at a high-performance computing centre, including documented failure modes in air-gapped and offline profile-download scenarios relevant to regulated on-prem stacks. |
| a25 | The Aloe Family Recipe for Open and Specialized Healthcare LLMs | arXiv | 2025-05 | Documents the training recipe and benchmark performance of an open-weight healthcare-specialised model family, giving a measured comparison point for open-weight clinical deployment against proprietary hosted alternatives. |