Research · Tech Industry & Practitioner
Back to sweepResearch sweep · deep · 2023 – 2026
On-prem and open-weight models in regulated, high-security enterprises
Adoption of on-premises and open-weight AI models in financial services, defence and healthcare, September 2023–September 2026: the split between hosted API, private cloud and on-prem deployment, routing and gateway tooling (OpenRouter, LiteLLM, vLLM, NVIDIA NIM), sovereign model providers (Mistral, Cohere, Aleph Alpha) and their government agreements including Cohere and Mistral with the UK Government, and how regulated firms actually use models in high-security environments
- Claude Fable 5.1
- financial
- frontier
- academic
- vc
- blogs
- tech
Synthesised 2026-09-26
Narrative
The clearest primary-source evidence in this lane comes from regulators rather than vendors. The Bank of England and FCA's third AI and Machine Learning Survey, published in November 2024, found 75 percent of UK financial firms using AI, up from 58 percent in 2022, with foundation models now accounting for 17 percent of use cases and fully autonomous decision-making still confined to 2 percent of deployments. The Bank's April 2025 Financial Stability in Focus note follows up with a systemic-risk argument: concentration among a handful of third-party model and cloud providers is now something supervisors track explicitly, which is the regulatory logic underpinning sovereignty and on-prem procurement rather than a marketing frame.
On tooling, ThoughtWorks' Technology Radar documents LiteLLM's move from a thin multi-provider wrapper into a governance-capable gateway, handling retries, budget controls and edge-level guardrails, and separately tracks vLLM as a standard self-hosted inference worker via FastChat. CNCF's own 2025 annual survey, published January 2026, reports 82 percent of container users running production AI workloads on Kubernetes, with GPU-aware Dynamic Resource Allocation reaching general availability and a Gateway API Inference Extension purpose-built for routing inference traffic by model name and endpoint health. These are the concrete platform components sitting under any regulated firm's on-prem or private-cloud model serving, and they are foundation-published survey data, not vendor case studies.
Sovereign-model government agreements are documented unevenly. Cohere's own blog post claims a memorandum of understanding with the UK government on secure AI collaboration, corroborated independently by BetaKit's reporting but without disclosed monetary terms. Mistral's clearest measurable outcome is not with the UK but with France: a defence framework agreement awarded by the Ministry of the Armed Forces on 8 January 2026, per trade coverage. Aleph Alpha's trajectory undercuts a simple national sovereignty narrative: after abandoning frontier-model development in 2024 to focus on its PhariaAI platform for German public agencies, it is being acquired by Cohere with a $600 million Schwarz Group investment, reported by Bloomberg in April 2026, meaning Germany's sovereign-AI bet is consolidating into a North American vendor's stack.
High-security deployment evidence is strongest at Los Alamos National Laboratory, where the Department of Energy and NNSA confirm the Venado supercomputer, running on NVIDIA GH200 Grace Hopper Superchips, moved onto a classified network in 2025 to run OpenAI's o1 and o3 models against controlled unclassified and ITAR-restricted data, independently corroborated by The Register. IBM's Granite-based defence model, covered by DefenseScoop in October 2025, targets the same air-gapped and classified segment as a packaged product. Set against McKinsey's March 2025 global survey of 1,993 organisations, where only 5.5 percent report material financial returns from AI despite 88 percent adoption, the regulated-sector figures in this lane look consistent with a broader pattern: adoption metrics run well ahead of measured value or independently audited deployment detail.
Sources
| ID | Title | Outlet | Date | Significance |
|---|---|---|---|---|
| p1 | LiteLLM | ThoughtWorks Technology Radar | 2025 | ThoughtWorks describes LiteLLM's evolution from a thin multi-provider abstraction into a governance-capable AI gateway (retries, budget controls, access control, edge guardrails), and names it a default choice for teams mixing hosted and self-hosted models. |
| p2 | FastChat | ThoughtWorks Technology Radar | 2025 | Documents practitioner use of vLLM as a self-hosted inference worker behind a routing layer, evidence that on-prem model serving is being treated as a standard architectural component, not a bespoke build. |
| p3 | Technology Radar Volume 32 | ThoughtWorks Technology Radar | 2025-04 | Primary-source radar document tracking which LLM serving, gateway and evaluation tools ThoughtWorks' delivery teams were assessing or adopting across client engagements in early 2025, ahead of most vendor marketing on the same tools. |
| p4 | DORA | State of AI-assisted Software Development 2025 | DORA (Google Cloud) | 2025 | Survey-based (nearly 5,000 technology professionals) finding that AI adoption in software delivery reached 90 percent but acts as an amplifier of existing platform and process quality rather than a uniform productivity gain, directly relevant to how regulated engineering organisations should read AI tooling claims. |
| p5 | Emerging Patterns in Building GenAI Products | martinfowler.com | 2024 | Practitioner-authored catalogue of RAG, gateway and evaluation patterns used to build LLM products, forming the architectural vocabulary (retrieval, guardrails, routing) that regulated deployments in finance and healthcare are built from. |
| p6 | Scaling AI With Adaptive Governance | MIT Sloan Management Review | 2025 | Based on interviews with AI governance leaders at Barclays, Lloyds Bank, Danske Bank, Nasdaq and the Abu Dhabi Department of Finance, this is a rare named-institution account of how banks actually structure controls over model deployment, not a vendor case study. |
| p7 | Match Your AI Strategy to Your Organization's Reality | Harvard Business Review | 2026-01 | Sets out a framework (focused differentiation, vertical integration, collaborative ecosystem, platform leadership) for choosing deployment posture, useful for distinguishing when sovereignty or on-prem control is a genuine strategic fit versus a compliance reflex. |
| p8 | Kubernetes Established as the De Facto 'Operating System' for AI as Production Use Hits 82% in 2025 CNCF Annual Cloud Native Survey | CNCF | 2026-01 | CNCF's own annual survey (analyst/foundation-published, self-reported by member organisations) puts Kubernetes-based production AI workloads at 82 percent and documents GA of Dynamic Resource Allocation for GPU scheduling and the Gateway API Inference Extension for routing inference traffic, the concrete platform layer under regulated on-prem model serving. |
| p9 | The platform under the model: How cloud native powers AI engineering in production | CNCF | 2026-03 | Describes how CNCF TAG-adjacent tooling (GPU scheduling, inference routing, observability) is being assembled into production AI platforms, the infrastructure layer regulated firms use to run open-weight models on their own hardware. |
| p10 | Developers remain willing but reluctant to use AI: The 2025 Developer Survey results are here | Stack Overflow | 2025-12 | Stack Overflow's own survey (self-reported, large developer sample) records AI tool usage at 84 percent against trust in output accuracy of only 29 percent, a data point that should discount vendor claims of frictionless adoption inside regulated engineering teams. |
| p11 | Open-Source Models Maintained 30% Market Share In 2025, But Chinese Models Grew At The Cost Of Other Open Models: OpenRouter Data | OfficeChai | 2025-12 | Secondary breakdown of OpenRouter's own data showing the shift within the open-weight share from Llama and Mistral toward DeepSeek and Qwen, relevant to which open-weight families dominate self-hosted deployments and licence terms decision-makers must check individually. |
| p12 | Mistral AI wins French defence AI framework agreement | Gend | 2026-01 | Reports a named, dated procurement outcome, France's Ministry of the Armed Forces awarding Mistral a framework agreement on 8 January 2026, one of the few sovereignty cases in this space with a measurable government contract rather than a marketing claim. |
| p13 | Cohere to Acquire Aleph Alpha With Backing From Schwarz Group | Bloomberg | 2026-04 | Independent financial reporting on Cohere's acquisition of Aleph Alpha with a $600 million Schwarz Group investment, evidence that Germany's sovereign-model bet consolidated into a North American vendor's stack rather than sustaining a purely domestic frontier-model alternative. |
| p14 | Air-Gap Deployment - NVIDIA NIM for Large Language Models | NVIDIA (technical documentation) | 2025 | Primary vendor documentation for deploying NIM inference containers with no external network access, the concrete mechanism (offline image pulls, local model caches) behind claims of air-gapped model serving in defence and classified environments. |
| p15 | NNSA's Los Alamos National Laboratory launches frontier AI models on the Venado supercomputer | U.S. Department of Energy / NNSA | 2025 | Government-published account of a national laboratory running frontier models on a GH200-based supercomputer moved onto a classified network for handling CUI, UCNI and ITAR-controlled data, a documented, named on-prem deployment in a high-security setting rather than an anecdote. |
| p16 | OpenAI inks deal with Los Alamos lab to cram o1 into Venado | The Register | 2025-01 | Independent trade-press account of the same LANL deployment that adds sceptical framing on classification timelines and hardware (GH200 Grace Hopper Superchips), useful for cross-checking the government press release. |
| p17 | A first look at IBM's new large language model that's fine-tuned for defense applications | DefenseScoop | 2025-10 | Independent defence-trade reporting on IBM's Granite-based defence model, built for air-gapped, classified and edge deployment and trained partly on Janes data, a named-vendor product aimed squarely at the on-prem defence segment this lane tracks. |
| p18 | The state of AI: How organizations are rewiring to capture value | McKinsey & Company | 2025-03 | Analyst-estimated global survey (1,993 respondents, 105 countries, fielded June to July 2025) finding 88 percent of organisations use AI in at least one function but only 5.5 percent report material financial returns, a cross-industry benchmark against which the regulated-sector figures in this lane should be read. |
| p19 | Transforming Financial Analysis with NVIDIA NIM | NVIDIA Technical Blog | 2024 | Vendor-published technical account of NIM microservices applied to financial-analysis workloads, useful for the inference-stack detail (containerised model serving, Triton backend) but a vendor claim rather than an independently verified bank deployment. |