Research · Tech Industry & Practitioner

Back to sweep

Research sweep · deep · 2023 – 2026

On-prem and open-weight models in regulated, high-security enterprises

Adoption of on-premises and open-weight AI models in financial services, defence and healthcare, September 2023–September 2026: the split between hosted API, private cloud and on-prem deployment, routing and gateway tooling (OpenRouter, LiteLLM, vLLM, NVIDIA NIM), sovereign model providers (Mistral, Cohere, Aleph Alpha) and their government agreements including Cohere and Mistral with the UK Government, and how regulated firms actually use models in high-security environments

  • Claude Fable 5.1
  • financial
  • frontier
  • academic
  • vc
  • blogs
  • tech

Synthesised 2026-09-26

Narrative

The clearest primary-source evidence in this lane comes from regulators rather than vendors. The Bank of England and FCA's third AI and Machine Learning Survey, published in November 2024, found 75 percent of UK financial firms using AI, up from 58 percent in 2022, with foundation models now accounting for 17 percent of use cases and fully autonomous decision-making still confined to 2 percent of deployments. The Bank's April 2025 Financial Stability in Focus note follows up with a systemic-risk argument: concentration among a handful of third-party model and cloud providers is now something supervisors track explicitly, which is the regulatory logic underpinning sovereignty and on-prem procurement rather than a marketing frame.

On tooling, ThoughtWorks' Technology Radar documents LiteLLM's move from a thin multi-provider wrapper into a governance-capable gateway, handling retries, budget controls and edge-level guardrails, and separately tracks vLLM as a standard self-hosted inference worker via FastChat. CNCF's own 2025 annual survey, published January 2026, reports 82 percent of container users running production AI workloads on Kubernetes, with GPU-aware Dynamic Resource Allocation reaching general availability and a Gateway API Inference Extension purpose-built for routing inference traffic by model name and endpoint health. These are the concrete platform components sitting under any regulated firm's on-prem or private-cloud model serving, and they are foundation-published survey data, not vendor case studies.

Sovereign-model government agreements are documented unevenly. Cohere's own blog post claims a memorandum of understanding with the UK government on secure AI collaboration, corroborated independently by BetaKit's reporting but without disclosed monetary terms. Mistral's clearest measurable outcome is not with the UK but with France: a defence framework agreement awarded by the Ministry of the Armed Forces on 8 January 2026, per trade coverage. Aleph Alpha's trajectory undercuts a simple national sovereignty narrative: after abandoning frontier-model development in 2024 to focus on its PhariaAI platform for German public agencies, it is being acquired by Cohere with a $600 million Schwarz Group investment, reported by Bloomberg in April 2026, meaning Germany's sovereign-AI bet is consolidating into a North American vendor's stack.

High-security deployment evidence is strongest at Los Alamos National Laboratory, where the Department of Energy and NNSA confirm the Venado supercomputer, running on NVIDIA GH200 Grace Hopper Superchips, moved onto a classified network in 2025 to run OpenAI's o1 and o3 models against controlled unclassified and ITAR-restricted data, independently corroborated by The Register. IBM's Granite-based defence model, covered by DefenseScoop in October 2025, targets the same air-gapped and classified segment as a packaged product. Set against McKinsey's March 2025 global survey of 1,993 organisations, where only 5.5 percent report material financial returns from AI despite 88 percent adoption, the regulated-sector figures in this lane look consistent with a broader pattern: adoption metrics run well ahead of measured value or independently audited deployment detail.


Sources

ID Title Outlet Date Significance
p1 LiteLLM ThoughtWorks Technology Radar 2025 ThoughtWorks describes LiteLLM's evolution from a thin multi-provider abstraction into a governance-capable AI gateway (retries, budget controls, access control, edge guardrails), and names it a default choice for teams mixing hosted and self-hosted models.
p2 FastChat ThoughtWorks Technology Radar 2025 Documents practitioner use of vLLM as a self-hosted inference worker behind a routing layer, evidence that on-prem model serving is being treated as a standard architectural component, not a bespoke build.
p3 Technology Radar Volume 32 ThoughtWorks Technology Radar 2025-04 Primary-source radar document tracking which LLM serving, gateway and evaluation tools ThoughtWorks' delivery teams were assessing or adopting across client engagements in early 2025, ahead of most vendor marketing on the same tools.
p4 DORA | State of AI-assisted Software Development 2025 DORA (Google Cloud) 2025 Survey-based (nearly 5,000 technology professionals) finding that AI adoption in software delivery reached 90 percent but acts as an amplifier of existing platform and process quality rather than a uniform productivity gain, directly relevant to how regulated engineering organisations should read AI tooling claims.
p5 Emerging Patterns in Building GenAI Products martinfowler.com 2024 Practitioner-authored catalogue of RAG, gateway and evaluation patterns used to build LLM products, forming the architectural vocabulary (retrieval, guardrails, routing) that regulated deployments in finance and healthcare are built from.
p6 Scaling AI With Adaptive Governance MIT Sloan Management Review 2025 Based on interviews with AI governance leaders at Barclays, Lloyds Bank, Danske Bank, Nasdaq and the Abu Dhabi Department of Finance, this is a rare named-institution account of how banks actually structure controls over model deployment, not a vendor case study.
p7 Match Your AI Strategy to Your Organization's Reality Harvard Business Review 2026-01 Sets out a framework (focused differentiation, vertical integration, collaborative ecosystem, platform leadership) for choosing deployment posture, useful for distinguishing when sovereignty or on-prem control is a genuine strategic fit versus a compliance reflex.
p8 Kubernetes Established as the De Facto 'Operating System' for AI as Production Use Hits 82% in 2025 CNCF Annual Cloud Native Survey CNCF 2026-01 CNCF's own annual survey (analyst/foundation-published, self-reported by member organisations) puts Kubernetes-based production AI workloads at 82 percent and documents GA of Dynamic Resource Allocation for GPU scheduling and the Gateway API Inference Extension for routing inference traffic, the concrete platform layer under regulated on-prem model serving.
p9 The platform under the model: How cloud native powers AI engineering in production CNCF 2026-03 Describes how CNCF TAG-adjacent tooling (GPU scheduling, inference routing, observability) is being assembled into production AI platforms, the infrastructure layer regulated firms use to run open-weight models on their own hardware.
p10 Developers remain willing but reluctant to use AI: The 2025 Developer Survey results are here Stack Overflow 2025-12 Stack Overflow's own survey (self-reported, large developer sample) records AI tool usage at 84 percent against trust in output accuracy of only 29 percent, a data point that should discount vendor claims of frictionless adoption inside regulated engineering teams.
p11 Open-Source Models Maintained 30% Market Share In 2025, But Chinese Models Grew At The Cost Of Other Open Models: OpenRouter Data OfficeChai 2025-12 Secondary breakdown of OpenRouter's own data showing the shift within the open-weight share from Llama and Mistral toward DeepSeek and Qwen, relevant to which open-weight families dominate self-hosted deployments and licence terms decision-makers must check individually.
p12 Mistral AI wins French defence AI framework agreement Gend 2026-01 Reports a named, dated procurement outcome, France's Ministry of the Armed Forces awarding Mistral a framework agreement on 8 January 2026, one of the few sovereignty cases in this space with a measurable government contract rather than a marketing claim.
p13 Cohere to Acquire Aleph Alpha With Backing From Schwarz Group Bloomberg 2026-04 Independent financial reporting on Cohere's acquisition of Aleph Alpha with a $600 million Schwarz Group investment, evidence that Germany's sovereign-model bet consolidated into a North American vendor's stack rather than sustaining a purely domestic frontier-model alternative.
p14 Air-Gap Deployment - NVIDIA NIM for Large Language Models NVIDIA (technical documentation) 2025 Primary vendor documentation for deploying NIM inference containers with no external network access, the concrete mechanism (offline image pulls, local model caches) behind claims of air-gapped model serving in defence and classified environments.
p15 NNSA's Los Alamos National Laboratory launches frontier AI models on the Venado supercomputer U.S. Department of Energy / NNSA 2025 Government-published account of a national laboratory running frontier models on a GH200-based supercomputer moved onto a classified network for handling CUI, UCNI and ITAR-controlled data, a documented, named on-prem deployment in a high-security setting rather than an anecdote.
p16 OpenAI inks deal with Los Alamos lab to cram o1 into Venado The Register 2025-01 Independent trade-press account of the same LANL deployment that adds sceptical framing on classification timelines and hardware (GH200 Grace Hopper Superchips), useful for cross-checking the government press release.
p17 A first look at IBM's new large language model that's fine-tuned for defense applications DefenseScoop 2025-10 Independent defence-trade reporting on IBM's Granite-based defence model, built for air-gapped, classified and edge deployment and trained partly on Janes data, a named-vendor product aimed squarely at the on-prem defence segment this lane tracks.
p18 The state of AI: How organizations are rewiring to capture value McKinsey & Company 2025-03 Analyst-estimated global survey (1,993 respondents, 105 countries, fielded June to July 2025) finding 88 percent of organisations use AI in at least one function but only 5.5 percent report material financial returns, a cross-industry benchmark against which the regulated-sector figures in this lane should be read.
p19 Transforming Financial Analysis with NVIDIA NIM NVIDIA Technical Blog 2024 Vendor-published technical account of NIM microservices applied to financial-analysis workloads, useful for the inference-stack detail (containerised model serving, Triton backend) but a vendor claim rather than an independently verified bank deployment.

We use analytics cookies to understand site usage and improve the service. We do not use marketing cookies.