Research · Frontier Lab & Model News
Back to sweepResearch sweep · deep · 2022 – 2026
AI and the Labour Market, Signal, Noise and Structural Change
Evidence of AI-related labour-market change from November 2022–August 2026: employment, vacancies, wages and productivity; occupational and demographic distribution; job creation, displacement and work redesign; and competing economic interpretations using ONS, BLS, Eurostat, OECD, ILO and IMF evidence.
- GPT-5.6-sol
- financial
- frontier
- academic
- vc
- blogs
- tech
Synthesised 2026-08-14
Narrative
Frontier-lab releases shifted materially from conversational assistance towards tool-using reasoning and software agents during 2025. OpenAI’s o3 and o4-mini, Anthropic’s Claude 4, Google DeepMind’s Gemini 2.5, Mistral’s Devstral, and xAI’s Grok 4 all presented coding, multi-step reasoning and tool use as central capabilities rather than peripheral features. Meta’s Llama 4 made multimodal mixture-of-experts models available as open weights, while Mistral positioned an Apache-licensed coding agent as an alternative route to deployment.
The evidence most directly relevant to work comes from Anthropic’s usage research, not from release benchmarks. Its Economic Index reports show Claude use concentrated in computer and mathematical tasks, with growing use of coding-oriented systems and some evidence of higher delegation. These are leading indicators of task redesign and potential exposure, not evidence that employment, vacancies, wages or aggregate productivity have already changed because of AI: the samples cover Claude users and transcripts, are not representative of workers or firms, and frequently depend on model-assisted classifications and counterfactual time estimates.
Independent evidence tempers the labs’ productivity narrative. METR’s randomised study of 16 experienced open-source developers working on familiar repositories found that access to contemporary AI tools made tasks 19% slower on average, despite participants and experts expecting gains. METR’s time-horizon work nevertheless records rapidly improving autonomous completion of bounded software tasks, while its 2026 risk report documents continued testing of named frontier systems. The appropriate labour-market interpretation is therefore mixed: technical capability and API availability have advanced quickly, especially in software and professional knowledge work, but realised workplace effects remain contingent on task selection, integration costs, verification, worker experience and organisational redesign.
Sources
| ID | Title | Outlet | Date | Significance |
|---|---|---|---|---|
| t1 | Introducing OpenAI o3 and o4-mini | OpenAI | 2025-04 | OpenAI’s April 2025 release describes reasoning models with integrated web, code, file and image tools, marking a move from answer generation towards agentic execution of professional tasks. |
| t2 | Introducing GPT-5 for developers | OpenAI | 2025-08 | This API announcement frames GPT-5 as OpenAI’s model for coding and agentic tasks and documents new controls for tool calling, reasoning effort and cost-performance choices. |
| t3 | GPT-5 System Card | OpenAI | 2025-08 | OpenAI’s system card records the GPT-5 model family, its routing design, safety evaluations and the decision to apply high biological and chemical capability safeguards to GPT-5-thinking. |
| t4 | Introducing GPT-5.2 | OpenAI | 2025-12 | OpenAI presents GPT-5.2 as a model for professional work and long-running agents, including performance claims on GDPval, coding and tool use that are relevant to task-level workplace substitution claims. |
| t5 | Introducing Claude 4 | Anthropic | 2025-05 | Anthropic’s Claude 4 announcement documents Opus 4 and Sonnet 4 as coding and long-running agent models, including tool use during extended reasoning and persistent-memory features. |
| t6 | Claude 4 System Card | Anthropic | 2025-05 | Anthropic’s primary technical safety document reports evaluations, threat models and Responsible Scaling Policy treatment for Claude Opus 4 and Sonnet 4. |
| t7 | Anthropic Economic Index: Insights from Claude 3.7 Sonnet | Anthropic | 2025-03 | Using observed Claude.ai interactions after the Claude 3.7 release, this report finds increasing coding, education, science and healthcare use, but measures product usage rather than labour-market outcomes. |
| t8 | Anthropic Economic Index: AI’s impact on software development | Anthropic | 2025-04 | This analysis of 500,000 coding-related Claude interactions distinguishes ordinary Claude.ai use from Claude Code agent use and provides a direct leading indicator for work redesign in software occupations. |
| t9 | Estimating AI productivity gains from Claude conversations | Anthropic | 2025-11 | Anthropic extrapolates from 100,000 sampled conversations and model-estimated task times to a possible 1.8 percentage point annual productivity-growth increase, while explicitly stating that this is not a forecast of realised productivity. |
| t10 | Anthropic Economic Index report: Economic primitives | Anthropic | 2026-01 | This report adds task complexity, skill, purpose, autonomy and success measures to transcript analysis, finding concentrated use in high-human-capital tasks but acknowledging that observed conversations do not straightforwardly map to real-world job change. |
| t11 | Anthropic Economic Index report: Cadences | Anthropic | 2026-06 | Anthropic’s June 2026 report combines usage measures with a user survey and explicitly documents severe occupational selection, with computer and mathematical workers heavily over-represented among respondents. |
| t12 | Gemini 2.5: Our most intelligent AI model | Google DeepMind | 2025-03 | Google’s Gemini 2.5 announcement presents a reasoning-first model with claimed advances in coding, mathematics and science, showing the rapid spread of test-time reasoning across frontier systems. |
| t13 | Model cards | Google DeepMind | 2026-07 | Google DeepMind’s model-card index provides a structured record of Gemini releases and updates, including Gemini 2.5 Pro, Deep Think and Computer Use, useful for tracking capability diffusion into agentic workflows. |
| t14 | Advancing Gemini’s security safeguards | Google DeepMind | 2025-05 | This technical safety post identifies indirect prompt injection as a practical risk for agents that access workplace emails, documents and websites, a constraint on autonomous work deployment. |
| t15 | The Llama 4 herd: The beginning of a new era of natively multimodal AI innovation | Meta AI | 2025-04 | Meta introduces open-weight Llama 4 Scout and Maverick as multimodal mixture-of-experts models with long context, expanding the set of systems organisations can run or customise outside closed-model APIs. |
| t16 | Everything we announced at our first-ever LlamaCon | Meta AI | 2025-04 | Meta’s LlamaCon announcement adds a limited-preview Llama API, customisation tools and security evaluation tools, indicating a broader enterprise deployment pathway for open-weight models. |
| t17 | Au Large | Mistral AI | 2024-02 | Mistral’s Mistral Large launch is an early European frontier-model release focused on multilingual reasoning, code generation and API distribution, relevant to the widening supplier base after 2022. |
| t18 | Devstral | Mistral AI | 2025-05 | Mistral and All Hands AI describe Devstral as an Apache-licensed agentic model trained on real GitHub issues, directly targeting software-engineering work rather than isolated code completion. |
| t19 | Latest news | Mistral AI | 2025-07 | Mistral’s primary news archive records the rapid 2025 rollout of agents, coding models, enterprise products, OCR and reasoning systems, helping establish the pace and breadth of deployment-oriented releases. |
| t20 | Grok 4 | xAI | 2025-07 | xAI’s Grok 4 release claims improved reasoning, native tool use and agentic benchmark performance, including a high-cost Grok 4 Heavy configuration based on parallel test-time compute. |
| t21 | Grok 4 Model Card | xAI | 2025-08 | xAI’s model card is the primary safety and capability documentation for Grok 4, covering concerning propensities, dual-use evaluations and deployment contexts. |
| t22 | Grok 4 Fast | xAI | 2025-09 | This release shows the commercial push to lower the cost of reasoning and search-enabled agents, a key condition for broad workplace adoption even when absolute model capability changes little. |
| t23 | Measuring AI Ability to Complete Long Tasks | METR | 2025-03 | METR’s reproducible time-horizon evaluation estimates that the duration of software tasks frontier agents can complete with 50% reliability has doubled roughly every seven months, but its extrapolation remains a forecast rather than labour-market evidence. |
| t24 | Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity | METR | 2025-07 | METR’s randomised controlled trial of 16 experienced developers and 246 tasks found AI access made work 19% slower on average, providing an important counterweight to vendor productivity claims. |
| t25 | Frontier Risk Report (February to March 2026) | METR | 2026-06 | METR’s external risk report records evaluations of frontier systems and their performance under substantial agentic task budgets, offering independent evidence on capability progression and remaining limits. |