Research · Academic & arXiv
Back to sweepResearch sweep · deep · 2022 – 2026
AI and the Labour Market, Signal, Noise and Structural Change
Evidence of AI-related labour-market change from November 2022–August 2026: employment, vacancies, wages and productivity; occupational and demographic distribution; job creation, displacement and work redesign; and competing economic interpretations using ONS, BLS, Eurostat, OECD, ILO and IMF evidence.
- GPT-5.6-sol
- financial
- frontier
- academic
- vc
- blogs
- tech
Synthesised 2026-08-14
Narrative
Foundational post-ChatGPT research separates potential exposure from realised adoption. Eloundou, Manning, Mishkin and Rock’s 2023 GPTs are GPTs estimates task exposure rather than job loss, whereas Humlum and Vestergaard’s Danish linked survey-register work measures workplace adoption and administrative earnings and hours. Bick, Blandin and Deming find rapid US take-up, but only a small share of aggregate work time assisted, while Brynjolfsson, Li and Raymond’s staggered rollout establishes a causal productivity gain for customer-support workers, concentrated among novices and lower-performing staff.
The strongest evidence on labour demand is mixed and remains preliminary. Liu, Wang and Yu’s 285-million-posting difference-in-differences design reports a relative 9% decline in postings for highly substitutable occupations after November 2022, growing over time and concentrated in administration and professional services; Chandar’s CPS study finds no corresponding average employment or earnings divergence across highly exposed occupations through early 2025. Several studies instead identify adjustment in hiring composition: Hosseini Maasoum and Lichtinger use 65 million US résumés to associate firm-level GenAI integration with slower junior hiring, while Mahieu finds a roughly 23% fall in entry-level vacancies between high- and low-exposure occupations in Flanders, with no comparable experienced-worker decline.
Experimental productivity evidence supports work redesign rather than a settled conclusion about aggregate displacement. Dillon, Jaffe, Immorlica and Stanton’s six-month experiment across 66 firms found two fewer weekly email hours among users but no detected change in task quantity or composition. Ni and colleagues’ Alibaba experiment finds faster service and better customer-rated quality, especially for lower performers, but no objective-quality improvement and deterioration among top agents. METR’s HCAST and RE-Bench evidence documents rising autonomous software and AI-research task capability, but it is a capability benchmark, not evidence that employers have substituted agents for workers; METR’s own developer RCT update cautions that benchmark success can fail to produce mergeable work in real repositories.
The most credible interpretation through August 2026 is therefore one of early, uneven reorganisation. Administrative records in Denmark rule out earnings or recorded-hours effects above 2% two years after ChatGPT despite widespread adoption, whereas online vacancies and résumé data suggest that firms may first reduce or redesign junior recruitment. Job-posting analyses cannot by themselves distinguish AI causation from the post-pandemic correction, high interest rates, sectoral weakness or changing recruitment practices. Biswas’s six-country software-posting study makes that disconfirming point directly: developer-posting contraction began after the 2022 peak, roughly two years before the late-2024 rise in AI-related demand.
Sources
| ID | Title | Outlet | Date | Significance |
|---|---|---|---|---|
| a1 | GPTs are GPTs: An Early Look at the Labor Market Impact Potential of Large Language Models | arXiv | 2023-03 | Eloundou, Manning, Mishkin and Rock provide the foundational task-exposure framework, estimating potential LLM effects across US occupations rather than realised employment outcomes. |
| a2 | Generative AI at Work | National Bureau of Economic Research | 2023-04 | Brynjolfsson, Li and Raymond use a staggered workplace rollout among 5,179 customer-support agents to estimate a causal 14% productivity increase, concentrated among novice and lower-skilled workers. |
| a3 | The Adoption of ChatGPT | IZA Discussion Paper | 2024-05 | Humlum and Vestergaard link a large Danish survey experiment to register data, documenting adoption patterns and employer restrictions among workers in 11 exposed occupations. |
| a4 | The Rapid Adoption of Generative AI | National Bureau of Economic Research | 2024-09 | Bick, Blandin and Deming provide nationally representative US survey evidence on work and household use, including estimates of work hours assisted and self-reported time savings. |
| a5 | Generative AI Impact on Labor Market: Analyzing ChatGPT's Demand in Job Advertisements | arXiv | 2024-12 | Ahmadi, Khosh Kheslat and Akintomide use US advertisements from May to December 2023 to identify ChatGPT-related skill clusters, measuring employer demand for AI skills rather than realised jobs. |
| a6 | Augmenting or Automating Labor? The Effect of AI Development on New Work, Employment, and Wages | arXiv | 2025-03 | This historical US study distinguishes labour-augmenting from automating AI exposure between 2015 and 2022, offering a pre-ChatGPT baseline for interpreting wage inequality and new-work creation. |
| a7 | Generative AI Adoption and Higher Order Skills | arXiv | 2025-03 | Gulati, Marchetti, Puranam and Sevcenko analyse postings at 378 US public firms recruiting for GenAI skills, finding higher advertised cognitive and social-skill requirements in GenAI roles. |
| a8 | Shifting Work Patterns with Generative AI | National Bureau of Economic Research | 2025-05 | Dillon, Jaffe, Immorlica and Stanton report a randomised experiment across 66 firms and 7,137 knowledge workers, finding time savings in email but no detected changes in task volume or composition. |
| a9 | Still Waters, Rapid Currents: Early Labor Market Transformation under Generative AI | National Bureau of Economic Research | 2025-05 | Humlum and Vestergaard combine Danish adoption surveys with administrative worker and workplace records, finding precise null effects on earnings and hours while documenting task reorganisation and occupational switching. |
| a10 | Tracking Employment Changes in AI-Exposed Jobs | SSRN | 2025-06 | Chandar uses US CPS data from Q4 2022 to Q1 2025 and finds no average employment or earnings-growth gap by GenAI exposure, while showing heterogeneity between software and customer-service occupations. |
| a11 | Generative AI as Seniority-Biased Technological Change: Evidence from U.S. Résumé and Job Posting Data | SSRN | 2025-08 | Hosseini Maasoum and Lichtinger analyse 65 million US résumés across more than 280,000 firms and associate GenAI-integrator hiring with lower junior employment driven by slower hiring. |
| a12 | Labor Demand in the Shadow of Generative AI: Evidence from the U.S. Job Posting Data | SSRN | 2025-09 | Liu, Wang and Yu apply difference-in-differences and event studies to 285 million US postings from 2018Q1 to 2025Q2, reporting relative declines in high-substitution occupations while controlling for interest-rate changes. |
| a13 | Generative AI and Firm Productivity: Field Experiments in Online Retail | arXiv | 2025-10 | Fang, Yuan, Zhang, Donati and Sarvary report large-scale randomised experiments across seven retail workflows, providing causal evidence on firm-level productivity and heterogeneous gains for smaller sellers and less experienced consumers. |
| a14 | Generative AI in Action: Field Experimental Evidence from Alibaba's Customer Service Operations | arXiv | 2026-02 | Ni and colleagues use a large Alibaba customer-service experiment to separate access from use, finding faster service and heterogeneous quality effects across worker performance levels. |
| a15 | Hiring Up, Not Down: Generative AI and the Composition of Labor Demand | SSRN | 2026-05 | Mahieu uses administrative vacancy data for Flanders, 2021 to 2025, and estimates a roughly 23% lower entry-level vacancy rate between the 25th and 75th exposure percentiles at peak adoption. |
| a16 | Generative AI and the Reorganization of Labor Demand | arXiv | 2026-05 | Wang, Wei and Wang use nationwide US posting data to decompose labour-demand adjustment into hiring reallocation and within-job task redesign, with distinct patterns by seniority. |
| a17 | Human Capital, AI, and Labor Commoditization | arXiv | 2026-06 | Siddiq and Zhang analyse Upwork data around ChatGPT's release and report that, in more AI-exposed categories, labour demand places less weight on human-capital signals and more on price. |
| a18 | Generative AI and the Informational Value of Educational Credentials: Evidence from the Master's Margin | SSRN | 2026-07 | Cortes, Dellarocas and Wang use Revelio worker-flow data from 2018 to 2025 to examine whether AI exposure changes employers' use of master's credentials in hiring. |
| a19 | Junior in Title, Senior in Practice: Job Posting Evidence on Entry-Level Software Hiring in the Age of Generative AI | SSRN | 2026-07 | Biswas offers a disconfirming descriptive study across six national software-posting series, finding that the hiring decline predated late-2024 AI-demand growth and warning against simple AI attribution. |
| a20 | HCAST: Human-Calibrated Autonomy Software Tasks | METR | 2025-03 | METR's HCAST supplies human-calibrated measures of agents' autonomous performance on more than 180 software, cybersecurity, machine-learning and reasoning tasks, informing capability rather than labour-market impact. |
| a21 | Evaluating frontier AI R&D capabilities of language model agents against human experts | METR | 2024-11 | METR's RE-Bench work compares frontier agents with human experts on ML research-engineering tasks, giving a bounded leading indicator for possible work substitution in technical research. |
| a22 | Research Update: Algorithmic vs. Holistic Evaluation | METR | 2025-08 | METR connects its developer productivity randomised trial with agent-benchmark results, showing why automatic task success need not yield mergeable work in realistic software development. |
| a23 | How Does Time Horizon Vary Across Domains? | METR | 2025-07 | Kwa and Cheng report METR estimates of frontier-model task time horizons across HCAST, RE-Bench and related suites, while explicitly limiting inference to software and research domains. |
| a24 | Generative AI at Work: From Exposure to Adoption across 35 European Countries | arXiv | 2026-04 | This study uses the 2024 European Working Conditions Survey of more than 36,600 workers to separate occupational exposure from self-reported adoption and early task restructuring across Europe. |