Folder Explainer

Back to folder
GPT-5.6 Sol · Sync

Research Explainer · Vella (2026)

AI helps engineers move faster, but turns more of the job into supervision

Over six months, professional engineers reported spending less time writing code and shifting towards verification. Productivity remained positive, even as flow and other aspects of developer experience deteriorated for a growing minority.

Published May 2026

82% of participants reporting less time spent writing code by the second questionnaire

0.26 → 0.53 rise in the verification-minus-creation balance across the matched task-focus cohort

84% of continuing users reporting improved productivity at both survey waves

14% → 27% increase in matched participants reporting a worse experience on at least one dimension

Annie Vella and Kelly Blincoe followed professional software engineers through two online questionnaires, administered in October 2024 and April 2025. The six-month interval matters because AI coding tools were changing quickly, as were engineers’ habits. A single survey could describe enthusiasm or frustration on one particular day. Repeating the questions allowed the researchers to examine whether the same people’s perceptions moved.

The questionnaires combined rating scales with open-ended accounts. They measured perceived changes across designing, writing code, refactoring, reviewing, testing and debugging, alongside productivity and three dimensions of developer experience: cognitive load, feedback loops and flow state. Quantitative analysis used complete cases and non-parametric tests suited to ordered ratings. The authors also analysed written responses thematically, looking for recurring patterns in how engineers described their work.

The matched design is the study’s main strength, but it became progressively narrower. Of 224 initial responses, 158 participants were eligible at the first wave. The second produced 101 eligible participants, and 95 completed both. Item-level missingness reduced some analyses further, to 88 for task focus and 94 for developer experience. This is stronger than comparing two unrelated crowds, but not quite the same as following everyone.

Study stageQ1Q2 or matched
Recorded responses224111
Eligible participants158101
Matched longitudinal cohort9595
Matched task-focus analysis8888
Matched developer-experience analysis9494
Matched productivity analysis9595
Vella and Blincoe (2026). Recruitment, eligibility and matched samples across the longitudinal study.
Two side-by-side stacked bar charts compare engineers' reported task-focus changes at baseline and six-month follow-up across designing, writing, refactoring, testing, debugging, and reviewing code.
Engineers reported spending less time on writing and refactoring. At follow-up, reviewing code was the only task where more respondents reported increased rather than reduced time, signalling a move from creation to verification.

Figure 5 provides the clearest view of the task shift. Writing code had the lowest mean at both waves on a five-point scale where three meant no perceived change. By the second questionnaire, 82% said they spent less time writing code and only 2% said they spent more. Refactoring also remained below neutral. Designing and debugging changed little, while reviewing was the only task above neutral at both time points.

The six individual task changes did not remain statistically significant after correction for multiple comparisons. The authors therefore grouped designing, writing and refactoring as creation, and reviewing, testing and debugging as verification. The balance between those groups rose from 0.26 to 0.53. That shift was statistically significant, with an adjusted probability value of 0.006 and a moderate rank-biserial effect size of 0.39.

The averages conceal mixed personal trajectories. Writing remained stable for 56% of the task-focus cohort, while 30% moved towards still less writing. Testing moved upwards for 42%, and reviewing for 40%. Only 8% simultaneously reduced writing and increased reviewing, partly because most were already near the floor for writing at the first wave. The work moved, but not in one tidy procession.

Task or balanceQ1 meanQ2 meanChangeAdjusted p
Designing2.942.89-0.061.000
Writing code2.101.92-0.180.263
Refactoring code2.482.39-0.091.000
Reviewing code2.933.16+0.230.333
Testing2.522.77+0.250.263
Debugging2.842.840.001.000
Verification minus creation0.260.53+0.270.006
Vella and Blincoe (2026), Tables 2 and 4. Mean perceived task focus among 88 matched participants, where 1 means much less time, 3 no change and 5 much more time.

Less writing did not simply release engineers for leisurely architecture work. Participants also reported spending less time on several other conventional tasks, including testing. Their written accounts instead described a category that the usual software-development labels handle poorly: directing an assistant, inspecting its suggestions and repairing what it gets wrong.

The authors call this supervisory engineering work. Directing means expressing intent, supplying context and reformulating prompts when the model wanders. Evaluating means reading generated output and deciding what to accept, modify or reject. Correcting means fixing errors, integrating the result with an existing codebase and preserving consistency. These activities overlap with review, but they begin earlier and recur throughout generation.

Participants described faster boilerplate, debugging help and easier entry into unfamiliar technologies. They also described prompt loops, plausible mistakes, context switching and the need to corroborate answers. Trust therefore became something engineers actively calibrated rather than simply granted. The assistant can draft the change. It still cannot inherit the consequences.

The proposed supervisory work category contains three recurring responsibilities.

  1. DirectSpecify intent, provide relevant context, constrain the task and intervene when the model loses track.
  2. EvaluateRead the output critically and decide whether it should be accepted, changed or discarded.
  3. CorrectRepair errors, integrate generated code and maintain the standards and structure of the surrounding system.

Perceived productivity remained remarkably positive. At both waves, 84% said AI had improved their productivity, and 77% of the matched cohort gave exactly the same rating six months later. The mean moved only from 4.08 to 4.03 on the five-point scale. Participants associated the tools with faster starts, reduced repetitive effort and greater willingness to attempt unfamiliar work.

Developer experience moved less comfortably. Feedback loops improved significantly, suggesting that immediate suggestions and explanations helped engineers obtain information faster. Cognitive load and flow state declined, although neither change was statistically significant after correction. More tellingly, the share of matched participants with a negative rating on at least one experience dimension rose from 14% to 27%. Within that negative group, flow problems increased from 54% to 76%.

These findings do not show that AI caused worse experience, nor that objective output increased. Both productivity and experience were self-reported. They do show that feeling faster can coexist with more interruption, vigilance and correction. A developer can complete work sooner while enjoying the work less. Speed and a good working day are not synonyms.

MeasureQ1Q2ChangeAdjusted p
Feedback loops3.733.95+0.210.038
Cognitive load3.743.60-0.150.195
Flow state3.543.36-0.180.195
Perceived productivity4.084.03-0.050.326
Negative experience cohort14%27%+13 percentage pointsNot reported
Vella and Blincoe (2026), Table 6 and productivity results. Means use five-point perception scales among matched participants.

Forty per cent of eligible first-wave participants did not enter the matched cohort. Tests found no significant demographic differences between those retained and those lost, but unmeasured differences may remain. Engineers who abandoned AI tools or became disillusioned could be especially likely to disappear. The headline 84% therefore means 84% of continuing users, not 84% of everyone who tried an assistant.

The sample was 85% men, predominantly English-speaking and recruited through professional networks, online communities and nine participating organisations. The measures captured recollections and perceptions rather than delivery time, defects or other objective performance. Social desirability, effort justification, recent experiences and changing personal standards may all affect the ratings. The first author’s professional background may also have shaped the questions and interpretation, despite review with the research supervisor.

The tools themselves refused to sit still. Among matched participants, 82% changed their tool combinations, and average use rose from 1.9 to 2.9 tools. The study therefore cannot separate accumulated experience from product improvements or switching. Its evidence belongs to late 2024 and early 2025, before more autonomous agents became common. The safest conclusion is not that AI removes engineering effort, but that it moves the effort somewhere less visible.

Practical responses supported by the study include:

  • Measure experienceTrack flow, cognitive load and correction effort alongside perceived speed or output.
  • Reward judgementTreat verification, trust calibration and careful rejection of bad output as productive engineering work.
  • Teach supervisionDevelop skills in directing, evaluating and correcting AI output without sacrificing underlying technical understanding.
  • Keep claims boundedDistinguish continuing-user perceptions from objective productivity and from outcomes across all potential users.

WHAT CHANGED

AI coding assistants reduced perceived time spent writing code, but the saved effort did not simply vanish. It reappeared as direction, evaluation, correction and trust management, while positive productivity perceptions increasingly coexisted with disrupted flow. Organisations counting generated output may therefore miss the part of the job that has become harder to see.

Reference

Vella, A., & Blincoe, K. (2026). The Impact of AI Coding Assistants on Software Engineering: A Longitudinal Study. arXiv preprint arXiv:2605.23135. https://doi.org/10.48550/arXiv.2605.23135

We use analytics cookies to understand site usage and improve the service. We do not use marketing cookies.