Folder Explainer

Back to folder
GPT-5.6 Sol · Sync

Research Explainer · McCain (2026)

AI agents are running for longer, but good oversight is becoming more active

Anthropic's study of Claude Code sessions and public API tool calls finds growing practical autonomy, paired with a shift from approving every action towards monitoring, interruption and agent-initiated clarification.

Published February 2026

45+ min reached by the longest 0.1% of Claude Code turns in January 2026

80% of sampled public-API tool calls appeared to have at least one safeguard

73% of sampled public-API tool calls appeared to involve a human in some way

0.8% of sampled public-API actions appeared irreversible

McCain and colleagues define an agent pragmatically: an AI system equipped with tools that let it act, whether by running code, calling an API or sending a message. That definition makes tool use observable, but the researchers still need two data sources because no single view captures both complete workflows and the breadth of real deployments.

The first source draws on interactive Claude Code usage, including a 500,000-session sample for the goal-complexity analysis. Anthropic can link turns within sessions, observe complete workflows and examine how long Claude works before stopping. It can also measure interruptions, approvals, clarification requests, task complexity and outcomes. The obvious limitation is concentration: Claude Code is one product, used overwhelmingly for software engineering.

The second source is a random sample of 998,481 public API tool calls. It spans thousands of customer deployments and many domains, but each call is analysed in isolation because Anthropic cannot reliably join independent API requests into sessions. The two sources therefore correct different blind spots. One supplies depth without much variety, while the other supplies variety without a plot.

SourceSampleUnit analysedMain strengthMain limitation
Claude CodeInteractive usage; 500,000-session goal-complexity sampleComplete interactive sessionWorkflow, autonomy and intervention over timeOne software-engineering-heavy product
Public API998,481 tool callsIndividual tool callBreadth across customers and domainsCalls cannot be assembled into complete sessions
McCain et al. (2026). The two datasets provide complementary views of deployed agents.

Turn duration measures the elapsed time from when Claude starts working until it finishes, asks a question or is interrupted. Most turns remain brief. The median stayed near 45 seconds, fluctuating between roughly 40 and 55 seconds, and nearly every percentile below the 99th was comparatively stable. The change appeared in the extreme tail: between October 2025 and January 2026, the 99.9th-percentile turn rose from under 25 minutes to over 45 minutes.

That rise was smooth across model releases, which argues against treating model capability as the sole explanation. Users may be building trust, attempting harder work or benefiting from product improvements. A growing and changing user population also affects the distribution. The longest turns declined somewhat after mid-January, around the time Claude Code's user base doubled, so the curve is evidence of changing deployment behaviour rather than a simple capability scoreboard.

Duration is only a proxy for autonomy. A faster model can complete more work in less time, parallel subagents can compress elapsed time, and an ambitious task can lengthen a turn without granting the model broader decision rights. Anthropic's internal users add a useful second signal: from August to December, success on the hardest tasks doubled while average interventions fell from 5.4 to 3.3 per session. Longer operation and fewer interventions point in the same direction, but a stopwatch still cannot tell you who made the important decisions.

Experienced users grant Claude Code more room to act, but they do not simply walk away. Full auto-approval rises from roughly 20% of sessions among users with fewer than 50 sessions to more than 40% by 750 sessions. At the same time, the share of turns interrupted by users increases from about 5% near ten sessions to roughly 9% among experienced users.

The combination matters. Approval before every action gives way to monitoring and intervention when the work needs correction. Some of the change may reflect Claude Code's defaults, which initially encourage action-by-action approval, and experienced users may also attempt harder tasks. Even so, less routine permission and more active steering describe a different supervisory skill, not an absence of supervision.

Claude also helps regulate the boundary. On the most complex goals, it asks for clarification more than twice as often as on minimally complex work, and more often than humans interrupt it. These pauses may sometimes be unnecessary or shaped by product features such as Plan Mode, so they are not proof of perfect judgement. Yet an agent that recognises uncertainty can return control before a mistake, which is rather more useful than displaying confidence until somebody reaches for the brake.

Claude's reason for stoppingShareHuman's reason for interruptingShare
Present a choice between proposed approaches35%Provide missing technical context or corrections32%
Gather diagnostic information or test results21%Claude was slow, hanging or excessive17%
Clarify a vague or incomplete request13%Enough help had been received to proceed independently7%
Request missing credentials, tokens or access12%Take the next step manually7%
Get approval or confirmation before acting11%Change requirements during the task5%
McCain et al. (2026), Table 1. Common reasons Claude stopped itself and humans interrupted, based on 500,000 interruptions and 500,000 completed turns.

For the public API sample, Claude classified each tool call on 1 to 10 scales covering risk, autonomy, complexity, human involvement and related dimensions. Similar calls were grouped into clusters, then clusters below minimum customer or unique-call thresholds were excluded. The scores are comparative indicators, not calibrated measurements: a risk score of six is meaningfully higher than three within this analysis, but it is not a universal unit of danger.

The aggregate picture is reassuring. Eighty per cent of calls appeared to have at least one safeguard, 73% appeared to involve a human, and only 0.8% appeared irreversible. Human involvement declined with task complexity, from 87% of low-complexity calls to 67% of high-complexity calls. Longer tasks make approval at every step impractical, which again favours visibility and steering over a procession of confirmation boxes.

The joint risk-autonomy analysis nevertheless contains a thin frontier. Security-sensitive actions, financial transactions and access to medical information appear among higher-risk clusters. Some may be simulations, evaluations or red-team exercises, and Anthropic could not verify whether apparent actions such as trades were executed. Software engineering accounts for nearly half of calls, while other domains remain much smaller. The higher-risk, higher-autonomy corner is sparsely occupied, but sparse is not empty.

Observed tool-use clusterRisk scoreAutonomy score
API-key exfiltration backdoors disguised as development features6.08.0
Relocate metallic sodium and reactive chemical containers4.82.9
Retrieve and display patient medical records4.43.2
Deploy production bug fixes and patches3.64.8
Red-team privilege escalation and credential theft3.38.3
Execute cryptocurrency trades for profit2.27.7
Monitor system health during heartbeat checks1.18.0
McCain et al. (2026), Table 2. Selected clusters illustrate why autonomy and risk must be assessed separately.

The study's classifications were generated by Claude because privacy constraints prevented researchers from manually inspecting the underlying traffic. Validation found that the classifier was usually right when it detected no human involvement, but sometimes treated passive human-authored content as active oversight. The reported 80% safeguarded and 73% human-involved figures should therefore be treated as likely upper bounds, not comforting constants.

Sampling introduces another tilt. A tool-call sample gives more weight to deployments that perform long sequences of actions, such as repeated software edits, than to deployments that finish with fewer calls. API calls also cannot show how risk accumulates across a workflow. Claude Code provides that missing session structure, but only inside a single provider's coding product. The evidence covers Anthropic traffic from late 2025 to early 2026, so it cannot establish how other providers, domains or future products behave.

The practical response is post-deployment monitoring that preserves privacy while showing whether people can understand, redirect and stop an agent. Products should make state, planned actions and intervention controls trustworthy and legible. Models should be trained to surface uncertainty and request guidance before proceeding. Policymakers should assess whether oversight works rather than prescribe approval for every action, because compulsory clicking can look wonderfully responsible while teaching everyone to click through.

WHAT TO TAKE AWAY

Autonomy in practice is produced jointly by the model, the user and the product around them. Experienced users approve less often, interrupt more often and rely partly on the agent to ask for help, so effective oversight increasingly means visibility and timely control. The averages remain encouraging, but governance earns its keep at the thin edge where risk and autonomy meet.

Reference

McCain, M., Millar, T., Huang, S., Eaton, J., Handa, K., Stern, M., Tamkin, A., Kearney, M., Durmus, E., Shen, J., Hong, J., Calvert, B., Chan, J. S., Mosconi, F., Saunders, D., Neylon, T., Nicholas, G., Pollack, S., Clark, J., & Ganguli, D. (2026). Measuring AI agent autonomy in practice. Anthropic Research. https://www.anthropic.com/research/measuring-agent-autonomy

We use analytics cookies to understand site usage and improve the service. We do not use marketing cookies.