Research · Academic & arXiv

Back to sweep

Research sweep · deep · 2025 – 2026

Observability for Agent Orchestration, Pipelines and Swarms

Observability for LLM agent orchestration, pipelines and multi-agent swarms between August 2025 and August 2026: emerging standards (OpenTelemetry GenAI semantic conventions, MCP tool tracing), the tooling landscape (LangSmith, Langfuse, Arize Phoenix, Braintrust, W&B Weave, Comet Opik, Datadog LLM Observability), the two-layer pattern of durable orchestration plus agent tracing (Temporal, LangGraph), and how production operators such as Stripe, Shopify, Airbnb and Anthropic instrument, evaluate and debug agent fleets at scale

  • Claude Fable 5
  • financial
  • frontier
  • academic
  • vc
  • blogs
  • tech

Synthesised 2026-08-25

Narrative

Academic engagement with agent observability splits into three threads: benchmark and evaluation science from METR, taxonomy-building software-engineering papers on "AgentOps," and a fast-growing empirical literature on multi-agent failure modes and attribution, much of it explicitly framed as a debugging or observability problem rather than a capability problem.

METR's task-suite papers are the most rigorous quantitative anchor in the lane. HCAST (Rein et al., arXiv:2503.17354) introduces 189 machine-learning, cybersecurity, software-engineering and reasoning tasks with 563 human baseline attempts totalling over 1,500 hours, and reports that current agents succeed 70-80% of the time on tasks that take humans less than one hour


Sources

ID Title Outlet Date Significance
a1 AgentFixer: From Failure Detection to Fix Recommendations in LLM Agentic Systems arxiv.org Retrieved by this lane's web search.
a2 [2507.21504] Evaluation and Benchmarking of LLM Agents: A Survey arxiv.org July 29, 2025 Retrieved by this lane's web search.
a3 When Tools Fail: Benchmarking Dynamic Replanning and Anomaly Recovery in LLM Agents arxiv.org Retrieved by this lane's web search.
a4 From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review arxiv.org Retrieved by this lane's web search.
a5 Evaluation and Benchmarking of LLM Agents: A Survey arxiv.org Retrieved by this lane's web search.
a6 CATArena: Evaluating Evolutionary Capabilities of Code Agents via Iterative Tournaments arxiv.org Retrieved by this lane's web search.
a7 Who Tests the Testers? Systematic Enumeration and Coverage Audit of LLM Agent Tool Call Safety arxiv.org Retrieved by this lane's web search.
a8 Semantic Conventions for Generative AI Agentic Systems (gen_ai.*) · Issue #35 · open-telemetry/semantic-conventions-genai github.com August 20, 2025 Retrieved by this lane's web search.
a9 Semantic Conventions for Generative AI Agentic Systems (gen_ai.*) · Issue #2664 · open-telemetry/semantic-conventions github.com August 20, 2025 Retrieved by this lane's web search.
a10 OpenTelemetry GenAI Semantic Conventions | MLflow AI Platform mlflow.org Retrieved by this lane's web search.
a11 AI Agent Observability - Evolving Standards and Best Practices | OpenTelemetry opentelemetry.io February 24, 2026 Retrieved by this lane's web search.
a12 Reasoning Provenance for Autonomous AI Agents: Structured Behavioral Analytics Beyond State Checkpoints and Execution Traces arxiv.org Retrieved by this lane's web search.
a13 Inside the LLM Call: GenAI Observability with OpenTelemetry | OpenTelemetry opentelemetry.io May 14, 2026 Retrieved by this lane's web search.
a14 Time Horizon 1.1 - METR metr.org January 29, 2026 Retrieved by this lane's web search.
a15 Notes on the Long Tasks METR paper, from a HCAST task contributor - LessWrong 2.0 viewer greaterwrong.com Retrieved by this lane's web search.
a16 Notes on the Long Tasks METR paper, from a HCAST ... lesswrong.com Retrieved by this lane's web search.
a17 METR Time Horizons | Epoch AI epoch.ai Retrieved by this lane's web search.
a18 GitHub - METR/RE-Bench · GitHub github.com Retrieved by this lane's web search.
a19 Are We There Yet? Evaluating METR’s Eval of AI’s Ability to Complete Tasks of Different Lengths empiricrafting.substack.com December 15, 2025 Retrieved by this lane's web search.
a20 BRIDGE: Predicting Human Task Completion Time From Model Performance arxiv.org Retrieved by this lane's web search.
a21 Interpreting the METR Time Horizons Post lesswrong.com April 29, 2025 Retrieved by this lane's web search.
a22 LangGraph vs Temporal: AI Agent Orchestration Compared langchain.com July 17, 2026 Retrieved by this lane's web search.
a23 LangGraph in production: Temporal's LangGraph Plugin adds Durable Execution | Temporal temporal.io July 16, 2026 Retrieved by this lane's web search.
a24 Kinde Orchestrating Multi-Step Agents: Temporal/Dagster/LangGraph Patterns for Long-Running Work kinde.com Retrieved by this lane's web search.
a25 LangGraph vs Temporal for AI Agents: Durable Execution Architecture Beyond For Loops | by Anubhav | Data Science Collective | Medium medium.com March 19, 2026 Retrieved by this lane's web search.

We use analytics cookies to understand site usage and improve the service. We do not use marketing cookies.