Agents struggle with long-horizon tasks

Last updated
Agent-1 is bad at even simple long-horizon tasks (page 7, Early 2026 section). Also: agents in Mid 2025 are 'impressive in theory but in practice unreliable.'

At a glance

  • Assessment: Confirmed
  • Confidence in assessment: 85%
  • Predicted timing: 2025
  • Primary source: ai-2027.com, pages 3 (Mid 2025) and 7 (Early 2026)

What AI 2027 Predicted

The scenario predicts that while agents become useful for short tasks, they struggle with sustained, long-horizon work — tasks that require hours or days of coherent effort. This is framed as a temporary limitation that improves rapidly.

How We Track This

We monitor:

  • METR time horizon benchmarks (50% and 80% thresholds)
  • Reports on agent task completion rates for multi-hour workflows
  • Enterprise adoption patterns for long-running agent tasks
  • Academic research on agent planning and coherence over time

Current Evidence

This prediction was accurate for 2025, but the frontier has moved quickly since then. METR’s current Time Horizon 1.1 data estimates Claude Opus 4.6 at a 50% time horizon of roughly 11h59m and an 80% time horizon of roughly 1h10m. That is not reliable full-work-day autonomy, but it does show that structured software/ML/cybersecurity tasks requiring many human-hours are increasingly within reach at lower reliability. Published evaluations include failures on multi-day open-ended projects, without establishing a universal failure rate, and METR cautions that time horizon is a task-difficulty measure rather than the amount of time an AI can act autonomously.

Sources:

Counterevidence & Limitations

  • NVIDIA’s AVO system ran a supervised seven-day attention-kernel search, explored more than 500 directions, and committed 40 versions. NVIDIA reported kernels up to 10.5% faster than FlashAttention-4 on evaluated DGX B200 configurations. This shows that persistent memory, supervision, and tool use can support productive week-long execution in a narrow engineering domain. It does not establish broad reliability across open-ended tasks or unsupervised deployment.
  • METR tasks are unusually well-specified and evaluable compared with much real work; the suite is not entirely synthetic.
  • Improvements in time horizon don’t necessarily translate to reliability on diverse, open-ended tasks — agents may improve on structured tasks while still failing on ambiguous real-world problems
  • Some companies report agent failures on production workloads even when benchmarks improve, suggesting a gap between controlled and deployed performance
  • The prediction is becoming harder to score as the situation evolves: the 2025 claim is confirmed, but the rapid improvement trajectory means the “struggle” characterization has a short shelf life
  • There is no standardized definition of “long-horizon” — different benchmarks use different task lengths, making cross-study comparison difficult

What Would Change Our Assessment

  • Historical note: This prediction specifically targeted 2025 and has been confirmed. Subsequent gains should be recorded without changing the historical 2025 judgment.
  • Watch for: Reliable measurements on longer and messier tasks, alongside complete-project outcomes and documented assistance. Beyond-range horizon estimates need additional validation.

Update History

DateUpdate
2026-09-06Current evidence and assessment explanation reviewed. Clarified measurement scope, source interpretation and limitations.
2026-08-24NVIDIA reported a supervised AVO agent completing a productive seven-day kernel-optimization run with more than 500 explored directions and 40 committed versions. This adds a concrete long-horizon success case, while the narrow task and system-level supervision limit generalization. Status and confidence remain unchanged.
2026-06-06Corrected the Opus 4.6 METR framing: current TH1.1 raw data estimates roughly 11h59m at the 50% horizon and roughly 1h10m at the 80% horizon, not 14.5h at 80% reliability. Status unchanged because the 2025-era “long-horizon struggle” claim remains historically confirmed, while newer capability gains are tracked as a rapidly changing limitation.
2026-03Prediction confirmed for 2025 timeframe. Agents still struggle with multi-day autonomous work, though rapid improvement visible in early 2026.
2026-01METR Time Horizon 1.1: Claude Opus 4.5 at ~4h49m (Jan 2026). Agents now sustain coherent work for nearly 5 hours — well beyond the “simple long-horizon tasks” the prediction describes. The “struggle” characterization from mid-2025 is increasingly outdated for frontier models, though it remains accurate for most publicly available agents.
2025-09Scale AI SWE-Bench Pro (Sep 19): models scoring 70%+ on standard SWE-bench drop to ~23% on long-horizon tasks involving multi-file refactors and cross-repository changes. Long-horizon capability gap now formally measured.
2025-07METR domain time-horizon analysis (July 14): 50% success horizon for frontier models at ~50 minutes on human expert tasks. Tasks requiring hours of sustained autonomous work remain beyond current capability.
2025-06SWE-bench and GAIA benchmarks confirm agents fail on tasks requiring sustained multi-hour work. Error accumulation over long horizons remains the primary failure mode.