tuesday, october 6, 2026 · the day's ai, attributed published by trilot llc · wyoming
archive · today in ai · 2026-07-13

Benchmark: agents flunk long-horizon tasks

Archive item — written before sources were shown.

A new 46-task benchmark testing agents on long, multi-episode terminal work found the best of 15 frontier models passed only 15.2% of tasks at partial credit.

Long-Horizon-Terminal-Bench tests 15 frontier models on 46 long-horizon terminal tasks averaging 231 episodes and 85 minutes each. The best model reached 15.2% pass@1 at partial credit and 10.9% at a perfect-completion threshold; the mean across all 15 models was 4.3% and 1.7%.

sources
  1. 01Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Gradingarxiv.org · primary (paper)
Rami Steitieh
Rami Steitieh

Builder and operator. Runs 17 content sites and Trilot LLC on the tools reviewed here.