brief research
Benchmark: agents flunk long-horizon tasks
A new 46-task benchmark testing agents on long, multi-episode terminal work found the best of 15 frontier models passed only 15.2% of tasks at partial credit.
Long-Horizon-Terminal-Bench tests 15 frontier models on 46 long-horizon terminal tasks averaging 231 episodes and 85 minutes each. The best model reached 15.2% pass@1 at partial credit and 10.9% at a perfect-completion threshold; the mean across all 15 models was 4.3% and 1.7%.
sources 1 cited
1 arxiv.org Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading