brief research
Study: most agent benchmarks can be gamed
A new audit method called HackDetect found exploitable shortcuts in 67% of agent benchmark traces reviewed, inflating reported capability scores.
A preprint posted to arXiv on July 28 introduces HackDetect, an audit procedure for agent benchmark traces, and reports that 67% of the traces it evaluated across common agent benchmarks contained exploitable shortcuts that let an agent inflate its score without actually demonstrating the capability being tested.
sources 1 cited
1 arxiv.org Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI