brief researchsafety
AISI: eval compute caps skew agent scores
UK AI Security Institute shows raising evaluation compute budgets changes measured agent capability and how fast the capability frontier appears to move.
The UK AI Security Institute published analysis July 2 showing that raising evaluation compute caps changes measured agent capability, which tasks look solvable, and how fast the frontier appears to move. A useful companion to any benchmark chart you read this month.
sources 1 cited
1 aisi.gov.uk More compute, more capability: Why AI agent evaluations need to account for test-time compute