archive
· today in ai · 2026-07-03
AISI: eval compute caps skew agent scores
Archive item — written before sources were shown.
UK AI Security Institute shows raising evaluation compute budgets changes measured agent capability and how fast the capability frontier appears to move.
The UK AI Security Institute published analysis July 2 showing that raising evaluation compute caps changes measured agent capability, which tasks look solvable, and how fast the frontier appears to move. A useful companion to any benchmark chart you read this month.
sources
