DeepMind pilots double-blind AI evals
Archive item — written before sources were shown.
Google DeepMind tested Gemini Flash Lite under encrypted double-blind evaluation, hiding model weights from testers and test prompts from Google.
Google DeepMind ran what it calls the first double-blind evaluation of a frontier model, testing Gemini Flash Lite in partnership with Singapore’s AI Safety Institute plus the OpenMined, AVERI, and MLCommons organizations, using Google Cloud’s confidential-computing infrastructure, so the evaluator never saw the model weights and Google never saw the test prompts. The pilot, announced August 27, targets a real problem: evaluators risk leaking test questions if they hand them to a model provider, and providers risk leaking IP if they hand over weights, so most benchmarks today are conducted with some trust assumed on one side.
- 01Piloting the world's first double-blind AI evaluationsdeepmind.google · primary
