Google DeepMind has published results from what it calls the world's first double-blind AI evaluation pilot, a methodology borrowed from clinical trials designed to reduce bias in how AI models are assessed.
In the double-blind setup, human evaluators do not know which model produced the outputs they are rating, and the model names are masked throughout the evaluation process. The approach aims to eliminate brand bias — the tendency for evaluators to rate outputs from well-known models more favorably.
The pilot tested the framework across multiple task categories and found measurable differences in results compared to standard open-label evaluations. In some cases, outputs from smaller or less-known models scored higher when evaluators could not see the model name.
DeepMind is publishing the methodology openly and inviting other labs to adopt it, arguing that more rigorous evaluation standards are needed as AI models become increasingly competitive.