Piloting the world's first double-blind AI evaluations
Quick Answer
Google DeepMind introduces the first double-blind evaluation for AI models, utilizing cryptographic methods to prevent benchmark contamination.
Quick Take
This initiative, in collaboration with the Singapore AI Safety Institute and others, aims to enhance trust in AI performance assessments by ensuring models cannot access test prompts beforehand.
Key Points
- Double-blind evaluations prevent AI models from seeing test prompts in advance.
- Collaboration includes Singapore AI Safety Institute, OpenMined, and MLCommons.
- Cryptographic safeguards enhance the integrity of AI model assessments.
- Benchmark contamination can artificially inflate AI performance scores.
- External partners help identify blind spots in AI model evaluations.
📖 Reader Mode
~2 min readAugust 27, 2026 Responsibility & Safety
William Isaac, Sol Messing and Kristian Lum
Building trust in proprietary model benchmarks using cryptographically secure environments
Imagine a student is set to take a high-stakes exam. If they accidentally peek at the test questions in advance, achieving a perfect score is influenced by this knowledge, making it a meaningless accomplishment. To truly measure what they know, they must have no visibility of the test questions until it's time to take the exam. That is the exact challenge the industry faces when evaluating advanced AI models. If a model has already seen the test questions - a problem known as benchmark contamination - the results can only be trusted to an extent.
Today, we’re introducing the world’s first double-blind evaluation of a proprietary, frontier class AI model, which keeps external evaluations confined to a cryptographic “box” where they can’t be used by models later to optimize performance ahead of testing. We're partnering with the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons, to test a Gemini Flash Lite model against confidential benchmarks in a privacy-preserving environment, increasing evaluation integrity.
At Google, we assess our AI systems using a broad spectrum of evaluations throughout model development and deployment, but we don’t rely on internal testing alone. To identify potential blindspots, we work with a diverse group of external partners, including specialized research labs, civil society and national AI Safety and Security Institutes (AISIs), using their unique expertise to stress-test our models.
As AI models become more capable, ensuring the model has not seen the test questions or prompts in advance is critical, as this can skew the results. Policymakers, researchers, and enterprises need to trust that AI benchmarks accurately reflect a model's true capabilities and safety, but if models are able to “peek” at the evaluation questions in advance, it can artificially inflate scores and undermine this trust.
Although zero-logging protocols and rigorous contractual safeguards have long kept external test prompts confidential, incorporating technical and cryptographic safeguards marks a major step forward in secure model evaluation.
How double-blind evaluations work
— Originally published at deepmind.google
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from Google DeepMind
See more →
Introducing Gemma 4 12B: a unified, encoder-free
Google DeepMind has introduced Gemma 4 12B, a unified, encoder-free multimodal model designed to enhance performance across various tasks. This model aims to streamline processes in AI applications by eliminating the need for traditional encoders, potentially improving efficiency and reducing costs for developers and researchers in the field.
