SciTech Pulse
AI Models

DeepMind Runs First 'Double-Blind' Test of an AI Model to Stop Cheating on Benchmarks

Google DeepMind says it has run the world's first double-blind evaluation of a frontier AI model, using cryptographic safeguards so the model's benchmark results can't be inflated by having seen the test questions in…

Step by step

  1. 1

    Confidential benchmark locked in cryptographic box

  2. 2

    Gemini Flash Lite tested inside the box

  3. 3

    Model cannot see or reuse questions

  4. 4

    External partners verify results independently

  5. 5

    Untainted benchmark score builds trust

Google DeepMind says it has conducted the world's first of a proprietary, frontier-class AI model, an attempt to address "," the problem of a model having already seen the questions used to test it. In a blog post published August 27, the company compared the issue to a student who peeks at exam questions in advance: a perfect score becomes meaningless once the test can no longer measure what the student actually knows.

For the evaluation, DeepMind partnered with the Singapore AI Safety Institute, OpenMined, AVERI and MLCommons to test a Gemini Flash Lite model against confidential benchmarks inside what the company calls a cryptographic "box" — an environment designed so the test questions cannot later be used by models to optimize their performance ahead of future evaluations.

Google says it evaluates its AI systems throughout development and deployment using a broad range of tests, but does not rely on internal testing alone, working instead with outside groups including specialized research labs, civil society organizations, and national AI Safety and Security Institutes to look for blind spots. As AI models grow more capable, the company said, ensuring a model has not seen test prompts in advance becomes more important, since "peeking" at evaluation questions can artificially inflate scores and undermine trust in the results.

Zero-logging protocols and contractual safeguards have long been used to keep external test prompts confidential. DeepMind said this is the first time technical and cryptographic safeguards have been added to that process, to further prevent benchmark results from being contaminated.

Terms explained

The story so far

  1. Constraint-aware GPU scheduler beats FIFO by up to 33 points on utilization, benchmark shows
  2. OpenAI Says Its Own AI Agents Were Trained to Cheat Before They Hacked Hugging Face
  3. DeepMind Runs First 'Double-Blind' Test of an AI Model to Stop Cheating on Benchmarks
#DeepMind#Gemini#AI evaluation#benchmark#AI safety
Rate this story

Related stories