A double-blind pilot puts closed-model evaluation inside a cryptographic boundary
DeepMind and external evaluation partners report a way to test a proprietary model without revealing either its weights or the evaluator's private prompts.

What happened
Google DeepMind and four partner organisations reported a live pilot in which Gemini 2.5 Flash Lite and confidential safety prompts were evaluated inside an attested confidential-computing environment.
Why it matters
The pilot addresses a real integrity problem for closed-model evaluation: an evaluator wants to protect unseen test material while a model owner wants to protect proprietary weights and inference code.
Google DeepMind announced a pilot with Singapore AISI, OpenMined, AVERI and MLCommons in which a proprietary model and private evaluation prompts were brought together inside a protected computing environment. The participants tested Gemini 2.5 Flash Lite on reserve material from the MLCommons AILuminate corpus and on Singapore-focused harmful-content prompts.
This is not a new capability result. The accompanying technical report publishes no benchmark scores. Its contribution is the evaluation process: the model owner provides weights and inference code, the evaluator provides prompts and test code, and both parties check a cryptographic attestation before releasing either asset into an ephemeral CPU/GPU enclave. Only an agreed result leaves the environment. In this use of “double-blind”, each party keeps its core asset hidden from the other; it is not the blinding design used in a clinical trial.
The immediate implication for Model Evaluation is that contractual no-logging promises may not be the only option for protecting a test set from post-evaluation leakage. Confidential computing could become useful where evaluators cannot receive frontier weights and model owners must not see sensitive cyber, government or safety prompts.
The pilot is not trust-free. The report says that not all proprietary model code was inspectable or allowlisted, individual Confidential Space builds were not independently reproducible, and Google services remained in the attestation path. It also identifies legal agreements, code review and human coordination as the main current bottleneck. Scaling from one H100 environment to confidential multi-node clusters remains future work.
Teams should therefore investigate the architecture, not revise a capability verdict. A next evidence gate would require independent reproduction, disclosed operational cost and published evaluation outcomes. Until then, the pilot is useful context for evals and benchmark-contamination controls, not proof that Gemini is safer or more capable.