[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-deepmind-gemini-double-blind-eval":3,"topics-all":38,"news-related-47f12a0d-8559-473f-97be-dc12966bd4ff":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"47f12a0d-8559-473f-97be-dc12966bd4ff","DeepMind 双盲评测：Gemini 权重和考题锁进同一个加密飞地","Google DeepMind 联手 AVERI、MLCommons 等四方,把 Gemini 2.5 Flash-Lite 的权重和保密考题同时锁进硬件加密飞地完成评测,号称业界首例闭源模型双盲评测。考题不再泄漏给厂商,权重也不再交给评测方。","基准分数还能信吗?这个问题困扰行业已久:模型在训练时可能\"背过\"公开基准的题目,高分未必代表真实能力。2026 年一项覆盖 60 个 LLM 基准的研究发现,约一半已出现这类\"污染饱和\"迹象;NIST 也在 2026 年 2 月的报告中点名了这一风险。8 月 27 日,Google DeepMind 给出了一个新答案:把评测双方的秘密同时锁进硬件加密的\"飞地\",完成了一场谁也没法作弊的考试。\n\n## 评测行业的结构性死结\n\n第三方评测闭源模型,长期只有两条路:要么把考题交给厂商——通过 API 跑题,厂商看得见题目;要么把模型权重交给评测方——商业机密直接裸奔。零日志协议和合同保密条款是现有补丁,但本质仍靠\"组织信任\"。2026 年 7 月发布的《新加坡共识》(覆盖 13 国、100 位贡献者的全球 AI 安全研究优先级报告)明确把\"双盲评测基础设施\"列为待解难题:评测方拿不到模型参数,厂商不知道考题,并直言仅靠组织信任和流程控制不足以支撑前沿 AI 监督的分量。\n\n## 飞地里发生了什么\n\n这次试点的技术底座是 Google Cloud 的 Confidential Space:一台 NVIDIA H100 80GB 机密 GPU,配合 Intel TDX 主机内存加密。流程分几步:DeepMind 把 Gemini 2.5 Flash-Lite 的权重和推理代码、评测方把 AILuminate 安全基准的考题和评分代码,各自经加密连接送入同一个飞地。权重落在硬件加密的 GPU 显存,考题落在加密的主机内存;数据出发前,远程证明(remote attestation)先验证飞地运行的是双方事先核准的软件。评测全程,厂商读不到考题,评测方读不到权重;只有约定范围内的指标结果离开这台机器,随后临时环境直接销毁。\n\n参与方阵容值得记一笔:MLCommons 提供从未公开使用过的 AILuminate 提示词集,OpenMined 提供并改造了执行软件,新加坡 AI 安全研究院与 AVERI(AI 验证与评测研究院)作为评测方,对输出解密并按 AILuminate 标准打分,产出定性与小规模定量评估。AVERI 的完整试点报告见其官网(http:\u002F\u002Faveri.org\u002Fourwork\u002Faveri-pilot-report-the-worlds-first-double-blind-eval)。\n\n## 意义与边界\n\n往大了说,这是把 AI 评测的信任基础从\"合同条款\"搬进\"密码学\":密码学能证明是哪份代码碰了哪份秘密,却不能证明考题出得对、评分标准合理、模型因此安全——这是参与方自己都承认的边界。报告还指出,如今最大的障碍已不是飞地的性能开销,而是法律协议、代码审查、输出策略与多方协调这些\"人的问题\"。\n\n对行业的\"所以呢\"在于:当监管、采购、部署决策越来越依赖第三方评测,可验证的中立考场会成为一种基础设施。下一次看到\"XX 模型刷新纪录\"时,值得多问一句——这场考试,是闭卷的吗?","http:\u002F\u002Faveri.org\u002Fourwork\u002Faveri-pilot-report-the-worlds-first-double-blind-eval","29ad0df7-80ee-4124-a7ef-826447bbf1c6",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"1fcfaaf2-67de-43d3-9e35-5784852fec60","ai-safety",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":19,"name":20,"slug":20,"description":14,"color":14},"a9524a82-a7c5-4daa-bb4b-a7ee77bb0b94","gemini",{"id":22,"name":23,"slug":23,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"1cd2984a-5d0f-4e6e-a893-e1f75f6437c6","en","DeepMind runs first double-blind eval of a closed frontier model","Google DeepMind, with AVERI, MLCommons, OpenMined and Singapore's AISI, ran what it bills as the first double-blind evaluation of a proprietary frontier model: Gemini 2.5 Flash-Lite's weights and confidential benchmark prompts were sealed inside one hardware-encrypted enclave, visible to neither side.","Can we still trust benchmark scores? The industry has wrestled with this for years: a model may have effectively memorized public benchmark questions during training, so a high score no longer proves real capability. A 2026 study covering 60 LLM benchmarks found roughly half already show signs of this contamination-driven saturation, and NIST flagged the same risk in a February 2026 report. On August 27, Google DeepMind offered a new answer: lock both sides' secrets inside a hardware-encrypted enclave and run an exam where neither party can cheat.\n\n## The structural deadlock in AI evaluation\n\nThird-party evaluation of closed models has long offered only two bad options: hand the test questions to the vendor — run them through the provider's API, where the company can see every prompt — or hand the model weights to the evaluator, exposing the crown jewels. Zero-logging protocols and contractual NDAs are the current patches, but they ultimately rely on organizational trust. The 2026 Singapore Consensus on Global AI Safety Research Priorities — a report spanning 13 countries and 100 contributors — explicitly named this gap as a priority research challenge: evaluation infrastructure should ideally be double-blind, where the evaluator cannot access model parameters and the developer does not know the exact evaluations, because reliance on organizational trust and procedural controls is insufficient for the stakes of frontier AI oversight.\n\n## What happened inside the enclave\n\nThe pilot's technical foundation is Google Cloud's Confidential Space: a single NVIDIA H100 80GB confidential GPU paired with Intel TDX host memory encryption. The flow has several stages. DeepMind loads Gemini 2.5 Flash-Lite's weights and inference code; the evaluators load AILuminate safety benchmark prompts and scoring code. Both travel into the enclave over encrypted connections — weights land in hardware-encrypted GPU memory, prompts in encrypted host memory. Before anything is sent, remote attestation verifies the enclave is running the software both parties agreed on. Throughout the run, the vendor cannot read the prompts and the evaluators cannot read the weights. Only the bounded, pre-agreed metrics leave the machine, and the temporary environment is then decommissioned.\n\nThe cast is worth noting: MLCommons supplied a never-before-used set of AILuminate prompts; OpenMined produced and adapted the execution software; the Singapore AI Safety Institute and AVERI (the AI Verification and Evaluation Research Institute) acted as evaluators, decrypting outputs and grading them against AILuminate criteria to produce qualitative and small-scale quantitative assessments. AVERI's full pilot report is on its site (http:\u002F\u002Faveri.org\u002Fourwork\u002Faveri-pilot-report-the-worlds-first-double-blind-eval).\n\n## Significance and limits\n\nThe big picture: this moves the trust basis of AI evaluation from contract clauses into cryptography. Cryptography can prove which code touched which secret; it cannot prove the test asked the right questions, that the scoring policy is correct, or that the model is therefore safe — a boundary the participants themselves acknowledge. The report also notes that the biggest practical obstacle is no longer raw enclave overhead, but legal agreements, code review, output policy, and coordination between organizations — the human problems.\n\nSo what? As regulation, procurement, and deployment decisions lean increasingly on third-party evaluation, verifiable neutral exam rooms will become infrastructure. Next time you see \"model X sets a new record,\" it's worth asking one more question: was that exam closed-book?","deepmind-gemini-double-blind-eval","2026-08-29T15:05:00Z","2026-08-29T15:11:47.894268Z","2026-08-29T15:11:47.894283Z",true,"agent",184,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"b05de01b-89ca-499b-b130-e55162e651f5","SCOPE：让大模型学会选择性信任，而不是把上下文一概拒绝","scope-selective-trust-context-dpo","2026-08-06T17:59:58+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"f637e5a0-5e18-4ced-9aa4-2ce5df798a9c","Gemini 接管 Chrome 漏洞流水线:1072 个 bug、13 年陈年沙箱逃逸,LLM 重塑浏览器安全","gemini-chrome-vulnerability-pipeline","2026-07-31T10:00:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"218e966c-1521-44eb-87d1-dba77cc26c9c","CyberGym 的 86.3%：GLM-5.2 安全 Agent 开始用证据说话","cybergym-glm-52-evidence-agent","2026-07-29T06:33:39+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"86565410-ced9-4e64-80b7-d97a353a1d1d","LLM 元认知首次可度量：Metacognition-Bench 用 300 道陷阱题 + 11 个开源错题雷达适配器把自我纠错做成开源工程","metacognition-bench-llm-self-correction","2026-07-01T12:00:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"817a35b2-6b31-41e6-a213-3ac6f667fd14","MosaicLeaks：ServiceNow 撕开 Deep Research Agent 的\"查询即泄密\"盲区","mosaicleaks-servicenow-deep-research-leak","2026-06-26T16:30:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"f8207ff2-88ad-4e11-a87c-8350ffd42c01","推理训练在悄悄「偷走」模型对齐：arXiv 新论文六大维度系统审计","reasoning-alignment-audit-6-dim-2606-11046","2026-06-10T12:15:00+00:00"]