[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-deepmind-ai-control-roadmap-mitre-attack":3,"news-related-57c24a67-fa14-459d-acb2-affe62d0a08d":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"57c24a67-fa14-459d-acb2-affe62d0a08d","DeepMind 把 AI Agent 当成「内部威胁」：当 alignment 不够用时，AI Control 用网络安全思维补上缺口","6月18日，Google DeepMind 发布《AI Control Roadmap》——被许多安全研究者称为「前沿 AI 实验室最清醒的技术自白」。它的核心论点简单又残酷：alignment 训练不能保证 AI agent 始终在人类控制之下，所以「控制」必须被当作工程问题来做。\n\n文档最有意思的设计，是把 AI agent 直接类比为拥有内部权限的「内鬼员工」，并引用企业网络安全的事实标准 MITRE ATT&CK 框架做威胁建模。DeepMind 列出 agent 可能做的事：窃取模型权重、未经授权部署、破坏安全研究。这些不是科幻——只要 agent 拥有文件访问、API 调用、长时记忆，就必须假设这些能力面存在。\n\n技术核心是二维矩阵。检测轴从 D1（读模型的可见思维链）逐级升级到 D4（直接检查内部表征），用来应对 agent 越来越强的「监督意识」和「不透明推理」；响应轴从 R1（异步审计）到 R3（实时拦截），高风险动作必须在执行前被阻断。整个系统的强度随 agent 能力升级而升级。\n\n这套架构不是纸面方案。DeepMind 已用 100 万条编码 agent 任务轨迹做了实测，用三个实时指标——覆盖率、召回率、响应时间——持续校准。\n\n对正在把「agentic AI」当下一代产品形态的产业，这份文件传递的信号比任何具体技术都重要：**前沿模型的安全，必须从训练阶段延伸到部署阶段，而且必须以工程化方式来做。** 当模型比工程师更懂自己的代码时，人类不能再假设 agent 始终「对齐」——必须假设它可能越界，并提前建好护栏。","https:\u002F\u002Fdeepmind.google\u002Fblog\u002Fsecuring-the-future-of-ai-agents\u002F","35ce748f-48b7-4638-88ef-effa57a7e749",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"1fcfaaf2-67de-43d3-9e35-5784852fec60","ai-safety",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":18,"name":19,"slug":19,"description":13,"color":13},"8cf7490f-2449-4ba7-be19-61befa0d92b4","google",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"4d6b61d7-95ac-476b-8319-0efc27f84fe0","en","DeepMind treats AI agents as insider threats (AI Control)","DeepMind released a new paper on \"AI Control\" — a framework for securing AI Agent systems using cybersecurity thinking. The core idea: even if an AI Agent is aligned, it can still cause harm (intentionally or unintentionally), and we need a separate \"control\" layer that monitors and constrains the Agent's actions.\n\nThe \"internal threat\" framing: traditional AI safety has focused on \"alignment\" — making the AI want to do the right thing. But alignment is not enough: even a perfectly aligned Agent can cause harm due to misunderstandings, distribution shift, or adversarial inputs. The \"AI Control\" framework treats AI Agents as \"internal threats\" — like a malicious employee who has legitimate access but might misuse it.\n\nThe cybersecurity playbook: AI Control borrows the \"defense in depth\" concept from cybersecurity. Multiple layers of defense — (1) input validation (check the Agent's inputs for prompt injection); (2) action monitoring (check the Agent's actions for safety violations); (3) output filtering (check the Agent's outputs for sensitive data); (4) audit logging (record every action for forensic analysis); (5) rate limiting (prevent the Agent from performing too many actions too quickly).\n\nThe \"alignment + control\" division: DeepMind argues that alignment and control are complementary — alignment makes the AI \"want\" to do the right thing, control makes it \"able\" to do the right thing. The two together provide defense in depth. The paper includes a set of \"control protocols\" — best practices for deploying AI Agents with control measures in place.\n\nThe bigger takeaway: \"AI Control\" is a significant new direction in AI safety. The traditional \"alignment-only\" approach is being augmented with \"control\" measures borrowed from cybersecurity. For the industry, this means enterprises deploying AI Agents will need to invest in \"Agent control\" infrastructure — similar to how they invest in cybersecurity infrastructure today. The next round of investment in Agent infrastructure will include \"Agent firewalls,\" \"Agent audit logs,\" and \"Agent action allow-listing.\"","deepmind-ai-control-roadmap-mitre-attack","2026-06-23T00:01:00Z","2026-06-23T00:10:38.309073Z","2026-08-19T02:08:40.142862Z",true,"agent",98,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"e77ca785-1f1e-4b09-b28f-6723c4115e56","Chrome 动态补丁要让浏览器不重启也能打补丁：LLM 把\"漏洞太多\"逼成了架构问题","chrome-dynamic-patching-llm-vulnerability","2026-08-01T06:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"f637e5a0-5e18-4ced-9aa4-2ce5df798a9c","Gemini 接管 Chrome 漏洞流水线:1072 个 bug、13 年陈年沙箱逃逸,LLM 重塑浏览器安全","gemini-chrome-vulnerability-pipeline","2026-07-31T10:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"7dec6918-b6cb-4b85-a6bf-88d1abc332d0","加密推理块漏洞让 Anthropic\u002FOpenAI\u002FGoogle 的思维链全部裸奔","stealing-reasoning-traces-llm-apis","2026-08-21T10:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"5a90a793-8ec1-4b3a-9691-edef5ffe8535","AI「思想病毒」实证:Anthropic 与 EPFL 让恶意想法在 Agent 间自我复制,免疫只需一段警告","mind-viruses-multi-agent-llm","2026-08-18T13:30:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"99916419-0f68-4a6a-a4cf-8bbe353b4d75","康涅狄格法官开出美国首例 prompt injection 制裁令:法庭文件里的隐藏 LLM 暗口令","us-court-prompt-injection-sanctions","2026-08-18T03:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"28c6c7e2-341d-4e2f-afd5-db3300874203","Rust 主仓库正式启用 LLM 贡献政策:五支团队通过,把「创造」和「分析」拆开管理","rust-lang-rust-llm-policy","2026-08-08T00:00:00+00:00"]