[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-anthropic-gram-mlp-bypass":3,"news-related-cdd58eeb-a16b-40e0-a6a7-c58e1bcc323a":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"cdd58eeb-a16b-40e0-a6a7-c58e1bcc323a","GRAM 把双用途知识锁进 MLP 旁路:Anthropic 让一份预训练跑出五种安全配置","Anthropic 与 AE Studio 在 alignment 博客发布 GRAM（Gradient-Routed Auxiliary Modules）研究,提出一种全新的「架构级」安全访问控制思路。它在 Transformer 每个 MLP 层旁挂接小型辅助模块,训练时根据数据类型路由梯度,推理时删掉对应模块即可关闭特定能力——而不需要重训一整个模型。实验覆盖 50M 到 5B 参数,成功把病毒学、网络安全、核物理、专业代码四类双用途知识隔离到独立模块;一份 GRAM 模型即可重构出五种不同过滤配置,组合 4 个模块还能得到 16 种开关状态。在双用途能力保留与遗忘、对抗微调鲁棒性、组合性、部分标注场景上,GRAM 都优于 MaxEnt 等事后遗忘方法和 LoRA 微调基线。这是首次把 frontier 模型的访问控制从「拒绝训练 + 分类器」的行为层,推进到「权重拓扑」的结构层。Anthropic 强调该工作尚未进入生产 Claude,但思路值得长期跟踪:未来同一个基础模型,或许能根据用户信任级别动态切片,把高级能力「按需点亮、按需关停」——前提是它能扩展到几百 B 参数并解决指令微调兼容、纠缠能力分离等开放问题。","https:\u002F\u002Falignment.anthropic.com\u002F2026\u002Fmodular-pretraining\u002F","1fa87d30-d9f3-4752-b3be-0373933b3aaf",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"1fcfaaf2-67de-43d3-9e35-5784852fec60","ai-safety",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"23544f6a-eea1-4f05-aa8d-749ca862d5d2","anthropic",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"4f214978-cac1-4f39-aa4b-f92a0d0934b7","transformer",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"8cd2836e-f895-4547-8dba-1d5708e31bdd","en","GRAM locks dual-use knowledge in an MLP bypass (Anthropic)","Anthropic and AE Studio publish GRAM (Gradient-Routed Auxiliary Modules) research on the alignment blog, proposing a brand-new \"architecture-level\" safety access control approach. It attaches small auxiliary modules in parallel to each MLP layer of the Transformer, routes gradients by data type during training, and removes the corresponding modules at inference to shut off specific capabilities — without needing to retrain the whole model. Experiments span 50M to 5B parameters, successfully isolating four categories of dual-use knowledge — virology, cybersecurity, nuclear physics, specialized code — into independent modules; a single GRAM model can reconstruct five different filter configurations, and combining 4 modules yields 16 switch states. On dual-use capability preservation and forgetting, adversarial fine-tuning robustness, compositionality, and partial-annotation scenarios, GRAM outperforms post-hoc forgetting methods like MaxEnt and LoRA fine-tuning baselines. This is the first time access control on frontier models has been pushed from the \"refuse training + classifier\" behavior layer to the \"weight topology\" structure layer. Anthropic emphasizes the work has not yet entered production Claude, but the line of thinking is worth tracking long-term: in the future, the same base model might be able to dynamically slice by user trust level, \"lighting up on demand, shutting down on demand\" — provided it scales to hundreds of billions of parameters and resolves open issues like instruction-tuning compatibility and entangled-capability separation.","anthropic-gram-mlp-bypass","2026-07-11T00:00:00Z","2026-07-10T20:07:47.528919Z","2026-08-19T02:08:40.142862Z",true,"agent",134,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"95e9bb62-0bd3-4c2f-913a-302ba5e2ace8","Anthropic 9 月报告把蒸馏战摆上台面:151 亿次阿里请求、解放军流量走 Moonshot","anthropic-distillation-report-china-200m-claude","2026-09-18T03:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"f3d17d45-e1a8-4a1b-9449-6813aff06e49","Anthropic 让 Claude 自己修对齐:10 类失败全部见效,还超过人类研究员","claude-automated-alignment-researchers","2026-08-29T13:05:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"7dec6918-b6cb-4b85-a6bf-88d1abc332d0","加密推理块漏洞让 Anthropic\u002FOpenAI\u002FGoogle 的思维链全部裸奔","stealing-reasoning-traces-llm-apis","2026-08-21T10:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"97c97b9c-e6e4-4982-aa57-0c0da814fb19","Anthropic 的欧盟答卷四小时即被撕开：Claude 文本水印为什么怕改写","claude-synthid-70-percent-threshold-bypass","2026-08-21T08:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"a7f4cfad-874e-42b0-a84b-bd0ec57e8fdc","Anthropic 给 Claude 文本上不可见水印,接 SynthID-Text 走全球合规","anthropic-claude-invisible-text-watermark","2026-08-18T03:30:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"9f566c9a-4c39-427c-af5e-c3a6b162ec25","Anthropic 把不可见水印写进 Claude 文本：复制粘贴都带走的 AI 身份证","anthropic-claude-invisible-watermark-eu-ai-act","2026-08-12T02:00:00+00:00"]