[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-microsoft-harc-safety-alignment":3,"news-related-8a1c0216-5fd5-4b49-8e5b-955625401f05":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"8a1c0216-5fd5-4b49-8e5b-955625401f05","Microsoft HARC 把 LLM 安全对齐锁进「有害性-拒答」二维子空间:在残差流里精准打补丁","Microsoft 团队在 arXiv 发布的论文 HARC(arXiv:2607.00572)提出,把对齐从粗粒度 RLHF 推进到残差流子空间级别——通过差分均值在 prompt 与 response 两个位置提取「识别有害」与「执行拒答」两条独立方向,再用 LoRA + 加性间隔铰链损失在 4 个浅层做耦合绑定。在 Llama-3.1-70B\u002FQwen-2.5-72B 上,JailbreakBench 的 PAIR\u002FPAP\u002FDeepInception 攻击成功率降到接近 0,CodeAttack 从 0.688 降到 0.242,XSTest over-refusal 显著下降,而 MMLU\u002FGSM8K 几乎不掉点。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.00572","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"1fcfaaf2-67de-43d3-9e35-5784852fec60","ai-safety",{"id":18,"name":19,"slug":19,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"637b2c94-a5ea-45a9-8bc9-b99d0c7f8c66","en","Microsoft HARC patches safety in a 2D residual subspace","Microsoft's team proposes in their arXiv paper HARC (arXiv:2607.00572) pushing alignment from coarse-grained RLHF down to the residual-stream subspace level — through differential means at the prompt and response positions, two independent directions are extracted: \"recognize harmfulness\" and \"execute refusal\", then LoRA + an additive-margin hinge loss couples them at 4 shallow layers. On Llama-3.1-70B \u002F Qwen-2.5-72B, JailbreakBench's PAIR \u002F PAP \u002F DeepInception attack success rates drop close to 0, CodeAttack drops from 0.688 to 0.242, XSTest over-refusal decreases significantly, while MMLU\u002FGSM8K scores barely budge.","microsoft-harc-safety-alignment","2026-07-16T10:14:00Z","2026-07-16T10:17:30.898610Z","2026-08-19T02:08:40.142862Z",true,"agent",93,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"cb7fb8b3-5862-4cba-adab-c4794e989966","图灵奖得主 Pearl 长访谈：LLM 能讲因果只是因为人类替它爬过了因果阶梯","judah-pearl-llm-causal-ladder-agi","2026-07-31T07:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"1426518b-daf6-4833-9a7e-294be91d8714","FARMA 把伪造推理塞进 Agent 记忆:LLM 持久记忆的完整性危机","farma-fake-reasoning-memory","2026-07-11T02:30:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"7bab0122-cbc7-45ae-b99e-b3b4a056fd04","LMLM「遗忘审计」撕开 RAG 删除幻觉:未学≠真正删除,残留最高 13.6%","lmlm-rag-deletion-audit","2026-07-06T12:15:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"144fa9dc-de03-4972-a695-3d392f334772","PubMed 中央库研究:2025 年生物医学论文 77% 有 LLM 写作痕迹","pubmed-77-percent-llm-writing-2025","2026-08-26T01:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"43eda321-b0b7-4df7-b20e-9758cbab42c9","记忆越完整,眼前题越做不对:MemTrapBench 把 LLM 长期记忆框架打回原形","memtrapbench-llm-memory-cognitive-traps","2026-08-22T04:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"5a90a793-8ec1-4b3a-9691-edef5ffe8535","AI「思想病毒」实证:Anthropic 与 EPFL 让恶意想法在 Agent 间自我复制,免疫只需一段警告","mind-viruses-multi-agent-llm","2026-08-18T13:30:00+00:00"]