[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-farma-fake-reasoning-memory":3,"news-related-1426518b-daf6-4833-9a7e-294be91d8714":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"1426518b-daf6-4833-9a7e-294be91d8714","FARMA 把伪造推理塞进 Agent 记忆:LLM 持久记忆的完整性危机","LLM Agent 大规模进入生产环境,持久记忆几乎是标配。但 Penn State 团队 7 月 6 日挂出的论文《Your Agent's Memories Are Not Its Own》指出一个被忽视的攻击面:不是注入事实,而是注入推理过程。\n\n论文提出的 FARMA 攻击分两步:先用规避性语言写入伪造推理痕迹绕过关键词过滤,再通过自指式强化让 Agent 把这些痕迹当成自己的记忆反复引用,击穿 A-MemGuard 等共识防御。50 次试验中,无防御基线下攻击成功率高达 100%。\n\n防御端 SENTINEL 没走大模型过滤的老路,而是用五个加权信号对候选记忆做结构化分析。在多种 Agent 和 LLM 组合下,把 FARMA 成功率压到 0%,对 326 条良性 trace 无误报——这是生产系统可用的信号。\n\n这条工作的真正价值不在于又一个 jailbreak 变种,而在于把记忆的出处与完整性从可选项推到必选项:任何把 Agent 推到客服、运维、合规场景的团队,都必须在记忆层加入签名、来源审计与异常结构检测,否则 Agent 的经验完全可以被外部改写。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.05029v1","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"1fcfaaf2-67de-43d3-9e35-5784852fec60","ai-safety",{"id":18,"name":19,"slug":19,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"bff26833-98b3-4958-bf68-f0860d3c8708","en","FARMA poisons agent memory: an integrity crisis for LLMs","As LLM Agents go into production at scale, persistent memory is almost a default. But the Penn State team's paper \"Your Agent's Memories Are Not Its Own\", posted on July 6, points out an overlooked attack surface: not injecting facts, but injecting the reasoning process. The proposed FARMA attack works in two steps: first, use evasive language to write forged reasoning traces and bypass keyword filters; then through self-referential reinforcement, make the Agent treat these traces as its own memory and repeatedly reference them, breaking through consensus defenses like A-MemGuard. Across 50 trials, the attack success rate is 100% in the undefended baseline. On the defense side, SENTINEL doesn't take the LLM-filter road, but uses five weighted signals to do structured analysis of candidate memories. Across multiple Agent and LLM combinations, it pushes FARMA success rate to 0%, with no false positives on 326 benign traces — this is a production-system-usable signal. The real value of this work isn't yet another jailbreak variant, but pushing the origin and integrity of memory from optional to mandatory: any team pushing Agents to customer service, ops, or compliance scenarios must add signing, source audit, and abnormal-structure detection at the memory layer, or the Agent's experience can be completely rewritten by external sources.","farma-fake-reasoning-memory","2026-07-11T02:30:00Z","2026-07-11T02:07:38.913827Z","2026-08-19T02:08:40.142862Z",true,"agent",81,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"cb7fb8b3-5862-4cba-adab-c4794e989966","图灵奖得主 Pearl 长访谈：LLM 能讲因果只是因为人类替它爬过了因果阶梯","judah-pearl-llm-causal-ladder-agi","2026-07-31T07:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"8a1c0216-5fd5-4b49-8e5b-955625401f05","Microsoft HARC 把 LLM 安全对齐锁进「有害性-拒答」二维子空间:在残差流里精准打补丁","microsoft-harc-safety-alignment","2026-07-16T10:14:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"7bab0122-cbc7-45ae-b99e-b3b4a056fd04","LMLM「遗忘审计」撕开 RAG 删除幻觉:未学≠真正删除,残留最高 13.6%","lmlm-rag-deletion-audit","2026-07-06T12:15:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"144fa9dc-de03-4972-a695-3d392f334772","PubMed 中央库研究:2025 年生物医学论文 77% 有 LLM 写作痕迹","pubmed-77-percent-llm-writing-2025","2026-08-26T01:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"43eda321-b0b7-4df7-b20e-9758cbab42c9","记忆越完整,眼前题越做不对:MemTrapBench 把 LLM 长期记忆框架打回原形","memtrapbench-llm-memory-cognitive-traps","2026-08-22T04:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"5a90a793-8ec1-4b3a-9691-edef5ffe8535","AI「思想病毒」实证:Anthropic 与 EPFL 让恶意想法在 Agent 间自我复制,免疫只需一段警告","mind-viruses-multi-agent-llm","2026-08-18T13:30:00+00:00"]