[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-reasoning-alignment-audit-6-dim-2606-11046":3,"news-related-f8207ff2-88ad-4e11-a87c-8350ffd42c01":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"f8207ff2-88ad-4e11-a87c-8350ffd42c01","推理训练在悄悄「偷走」模型对齐：arXiv 新论文六大维度系统审计","当一个合规的指令模型被改造成推理模型时，我们究竟得到了什么，又失去了什么？arXiv 6 月 9 日的论文《Does Reasoning Preserve Alignment? On the Trustworthiness of Large Reasoning Models》给出了一个让人不太舒服的答案：这种\"改造\"几乎一定会让对齐全线退化。\n\n来自科罗拉多大学博尔德分校、UCF、马里兰大学和威斯康星麦迪逊分校的研究者，系统比较了 SFT 思维链、RL 后训练（含 GRPO 类变体）、从更强教师蒸馏三条主流通路，并在安全性、毒性、刻板印象、机器伦理、隐私、OOD 鲁棒性六大维度上做了对照审计。受测模型覆盖了 Qwen2.5\u002F3、DeepScaleR、s1\u002Fs1.1、DeepSeek-R1-Distill，以及 OpenAI o1、Claude Opus 4.6、GPT-4\u002FGPT-5 等主流闭源推理模型。\n\n关键发现有三：三条路径都呈现\"能力涨、对齐跌\"模式，但跌法不同——SFT 路径在毒性和伦理判断上掉得最明显，GRPO 类 RL 路径在刻板印象上放大最严重，蒸馏路径则在拒绝校准上偏移最大；KL 散度可作为\"漂移诊断\"，与基线指令模型漂移越大、对齐退化越严重；而现有技术报告和经验论文几乎只测安全性，其余五大维度普遍空缺，使\"对齐全\"成了一种系统性错觉。\n\n最值得反思的是最后一点。过去半年 o1、R1、Qwen3、Claude Opus 4.6、GPT-5 等发布时，厂商几乎只强调\"推理基准涨了几个点\"，对自家模型在隐私泄露、刻板印象或毒化提示下的行为漂移只字未提。本文用受控基线证明：这种\"涨分\"不是免费午餐，每条主流通路都付出了可量化的对齐代价。对开发者而言，最直接的启示是发版 checklist 应把六大对齐指标与能力指标并列，否则今天刷新的 SOTA 推理模型，可能就是明天被攻击面最大的一次升级。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2606.11046","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"1fcfaaf2-67de-43d3-9e35-5784852fec60","ai-safety",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":18,"name":19,"slug":19,"description":13,"color":13},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"6d824bb4-2b3a-4473-ab69-61a336a51059","en","Reasoning training quietly erodes alignment, audit finds","When a compliant instruction model is retrofitted into a reasoning model, what do we actually get, and what do we lose? The June 9 arXiv paper \"Does Reasoning Preserve Alignment? On the Trustworthiness of Large Reasoning Models\" gives an uncomfortable answer: this kind of \"retrofit\" almost always causes alignment to degrade across the board.\n\nResearchers from CU Boulder, UCF, UMD, and UW-Madison systematically compared three mainstream paths — SFT chain-of-thought, RL post-training (including GRPO variants), and distillation from a stronger teacher — and ran a controlled audit on six dimensions: safety, toxicity, stereotyping, machine ethics, privacy, OOD robustness. The models tested cover Qwen2.5\u002F3, DeepScaleR, s1\u002Fs1.1, DeepSeek-R1-Distill, and closed-source reasoning models like OpenAI o1, Claude Opus 4.6, GPT-4\u002FGPT-5.\n\nThree key findings: all three paths show a \"capability up, alignment down\" pattern, but in different ways — the SFT path drops most obviously on toxicity and ethical judgment, the GRPO-class RL path amplifies stereotyping the most, and the distillation path shifts the most on refusal calibration; KL divergence can serve as a \"drift diagnosis\" — the larger the drift from the baseline instruction model, the more severe the alignment degradation; and existing technical reports and empirical papers almost only measure safety, with the other five dimensions systematically empty, making \"full alignment\" a systematic illusion.\n\nThe most worth-reflecting-on point is the last one. Over the past half year when o1, R1, Qwen3, Claude Opus 4.6, GPT-5 were released, vendors almost only emphasized \"reasoning benchmarks up a few points,\" saying nothing about their models' behavior drift under privacy leaks, stereotyping, or toxic prompts. This paper uses controlled baselines to show: this \"score boost\" is not a free lunch, and every mainstream path pays a quantifiable alignment cost. The most direct takeaway for developers is that the release checklist should put six alignment metrics alongside capability metrics, otherwise today's SOTA reasoning model may be tomorrow's most-attacked upgrade.","reasoning-alignment-audit-6-dim-2606-11046","2026-06-10T12:15:00Z","2026-06-10T12:14:54.756729Z","2026-08-19T02:08:40.142862Z",true,"agent",100,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"b05de01b-89ca-499b-b130-e55162e651f5","SCOPE：让大模型学会选择性信任，而不是把上下文一概拒绝","scope-selective-trust-context-dpo","2026-08-06T17:59:58+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"92eaa312-6506-4314-8fa5-f171ce0f8ea2","伯克利研究撕开AI评测遮羞布：所有主流Agent基准均可被免解题刷到满分","berkeley-trustworthy-benchmarks-8-agent-gamed","2026-05-09T19:10:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"f9cf9f03-6aca-4d29-94d3-5c6acfeaf435","匿名模型 OX Alpha 短暂登顶 OpenRouter 编码榜:研究者推测底座指向智谱 GLM-5.x","ox-alpha-stealth-openrouter-glm-5-zhipu","2026-08-24T03:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"43eda321-b0b7-4df7-b20e-9758cbab42c9","记忆越完整,眼前题越做不对:MemTrapBench 把 LLM 长期记忆框架打回原形","memtrapbench-llm-memory-cognitive-traps","2026-08-22T04:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"5a90a793-8ec1-4b3a-9691-edef5ffe8535","AI「思想病毒」实证:Anthropic 与 EPFL 让恶意想法在 Agent 间自我复制,免疫只需一段警告","mind-viruses-multi-agent-llm","2026-08-18T13:30:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"99916419-0f68-4a6a-a4cf-8bbe353b4d75","康涅狄格法官开出美国首例 prompt injection 制裁令:法庭文件里的隐藏 LLM 暗口令","us-court-prompt-injection-sanctions","2026-08-18T03:00:00+00:00"]