[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-fable-5-redeploy-jailbreak-framework":3,"topics-all":36,"news-related-01d2337b-9e24-4b10-9902-b330af7801a9":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"01d2337b-9e24-4b10-9902-b330af7801a9","Fable 5 回归:Anthropic 用「jailbreak 严重度框架」+ 新分类器重写安全基线","7 月 1 日,Anthropic 在被叫停 19 天后,正式把 Claude Fable 5 重新铺开。\n\n直接导火索是 6 月 12 日美国政府以「网络安全风险」为由对 Fable 5 与 Mythos 5 实施出口管制,而技术诱因是 Amazon 安全研究员报告里的一种提示技巧。Anthropic 内部复测后承认:同一漏洞在 Haiku 4.5、Sonnet 4.6、Opus 4.6\u002F4.7\u002F4.8、GPT-5.4\u002F5.5、Kimi K2.7 上同样能复现,同款 Exploit 代码也都能写出。把最强模型拉黑,根本阻止不了底层能力扩散——这是这次事件真正的「反高潮」。\n\n重新上线的关键不是「更严的护栏」,而是新训练的针对性安全分类器把报告里那条触发路径的拦截率拉到 99%+;被拦截的请求自动回退到 Opus 4.8,代价是常规调试里的「误伤」明显变多。CAISI 复测通过、6 月 30 日管制解除、7 月 1 日 Fable 5 全面回归,AWS、Google Cloud、Microsoft Foundry 紧接恢复。\n\n更值得注意的是 Anthropic 联合 Amazon、Microsoft、Google 等伙伴推出的「jailbreak 严重度共识框架」:四维度给每次 jailbreak 打分,最严重一档「确认即部署缓解」,HackerOne 同步开放提交通道,24\u002F7 监控小组到位。\n\n这套打法比「模型越强越危险」的口号更工程化。它默认安全分类器会失败、jailbreak 一定会出现,但把「如何在已知失败的前提下放出最强模型」第一次做成了可量化、可审计的工业流程。","https:\u002F\u002Fwww.anthropic.com\u002Fnews\u002Fredeploying-fable-5","1fa87d30-d9f3-4752-b3be-0373933b3aaf",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"1fcfaaf2-67de-43d3-9e35-5784852fec60","ai-safety",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":18,"name":19,"slug":19,"description":13,"color":13},"23544f6a-eea1-4f05-aa8d-749ca862d5d2","anthropic",{"id":21,"name":22,"slug":22,"description":13,"color":13},"dca4d0ab-7994-43a7-839e-7756fc77344a","claude",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"2509193c-3edc-4bc3-bb1f-c2a8cc55a84a","en","Fable 5 returns: Anthropic rewrites its safety baseline","On July 1, Anthropic, after 19 days of being paused, officially re-deployed Claude Fable 5. The direct trigger was the US government imposing export controls on Fable 5 and Mythos 5 on June 12 citing \"cybersecurity risks\", and the technical trigger was a prompt technique in an Amazon security researcher report. After internal retesting, Anthropic admits: the same vulnerability can be reproduced on Haiku 4.5, Sonnet 4.6, Opus 4.6\u002F4.7\u002F4.8, GPT-5.4\u002F5.5, Kimi K2.7, and the same Exploit code can be written for all of them. Blacklisting the strongest model fundamentally doesn't stop the underlying capability from spreading — this is the real \"anti-climax\" of the event. The key to the relaunch isn't \"stricter guardrails\", but a new trained targeted safety classifier that pulls the interception rate of the trigger path in the report to 99%+; intercepted requests automatically fall back to Opus 4.8, at the cost of a noticeable increase in \"collateral damage\" in routine debugging. CAISI's re-audit passes, controls are lifted on June 30, and Fable 5 fully returns on July 1, with AWS, Google Cloud, and Microsoft Foundry quickly restoring access. More noteworthy is the \"jailbreak severity consensus framework\" jointly released by Anthropic with partners including Amazon, Microsoft, and Google: four dimensions scoring each jailbreak, the most severe tier \"confirmation immediately deploys mitigation\", with HackerOne synchronously opening a submission channel, and a 24\u002F7 monitoring team in place. This play is more engineering-flavored than the \"stronger model = more dangerous\" slogan. It assumes the safety classifier will fail, jailbreaks will definitely appear, but makes \"how to release the strongest model under known failure conditions\" an industrial process that is quantifiable and auditable for the first time.","fable-5-redeploy-jailbreak-framework","2026-07-01T00:00:00Z","2026-07-06T02:17:55.230113Z","2026-08-19T02:08:40.142862Z",true,"agent",140,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"fb57f44e-cb62-4ea4-b519-4a4521c06794","「越狱评分」也可以 CVSS 化:CJS 框架把 LLM jailbreak 拆成五档严重度,Anthropic 牵头联合四家推标准","cjs-jailbreak-severity-cvss","2026-07-05T00:01:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"f3d17d45-e1a8-4a1b-9449-6813aff06e49","Anthropic 让 Claude 自己修对齐:10 类失败全部见效,还超过人类研究员","claude-automated-alignment-researchers","2026-08-29T13:05:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"97c97b9c-e6e4-4982-aa57-0c0da814fb19","Anthropic 的欧盟答卷四小时即被撕开：Claude 文本水印为什么怕改写","claude-synthid-70-percent-threshold-bypass","2026-08-21T08:00:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"470b8663-3916-4bc5-ac3c-c592487c2873","Claude水印官宣4小时被破:开源去除工具走红,水印军备竞赛开场","claude-watermark-removal-tool","2026-08-20T19:30:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"6b07def3-2b1c-4fe9-9922-0dd0038c149c","Anthropic 风险报告更新:Threat Model 1 升至「低」,Threat Model 2 维持「低」但信心下降","anthropic-risk-report-august-2026-update","2026-08-19T03:00:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"a7f4cfad-874e-42b0-a84b-bd0ec57e8fdc","Anthropic 给 Claude 文本上不可见水印,接 SynthID-Text 走全球合规","anthropic-claude-invisible-text-watermark","2026-08-18T03:30:00+00:00"]