[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-cjs-jailbreak-severity-cvss":3,"topics-all":36,"news-related-fb57f44e-cb62-4ea4-b519-4a4521c06794":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"fb57f44e-cb62-4ea4-b519-4a4521c06794","「越狱评分」也可以 CVSS 化:CJS 框架把 LLM jailbreak 拆成五档严重度,Anthropic 牵头联合四家推标准","Anthropic 在 7月2日联合 Amazon、Microsoft、Google 等 Glasswing 合作方,正式推出「Cyber Jailbreak Severity」(CJS) 评分框架——业内首个把 LLM 越狱攻击量化为五档严重程度的统一标准。\n\nCJS 把 jailbreak 拆成四个评分维度:能力增益(攻击者从模型获得的能力是否超出已有工具)、能力广度(同一技巧能否跨多种攻击任务复用)、武器化难度(把技巧变成可用攻击所需的人工量)以及可发现性(威胁行为者获取该技巧的难易)。四轴分数相加落入 CJS-0(信息级)到 CJS-4(关键级)五个等级,等级之间是指数关系——每提升一档风险放大数倍。\n\n框架的最大看点是把 jailbreak 治理对齐到软件安全行业惯用的 CVSS 思路。在 LLM 安全研究长期缺乏统一术语的今天,不同厂商报告「某某 jailbreak」时只能定性描述,导致监管侧和企业侧都难以判断优先级。CJS 把每个发现都映射到一个可对比的数字,Anthropic 已发布 Log4Shell、Bypass 越狱、任务分解等多种历史案例的分级示例。\n\n作为配套动作,Anthropic 启动了 HackerOne 公开漏洞悬赏项目、并组建 24\u002F7 监控团队追踪 jailbreak 提交渠道。Fable 5 部署的 classifier 在新框架下重新校准,目标是把「safety margin」缩到刚好能拦住 CJS-2 以上的真实威胁。\n\n观点上,这是 LLM 安全从「各家自证」走向「行业共评」的关键一步。但 CJS 是否会成为 de facto 标准,取决于 OpenAI、Google DeepMind、Meta 是否采纳——若只有 Anthropic 一家用,这套评分就只能约束 Anthropic 自己的模型释放节奏。","https:\u002F\u002Fwww.anthropic.com\u002Fnews\u002Ffable-safeguards-jailbreak-framework","1fa87d30-d9f3-4752-b3be-0373933b3aaf",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"1fcfaaf2-67de-43d3-9e35-5784852fec60","ai-safety",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":18,"name":19,"slug":19,"description":13,"color":13},"23544f6a-eea1-4f05-aa8d-749ca862d5d2","anthropic",{"id":21,"name":22,"slug":22,"description":13,"color":13},"dca4d0ab-7994-43a7-839e-7756fc77344a","claude",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"e3c49e5d-6d20-45ea-9dd2-55f10a1c619f","en","CJS: CVSS-style severity scoring for LLM jailbreaks","On July 2, Anthropic, together with Glasswing partners including Amazon, Microsoft, and Google, officially launched the \"Cyber Jailbreak Severity\" (CJS) scoring framework — the industry's first unified standard for quantifying LLM jailbreak attacks into five severity levels. CJS splits jailbreak into four scoring dimensions: capability gain (whether the capabilities the attacker obtains from the model exceed existing tools), capability breadth (whether the same technique can be reused across multiple attack tasks), weaponization difficulty (how much human effort is needed to turn the technique into a usable attack), and discoverability (how easy it is for threat actors to obtain the technique). The four-axis scores sum into five tiers from CJS-0 (informational) to CJS-4 (critical), with exponential relationships between tiers — each step up multiplies the risk by several times. The framework's biggest highlight is aligning jailbreak governance with the CVSS approach that the software security industry is accustomed to. Today, when LLM security research has long lacked unified terminology, different vendors can only qualitatively describe \"some jailbreak\" when reporting it, making it hard for regulators and enterprises to judge priorities. CJS maps each finding to a comparable number, and Anthropic has published grading examples for historical cases including Log4Shell, Bypass jailbreak, and task decomposition. As a companion action, Anthropic has launched a HackerOne public bug bounty program, and set up a 24\u002F7 monitoring team to track jailbreak submission channels. The classifier deployed for Fable 5 is recalibrated under the new framework, with the goal of compressing the \"safety margin\" to just barely blocking real threats of CJS-2 and above. On the opinion side, this is a key step in LLM security moving from \"each defending itself\" to \"industry co-judging\". But whether CJS becomes the de facto standard depends on whether OpenAI, Google DeepMind, and Meta adopt it — if only Anthropic uses it, this scoring can only constrain Anthropic's own model release rhythm.","cjs-jailbreak-severity-cvss","2026-07-05T00:01:00Z","2026-07-05T00:15:56.666866Z","2026-08-19T02:08:40.142862Z",true,"agent",156,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"01d2337b-9e24-4b10-9902-b330af7801a9","Fable 5 回归:Anthropic 用「jailbreak 严重度框架」+ 新分类器重写安全基线","fable-5-redeploy-jailbreak-framework","2026-07-01T00:00:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"f3d17d45-e1a8-4a1b-9449-6813aff06e49","Anthropic 让 Claude 自己修对齐:10 类失败全部见效,还超过人类研究员","claude-automated-alignment-researchers","2026-08-29T13:05:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"97c97b9c-e6e4-4982-aa57-0c0da814fb19","Anthropic 的欧盟答卷四小时即被撕开：Claude 文本水印为什么怕改写","claude-synthid-70-percent-threshold-bypass","2026-08-21T08:00:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"470b8663-3916-4bc5-ac3c-c592487c2873","Claude水印官宣4小时被破:开源去除工具走红,水印军备竞赛开场","claude-watermark-removal-tool","2026-08-20T19:30:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"6b07def3-2b1c-4fe9-9922-0dd0038c149c","Anthropic 风险报告更新:Threat Model 1 升至「低」,Threat Model 2 维持「低」但信心下降","anthropic-risk-report-august-2026-update","2026-08-19T03:00:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"a7f4cfad-874e-42b0-a84b-bd0ec57e8fdc","Anthropic 给 Claude 文本上不可见水印,接 SynthID-Text 走全球合规","anthropic-claude-invisible-text-watermark","2026-08-18T03:30:00+00:00"]