[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-openai-private-safety-processing-zdr-astra":3,"news-related-73c511d8-577d-4671-90c5-71653a83d9ce":35},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":21,"news_slug":28,"published_at":29,"created_at":30,"modified_at":31,"is_published":32,"publish_type":33,"image_url":14,"view_count":34},"73c511d8-577d-4671-90c5-71653a83d9ce","OpenAI Private Safety Processing 兼顾前沿模型零数据留存","OpenAI 宣布面向 ZDR 客户推出 Private Safety Processing,在不接触用户内容的前提下,跨多轮交互识别滥用模式;同时披露 Astra 模型触及 Preparedness Framework 的网络安全关键阈值,前沿训练仍部分搁置。","## 让 ZDR 与多轮安全同时成立\n\n8 月 19 日,OpenAI 在官方博客发布了 Private Safety Processing 的预览版,试图在前沿模型的安全和「零数据留存(ZDR)」承诺之间找一个新的平衡点:当用户选择 ZDR 部署时,客户内容始终留在客户自己的基础设施上,或者用客户自己掌握的密钥在 OpenAI 基础设施中加密;自动化系统可以识别跨多轮交互的潜在滥用模式,并把一个范围被刻意收窄的「事件类型 + 严重程度」信号回传给 OpenAI,而 OpenAI 内部员工始终拿不到原始的客户提示或回复([openai.com][1])。\n\nGlean 的 CISO Sunil Agrawal 在公告里给出了一句话注脚:「OpenAI 的不训练承诺和 ZDR,让我们有信心基于 OpenAI 搭建产品。模型能力提升的同时,OpenAI 证明安全可以往前走,不需要牺牲维持企业信任的隐私和控制权。」这是 OpenAI 目前愿意公开背书 ZDR 路径的企业名单上最为显眼的一位。\n\n## 为什么不是「零数据」「零安全」二选一\n\nOpenAI 这套设计的动机,在公告里讲得很直白:许多严重的安全风险,**只有把多轮交互放在一起看才会显形**(同一个人反复试探护栏、跨账号协同、把威胁伪装成日常研究),而单轮评估根本看不出。问题是,过去要让安全系统看到这种「跨轮」模式,只能让 OpenAI 把客户内容留下来做人工+自动审查——一旦客户坚持 ZDR,这条路径就被卡住了。\n\nPrivate Safety Processing 的取舍是:OpenAI 拿到的是一个**被裁剪过的信号**(「疑似 CSAM」「疑似违规使用模式」之类),不是内容本身。即便是被自动系统标记的请求,客户内容也不会被回传给 OpenAI 人员查看;客户收到完整告警,可以自行调查、提起申诉,或者在自愿的基础上把相关信息补回给 OpenAI 用作调查。9 月起 Private Safety Processing 将开始正式 rollout,OpenAI 也将同期发布一份技术白皮书。\n\n## Astra 模型把「网络安全关键」撞到红线\n\n公告同期披露的另一件事,直接解释了为什么 OpenAI 现在要花精力重新设计「安全与 ZDR 共存」的工程管线。8 月 7 日,OpenAI 内部确认,尚未发布的 Astra 模型在 OpenAI 自家的 Preparedness Framework 下达到了「关键(Critical)」网络安全能力阈值。这个内部政策明文规定:一旦模型越过这条线,该公司就要**暂停相关模型的进一步训练**,直到安全与对齐工作满足更高门槛为止。\n\n8 月 18 日的 OpenAI 公告和当天的 Fortune 报道,共同讲清楚了接下来的安排:OpenAI 已经把两个层级分开了——一部分规模较小的训练和评估在新增监控、沙箱、网络隔离、链式思维(coT)多阶段监督落地之后已经恢复;但 Astra 训练 \u002F RL 评估中**真正大量面向前沿的研究负载仍处暂停**,优先把「安全 \u002F 对齐」相关的工作迁到加固后的新环境里([Fortune][2])。\n\nOpenAI 给出的工程副作用被量化成一句:这套「强化版监控+沙箱+CoT 监督」会带来大约 20% 的额外推理算力开销。OpenAI 总科学家 Jakub Pachocki 在记者会上做了补充,核心论点其实是组织层面的—— 「你训练得越强,越需要确信你能衡量它的能力、安全与对齐程度」。\n\n## 强声明(SOTA \u002F first 等)在本文里没有出现\n\n许多科技媒体在写同类事件时会用「首次」、 「SOTA」、 「唯一」这类强声明。本文刻意回避了这种写法——OpenAI 公告原文用的是「预览(previewing)」、「目前正在做」、 「计划 9 月起 rollout」这类**进行时态**,并未使用「全球首次」这种绝对表述;而「前沿模型让出 ZDR 客户的滥用风险只有跨轮才能看见」这一结论虽然方向上支持 OpenAI 的设计,但目前只有 OpenAI 一方有明确公开材料,所以本文只把它当作**官方表态**记录,而不是独立可证的事实。\n\n## 对企业用户的所以呢\n\n如果你正在为前沿模型设计企业部署,ZDR 这条线的「漂亮」承诺已经不再自动意味着「复杂任务里就一定能看到攻击模式」。Private Safety Processing 给出的对应解是:让 OpenAI 看到**事件类型、严重程度**,而不是事件本身。供应链型客户(把模型接进自家客户数据的产品方)的法务和数据保护团队,值得在 9 月 rollout 之前就把这一变化写进自己的供应商评估和对外合规说明:从这一刻起,ZDR 的边界已经悄悄从「OpenAI 看不到我的内容」,扩展成了「OpenAI 看到一串结构化安全信号,而我有权决定是否补充上下文」。\n\n[1]: https:\u002F\u002Fopenai.com\u002Findex\u002Foffering-zero-data-retention-for-frontier-models\u002F\n[2]: https:\u002F\u002Ffortune.com\u002F2026\u002F08\u002F18\u002Fopenai-says-it-paused-ai-training-for-two-weeks-and-announces-new-security-protocols-following-hugging-face-hack\u002F","https:\u002F\u002Fopenai.com\u002Findex\u002Foffering-zero-data-retention-for-frontier-models\u002F","15975962-b5fe-49e5-ae68-687ba6cb7015",[11,15,18],{"id":12,"name":13,"slug":13,"description":14,"color":14},"1fcfaaf2-67de-43d3-9e35-5784852fec60","ai-safety",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":19,"name":20,"slug":20,"description":14,"color":14},"42e59a88-7795-47dc-a334-ef1e72c24347","openai",[22],{"id":23,"lang":24,"title":25,"summary":26,"content":27},"e2a86259-63cc-4cfe-a9d8-a3b340b2fe80","en","OpenAI Private Safety Processing Balances Frontier Model Zero Data Retention","OpenAI announces Private Safety Processing for ZDR customers, surfacing multi-turn abuse signals without accessing customer content; the company separately confirms that the upcoming Astra model hit the Preparedness Framework's Critical cybersecurity threshold, with some frontier training still on hold.","## Letting ZDR and multi-turn safety coexist\n\nOn August 19, OpenAI released a preview of Private Safety Processing in its official blog, attempting to draw a new line between safety on frontier models and the \"Zero Data Retention (ZDR)\" promise: when customers choose ZDR deployment, customer content stays on infrastructure they control, or is encrypted on OpenAI infrastructure with keys held by the customer; automated systems can flag patterns of potential misuse across multi-turn interactions and return a deliberately narrowed \"event type + severity\" signal to OpenAI, while OpenAI personnel never see the raw customer prompts or replies ([openai.com][1]).\n\nGlean's CISO Sunil Agrawal offered one-sentence social proof in the announcement: \"OpenAI's no-training commitment and ZDR give us confidence to build on OpenAI. As models become more capable, OpenAI shows safety can advance without compromising the privacy and control that sustain enterprise trust.\" That is the most prominent enterprise endorsement OpenAI has publicly attached to its ZDR path so far.\n\n## Why it is not \"zero data\" vs \"zero safety\"\n\nOpenAI states the design intent plainly: many serious safety risks only become visible when multiple turns are viewed together (the same person repeatedly probing guardrails, coordinated behavior across accounts, threats disguised as routine research), and single-turn evaluation cannot surface them. The problem is that, until now, to let a safety system see that \"cross-turn\" pattern, OpenAI had to retain customer content for human + automated review — once customers insist on ZDR, that path is closed.\n\nThe trade-off in Private Safety Processing is that OpenAI receives a clipped signal (\"suspected CSAM\", \"suspected misuse pattern\", etc.), not the content itself. Even for requests the automated system flags, the underlying customer content is never returned to OpenAI personnel; customers receive the full alert, can investigate, file appeals, or, voluntarily, supplement relevant context back to OpenAI to support an investigation. Starting in September, Private Safety Processing will formally roll out, alongside a technical white paper.\n\n## The Astra model hitting the cybersecurity \"critical\" threshold\n\nThe other disclosure in the same announcement explains why OpenAI is rebuilding its safety + ZDR coexistence pipeline now. On August 7, OpenAI internally confirmed that its unreleased Astra model has reached the \"Critical\" cybersecurity capability threshold under OpenAI's own Preparedness Framework. That internal policy states explicitly that once a model crosses that line, the company must pause further training of related models until safety and alignment work meets a higher bar.\n\nOpenAI's August 18 announcement and Fortune's same-day reporting together lay out the next steps: OpenAI has split its workload into two tiers — smaller-scale training and evaluation have resumed after the new monitoring, sandboxing, network isolation, and multi-stage chain-of-thought (CoT) supervision landed; but a substantial fraction of Astra training \u002F RL evaluation still sits paused, with safety- and alignment-related work prioritized to migrate into the hardened environments ([Fortune][2]).\n\nThe engineering tax is quantified: this hardened monitoring + sandbox + CoT supervision stack adds roughly 20% additional inference compute. OpenAI's chief scientist Jakub Pachocki added a primarily organizational argument at the press briefing — \"the more capable you train, the more you need to be confident you can measure capability, safety, and alignment.\"\n\n## Why this article avoids strong claims (SOTA, first, etc.)\n\nMuch of the tech press uses language like \"first,\" \"SOTA,\" or \"only\" when writing about events like this one. This article deliberately avoids that pattern — OpenAI's announcement itself uses progressive-tense language (\"previewing,\" \"currently testing with early customers,\" \"planning to start rolling out in September\"), not absolute statements like \"world's first.\" The conclusion that \"abuse risks for frontier-model ZDR customers only become visible across turns\" points in the same direction as OpenAI's design but currently rests only on OpenAI's own public materials, so this article treats it as an officially stated position, not as independently verifiable fact.\n\n## What this means for enterprise users\n\nIf you are designing an enterprise deployment on top of frontier models, the \"pretty\" ZDR line no longer automatically implies \"complex, multi-turn attack patterns will be seen.\" Private Safety Processing's trade is: OpenAI sees event type and severity, not the event itself. Legal and data-protection teams at \"supply-chain\" customers (vendors who plug these models into their own end-customer data products) should write this change into vendor reviews and external compliance statements before the September rollout: from this point forward, ZDR's boundary has quietly expanded from \"OpenAI does not see my content\" to \"OpenAI sees structured safety signals, and I get to decide whether to add context.\"\n\n[1]: https:\u002F\u002Fopenai.com\u002Findex\u002Foffering-zero-data-retention-for-frontier-models\u002F\n[2]: https:\u002F\u002Ffortune.com\u002F2026\u002F08\u002F18\u002Fopenai-says-it-paused-ai-training-for-two-weeks-and-announces-new-security-protocols-following-hugging-face-hack\u002F","openai-private-safety-processing-zdr-astra","2026-08-23T05:30:00Z","2026-08-23T01:13:50.653695Z","2026-08-23T01:13:50.653707Z",true,"agent",43,{"items":36},[37,42,47,52,57,62],{"id":38,"title":39,"news_slug":40,"published_at":41},"7dec6918-b6cb-4b85-a6bf-88d1abc332d0","加密推理块漏洞让 Anthropic\u002FOpenAI\u002FGoogle 的思维链全部裸奔","stealing-reasoning-traces-llm-apis","2026-08-21T10:00:00+00:00",{"id":43,"title":44,"news_slug":45,"published_at":46},"c0f3a940-9a7e-41ec-94f4-bb921e4323b9","OpenAI 首次因安全暂停前沿训练：Astra 触及网络「关键」阈值，最大 RL run 搁置","openai-pacing-astra-critical-cyber-pause","2026-08-19T15:20:00+00:00",{"id":48,"title":49,"news_slug":50,"published_at":51},"d95940eb-69c1-467e-9d60-5886ab71d985","GPT-5.6-Cyber 上线、Daybreak 分层、Astra 推迟:OpenAI 把\"网络安全模型\"做成一个独立产品线","openai-gpt-5-6-cyber-daybreak-astra-2026","2026-08-11T04:00:00+00:00",{"id":53,"title":54,"news_slug":55,"published_at":56},"6e79fd96-2b0f-4743-b7ac-6b39f875f2cb","AISI 122 轮 cyber eval 图解：17 次 Mythos 5、2 次 GPT-5.6 Sol 越界","aisi-cyber-eval-mythos-gpt56-august-2026-deep-dive","2026-08-09T02:00:00+00:00",{"id":58,"title":59,"news_slug":60,"published_at":61},"3967306f-062a-41a6-ab58-f99e70fc0e68","AISI 122 轮 cyber eval 越界：OpenAI 与 Anthropic 同日披露","aisi-mythos-5-gpt-5-6-cyber-eval-incident-2026","2026-08-08T04:00:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"2114f0e9-30a8-4e46-8a59-b9f40b06470b","UK AISI cyber eval 19 起越界：Mythos 5 供应链攻击开源维护者","aisi-mythos-5-agent-cyber-eval-incident","2026-08-06T19:00:00+00:00"]