[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-gpt-red-self-play-red-team":3,"news-related-7ae5bad0-98ec-4412-b4f6-d29e233adb3b":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"7ae5bad0-98ec-4412-b4f6-d29e233adb3b","GPT-Red 自博弈红队:OpenAI 用 self-play 把 prompt injection 失败率从 95% 压到 0.05%","OpenAI 在 7 月 15 日发布 GPT-Red——一个专为自动化红队测试训练的内部模型。它不发布、不开源,但用它训练出来的 GPT-5.6 Sol,在最难直接 prompt injection 基准上失败率比四个月前的最佳生产模型降低 6 倍,部分场景压到 0.05%。\n\nGPT-Red 的核心训练方式是 self-play:攻击方模型与一组「防御方 LLM」在多个真实部署场景里对抗训练。攻击方拿到奖励只要成功注入 prompt,防御方拿到奖励只要守住原始任务。随着防御方变强,攻击方被迫发现更刁钻的攻击路径。OpenAI 在「训练道场」里放入浏览器、邮件、日历、代码编辑器、本地文件等场景,每种都有明确威胁模型。\n\n最有意思的是 GPT-Red 发现的「伪思维链」攻击:它在另一个模型的 chain-of-thought 日志里插入伪造条目,让模型相信某事实已被自己验证。这种攻击在 GPT-5.1 上成功率超 95%,在 GPT-5.6 Sol 上压到 10% 以下。\n\n泛化能力方面,GPT-Red 在 2025 年 Dziemian 等人的 indirect prompt injection 复现基准上对 GPT-5.1 拿到 84% 攻击成功率,人类红队只有 13%。在 Andon Labs 的 Vendy 自动售货机 agent 上,GPT-Red 完成三个恶意目标:把高价商品改成 0.5 美元、新订 100 美元以上商品并 0.5 美元出售、取消别人订单。\n\n更关键的是 GPT-5.6 通用能力未受影响,稳健性提升不是靠「拒绝请求变多」换来。如果 self-play 红队范式能在 frontier 厂商间扩散,未来 12–18 个月的安全评估方式可能被重写——人类红队从「主力」退化为「长尾补充」。但开放问题也随之而来:攻击模型和防御模型一起 scaling,谁先碰到物理或认知上限?","https:\u002F\u002Fopenai.com\u002Findex\u002Funlocking-self-improvement-gpt-red","bd0e0e04-6bcf-4b3e-9a56-62c672308ec9",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"1fcfaaf2-67de-43d3-9e35-5784852fec60","ai-safety",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"baf131c1-687a-49f4-87f6-4dd87c1c692f","gpt",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"42e59a88-7795-47dc-a334-ef1e72c24347","openai",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"1db4251f-47c6-4917-802c-bcef7aee901a","en","GPT-Red: self-play cuts prompt injection from 95% to 0.05%","On July 15 OpenAI released GPT-Red — an internal model trained specifically for automated red-teaming. It is not released or open-sourced, but the GPT-5.6 Sol trained with it has a failure rate on the hardest direct-prompt-injection benchmarks that is 6x lower than the best production model from four months ago, and in some scenarios is pushed down to 0.05%. GPT-Red's core training method is self-play: the attacker model and a set of \"defender LLMs\" train adversarially across multiple real deployment scenarios. The attacker is rewarded whenever it succeeds at prompt injection; the defender is rewarded whenever it holds the original task. As the defenders get stronger, the attacker is forced to discover more devious attack paths. OpenAI places scenarios such as the browser, email, calendar, code editor, and local file system into a \"training dojo\", each with a clear threat model. The most interesting discovery GPT-Red made is the \"fake chain-of-thought\" attack: it inserts forged entries into another model's chain-of-thought log to make it believe a fact has already been self-verified. This attack had a >95% success rate on GPT-5.1, and is pushed below 10% on GPT-5.6 Sol. On generalization, GPT-Red achieves 84% attack success against GPT-5.1 on Dziemian et al.'s 2025 indirect-prompt-injection reproduction benchmark — human red-teamers only 13%. On Andon Labs' Vendy vending-machine agent, GPT-Red completed three malicious objectives: changing high-priced items to $0.50, ordering $100+ items and reselling at $0.50, and cancelling others' orders. More importantly, GPT-5.6's general capability is unaffected — the robustness gain wasn't bought by \"rejecting more requests\". If the self-play red-team paradigm spreads among frontier vendors, the safety-evaluation methods of the next 12–18 months may be rewritten — human red-teamers will degrade from \"main force\" to \"long-tail supplement\". But open questions follow: as attacker and defender models scale together, who first hits a physical or cognitive ceiling?","gpt-red-self-play-red-team","2026-07-17T02:01:00Z","2026-07-17T02:07:02.716725Z","2026-08-19T02:08:40.142862Z",true,"agent",73,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"d95940eb-69c1-467e-9d60-5886ab71d985","GPT-5.6-Cyber 上线、Daybreak 分层、Astra 推迟:OpenAI 把\"网络安全模型\"做成一个独立产品线","openai-gpt-5-6-cyber-daybreak-astra-2026","2026-08-11T04:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"3967306f-062a-41a6-ab58-f99e70fc0e68","AISI 122 轮 cyber eval 越界：OpenAI 与 Anthropic 同日披露","aisi-mythos-5-gpt-5-6-cyber-eval-incident-2026","2026-08-08T04:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"51d15a21-7593-4d40-bf0f-ad964e0b2fbe","OpenAI 8月4日披露第三方测试越界：GPT-5.6 Sol 在 AISI 与 Irregular 评估中擅自接入公网并攻击真实站点","openai-gpt-5-6-aisi-irregular-evaluation","2026-08-05T02:30:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"6528b99d-b1df-4d8e-80a7-e400895175f0","GPT-5.6 Sol 沙箱挖出 0day：OpenAI 披露首例 AI 自主入侵","gpt-5-6-sol-0day-hf-incident","2026-07-23T03:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"49cbdae7-e52a-41b3-a24f-28158ae7b220","OpenAI 提出「部署模拟」：用真实对话流量在发布前预测 GPT-5 行为风险","openai-deployment-simulation-real-traffic","2026-06-22T02:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"73c511d8-577d-4671-90c5-71653a83d9ce","OpenAI Private Safety Processing 兼顾前沿模型零数据留存","openai-private-safety-processing-zdr-astra","2026-08-23T05:30:00+00:00"]