[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-openai-gpt-5-6-august-update-reasoning-slider":3,"news-related-d1c7b405-fe4e-40f9-9249-a12e2bba6913":38},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"d1c7b405-fe4e-40f9-9249-a12e2bba6913","GPT-5.6 八月更新：把「推理强度滑块」下放给 Plus\u002FPro，同时把免费用户拉进 Luna 时代","OpenAI 发布 GPT-5.6 Sol\u002FLuna 的八月更新,把推理强度滑块下放给 Plus\u002FPro,Luna 同步扩展到 Free 用户,旧版 GPT-5.5 Instant 全面替换;系统卡同时把新模型定性为 Cybersecurity 与 Bio\u002FChemical 的 High capability。","## 一句话\n\nOpenAI 今天上线 GPT-5.6 Sol \u002F Luna 的 8 月更新,核心改动有三层:一是 Plus\u002FPro 用户拿到一个能手动调「推理强度」的滑块,二是 Luna 这一档正式向 Free 用户开放,三是旧版 GPT-5.5 Instant 全面下线。Deployment Safety Hub 同日发布完整 system card,Cybersecurity 与 Bio\u002FChemical 两项被定性为 **High capability**(但未达 Critical)。([OpenAI Deployment Safety][1], [OpenAI 博客][2])\n\n## 这次到底改了哪些东西\n\n按 OpenAI Deployment Safety Hub 的官方描述:\n\n- **Free \u002F Go 用户**——拿到一个新的默认模型用于日常对话,这是 Luna 这一档首次覆盖免费层。\n- **Plus \u002F Pro 用户**——拿到更新后的 GPT-5.6 Sol,关键 UX 变化是新增了一个「推理强度滑块」,用户可以自己决定 ChatGPT 在每次回答上花多少算力。\n- **Codex \u002F ChatGPT Work**——这次**没有切**,仍在使用 7 月版本(OpenAI 在 system card 里显式区分了「August release」和「July release」),这是为了企业场景的稳定性考虑。\n- 旧版 **GPT-5.5 Instant** 被全面替换,免费层的默认模型间接完成了一次升级。\n\n这条产品路径有意思的点:**把「推理强度」这个变量第一次做成了用户可见的旋钮**,而不是只放在 API 参数里。在 OpenAI 的部署哲学里,这等于把「深度思考」与「快速回应」之间的取舍权,从开发者手里下放给普通用户。\n\n## 安全评估里藏了哪些细节\n\n这次 system card(8 月版)里值得关注的安全\u002F能力数字:\n\n1. **能力定性**:8 月版 Sol \u002F Luna 在 **Cybersecurity** 和 **Biological & Chemical** 两个领域都被定为 **High capability**,但**没有达到 Critical**。两套安全护栏沿用 7 月版的策略,没有新增。\n2. **Cyber range 综合通过率**:Sol(8 月版)**83.3%**,Luna(8 月版)**61.5%**,对照 7 月版 Sol 与 GPT-5.5 Thinking 的 92.3% 反而**略有下降**(尤其在 EDR Evasion、Leaked Token、Binary Exploitation 几个场景)。这是少数能拿来横向对比的具体数字。\n3. **Capture The Flag(内部评测)**:Sol 直接打到 **97.06%**,**饱和**了整张评测表。说明在「找出并利用漏洞」这一单项上,8 月版的 Sol 已经接近单层防线下的能力上限。\n4. **CVE-Bench(真实 Web 应用漏洞)**:Sol **高于 High 阈值**,Luna **低于阈值**。两个变体在同一基准上出现的能力差,值得作为「旗舰 vs 性价比」分档的实证。\n5. **U18 专项评估**:首次加入了面向 18 岁以下用户的专属评测,覆盖 self-harm、eating disorders、age-restricted goods、sexual content 等。Sol \u002F Luna 在 Eating Disorders 上拿到 **0.808 \u002F 0.810**,Age-restricted goods 上 **0.865 \u002F 0.857**,相比 GPT-5.5 Instant 6 月版有可见提升。\n\nOpenAI 自己给这次更新的总结是:模型**更容易发现并修复漏洞**,但**尚未具备可靠执行端到端 hardened 攻击的能力**。这也是为什么新版依然以「defender 友好」的姿态放出,而不是限制 access。\n\n## 为什么这次更新值得盯\n\n业内通常不把这种「小版本更新」当作重要事件,但 8 月版有几条**结构性变化**值得拎出来:\n\n- **产品层**——把「reasoning effort」做成用户可见的滑块,等于 OpenAI 第一次承认 **「深度思考 = 商品」**,这会影响未来所有 to-C 产品的设计:任何把 reasoning 当核心能力的应用,都会被用户拿来跟这个滑块对标。\n- **能力层**——Cyber capability 在 Sol 上**几乎饱和**,Capture The Flag 97%、CVE-Bench 超阈值;但 Luna 在同一基准上还**不到阈值**。这条分水岭说明:**同一个名字下,旗舰与中端的能力差,可能大到「能力分档」的程度**——开发者选型时不能只看名字。\n- **安全层**——明确把 Cyber 与 Bio\u002FChem 定为 **High**,这是 OpenAI 自家 Preparedness Framework 里「已具备缓解措施,但仍需持续监控」的等级;**没有定到 Critical**(Critical 的定义是「可独立完成零日漏洞的发现与利用」),所以并未触发额外 access 控制。\n- **企业层**——Codex \u002F ChatGPT Work **这次没切**,留在了 7 月版。这条细节常被忽略,但对企业 SLA 而言很关键:OpenAI 在悄悄给「to-B 稳定」与「to-C 能力」做版本分层。\n\n## 所以呢\n\n这次的 GPT-5.6 August Update,**真正重要的不是参数或跑分,而是 OpenAI 第一次把「推理强度」这件在 API 时代只是  参数的事,做成了普通用户能看见、能调的旋钮**。\n\n它会带来几个可观察的连锁反应:\n\n1. **面向 C 端的 AI 助手**,接下来 6 个月内会集体跟进 reasoning effort UI,**因为有 OpenAI 这个锚**——任何不带这个滑块的产品,会被用户感知为「不如 ChatGPT 灵活」。\n2. **能力分档被显性化**:Sol 与 Luna 在同一个安全评测里出现 High vs sub-High 的能力差,意味着「同一个 GPT-5.6 名号」下面其实有**能力断层**,开发者在 agent 设计时需要明确选 Sol 还是 Luna。\n3. **Cyber capability 进入「饱和前夜」**:Sol 在 CTF 内部评测已经 97%,CVE-Bench 超阈值,再往前一步就是 OpenAI 自己 Preparedness Framework 里 **Critical** 的定义——**这条边界,大概率是 2026 年下半年 OpenAI 系统卡最值得盯的一条线**。\n4. **企业版本分层会成为常态**:Codex \u002F Work 留 7 月版、ChatGPT 主产品换 8 月版,这种「企业用旧、用户用新」的做法会扩散,做 to-B 集成的开发者要习惯「名字相同,行为不同」。\n\n「更新」这个词在 OpenAI 这家公司里,**越来越不等于「升一档」,而是「把上一档切成两个版本,分别下放到不同的用户层」**。这种节奏一旦固定,产品、安全与商业三层会同时被重新定义。([OpenAI Deployment Safety][1], [OpenAI 博客][2])\n\n---\n\n[1]: https:\u002F\u002Fdeploymentsafety.openai.com\u002Fgpt-5-6-august-update\n[2]: https:\u002F\u002Fopenai.com\u002Findex\u002Fimproving-gpt-5-6-sol-in-chatgpt\u002F","https:\u002F\u002Fdeploymentsafety.openai.com\u002Fgpt-5-6-august-update","55a458a0-bca3-4d8a-a4ac-d6b0aeb9d2ab",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"baf131c1-687a-49f4-87f6-4dd87c1c692f","gpt",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"42e59a88-7795-47dc-a334-ef1e72c24347","openai",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"bb988d2f-d21d-498e-96fc-f7a7afe34ea9","en","GPT-5.6 August update: reasoning sliders for Plus and Pro","OpenAI has shipped the August update of GPT-5.6 Sol and Luna: a user-facing reasoning-effort slider for Plus\u002FPro, expanded Luna access for Free users, and full retirement of GPT-5.5 Instant. The same-day system card classifies both new variants as High capability in Cybersecurity and Bio\u002FChemical domains, but below Critical.","## TL;DR\n\nOpenAI shipped the August update of GPT-5.6 Sol and Luna today. Three changes matter: Plus and Pro users get a user-facing **reasoning-effort slider**, Luna is now available to Free-tier users, and the older GPT-5.5 Instant is fully retired. The same-day system card classifies the new Sol and Luna as **High capability** in both Cybersecurity and Biological & Chemical domains — but still below the Critical threshold. ([OpenAI Deployment Safety][1], [OpenAI blog][2])\n\n## What actually changed\n\nPer the official Deployment Safety Hub description:\n\n- **Free and Go users** get a new default model for everyday chats. This is the first time the Luna tier has been opened up to free users.\n- **Plus and Pro users** get the updated GPT-5.6 Sol. The most visible UX change is a new **reasoning-effort slider** that lets users decide how much compute ChatGPT spends on each response.\n- **Codex and ChatGPT Work** are **not** moving to the August model yet. They stay on the July release — OpenAI explicitly distinguishes the \"August release\" from the \"July release\" in the system card to keep enterprise environments stable.\n- The old **GPT-5.5 Instant** has been retired across the board, which means the Free tier's default model is indirectly upgraded.\n\nThe interesting shift: OpenAI is making \"reasoning depth\" a **user-visible knob** rather than just an API parameter (). In product terms, that is OpenAI moving the trade-off between \"deep thinking\" and \"fast answer\" from the developer to the end user.\n\n## What the safety card actually says\n\nThe August system card contains concrete numbers worth pulling out:\n\n1. **Capability classification.** Both August Sol and Luna are classified as **High capability** in Cybersecurity and Biological & Chemical domains, but **below Critical**. The same safeguards from the July release apply — no new mitigations were added.\n2. **Cyber range combined pass rate.** Sol (August) reaches **83.3%**, Luna (August) reaches **61.5%**, while the July Sol and GPT-5.5 Thinking sat at 92.3%. The August Sol is actually a slight regression in scenarios like EDR Evasion, Leaked Token, and Binary Exploitation.\n3. **Internal Capture The Flag.** Sol hits **97.06%** — saturating the eval set. On a single capability axis (find-and-exploit), the August Sol is near the ceiling of this evaluation.\n4. **CVE-Bench (real-world web vulnerabilities).** Sol is **above the High threshold**; Luna is **below**. That gap inside a single model family is a clean empirical \"flagship vs mid-tier\" split.\n5. **U18 evaluations.** OpenAI introduced dedicated under-18 evaluations covering self-harm, eating disorders, age-restricted goods, and sexual content. Sol and Luna score **0.808 \u002F 0.810** on Eating Disorders, **0.865 \u002F 0.857** on Age-restricted goods — visible improvement over the GPT-5.5 Instant June update.\n\nOpenAI's own framing: the updated models are **better at finding and fixing vulnerabilities** than at reliably executing end-to-end attacks against hardened targets. That is why access is opened up for defenders, with targeted safeguards and monitoring rather than new restrictions.\n\n## Why this update matters\n\nOpenAI usually downplays these \"small\" updates, but the August release has structural shifts worth tracking:\n\n- **Product layer.** Making reasoning-effort a user-visible slider is OpenAI's first admission that **\"depth of thought\" is a productizable commodity**. Anything that builds \"deep reasoning\" into the brand — from Cursor to Perplexity to Manus-style agents — will now be compared against this slider.\n- **Capability layer.** Sol nearly saturates the internal CTF eval (97%) and clears the CVE-Bench High threshold, while Luna is below the same threshold. A single model name (GPT-5.6) now hides a real capability cliff. Developers building agents need to pick Sol or Luna explicitly based on this gap, not based on the name.\n- **Safety layer.** Cyber and Bio\u002FChem are explicitly flagged as **High** under OpenAI's Preparedness Framework, which is the tier where mitigations exist but continuous monitoring is required. They are **not** at **Critical** (the threshold defined as \"find and develop functional zero-days in hardened systems without human intervention\"), so no extra access controls have been triggered yet.\n- **Enterprise layer.** Codex and ChatGPT Work stayed on July. That detail is easy to miss but matters for enterprise SLAs: OpenAI is quietly splitting the **to-B stability track** from the **to-C capability track** within the same model family.\n\n## So what\n\nThe real signal of this August update is **not** parameter size or benchmark deltas. It is OpenAI turning  — an API parameter that has been quietly important since o1 — into a UI slider that any ChatGPT user can grab.\n\nThe downstream consequences worth watching:\n\n1. **Every consumer AI assistant will add a reasoning-effort knob in the next six months**, simply because OpenAI has set the anchor. Products without one will feel \"less flexible\" by default.\n2. **The capability gap inside one model name is now official.** Sol and Luna land on opposite sides of the same High-cyber threshold. The lesson: **agent builders should pick by capability tier, not by model name**.\n3. **Cyber capability is approaching saturation.** Sol already at 97% on internal CTF and above High on CVE-Bench means the next step toward the Preparedness Framework's **Critical** definition is genuinely close. This is the single most important threshold to track in OpenAI's next two system cards.\n4. **Enterprise vs consumer version splitting is becoming the norm.** Codex\u002FWork on July, ChatGPT on August — that pattern will spread. Integrators should expect \"same model name, different behavior\" to be the default assumption.\n\nIn OpenAI's vocabulary, **\"update\" no longer means \"step up one tier\". It means \"split the previous tier into two versions and place each on a different user surface.\"** Once that cadence sticks, the product, safety, and business layers all get reshaped at the same time. ([OpenAI Deployment Safety][1], [OpenAI blog][2])\n\n---\n\n[1]: https:\u002F\u002Fdeploymentsafety.openai.com\u002Fgpt-5-6-august-update\n[2]: https:\u002F\u002Fopenai.com\u002Findex\u002Fimproving-gpt-5-6-sol-in-chatgpt\u002F","openai-gpt-5-6-august-update-reasoning-slider","2026-08-10T20:00:00Z","2026-08-10T20:06:02.309788Z","2026-08-10T20:06:02.309797Z",true,"agent",185,{"items":39},[40,45,49,54,59,64],{"id":41,"title":42,"news_slug":43,"published_at":44},"b9e635eb-ac5e-412d-8904-f113ad3fd5ec","微软宣布工程师 AI token 预算上限并把 OpenAI GPT-5.6 Sol 设为 GitHub Copilot 内部默认模型","microsoft-copilot-gpt-5-6-sol-default-token-budget-0806","2026-08-05T16:00:00+00:00",{"id":46,"title":47,"news_slug":47,"published_at":48},"4afbd081-d5dd-4595-80f8-dd72c69a136a","GPT-5.5 让模型在发布前先改自己跑的引擎：这不是新模型,是 OpenAI 的 release 范式更新","2026-08-03T18:00:00+00:00",{"id":50,"title":51,"news_slug":52,"published_at":53},"0fd9ee7a-5b8f-49d2-9032-57f763de40e3","OpenAI 下一代模型 Astra 一口气破解 10 个数学难题:从 27 年未决的非 sofic 群到 46 年未动的高维球体堆积","openai-astra-ten-math-proofs-2026","2026-08-01T10:00:00+00:00",{"id":55,"title":56,"news_slug":57,"published_at":58},"418a9ac0-18fd-49a4-b7a8-d29d1c1ba497","AI 承诺的四天工作制为什么没来：OpenAI \u002F Anthropic 内部工时真相","ai-four-day-work-week-myth-openai-90-hours","2026-08-16T03:30:00+00:00",{"id":60,"title":61,"news_slug":62,"published_at":63},"31f3215c-0892-419d-a610-fe815cc60bbe","GPT-5.6 降价 80% 把竞争拉进「同等智能成本」：DeepSeek V4 Flash 接招，国产模型卡出双线赛道","gpt-5-6-luna-price-cut-equal-intelligence-cost","2026-08-12T03:00:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"d95940eb-69c1-467e-9d60-5886ab71d985","GPT-5.6-Cyber 上线、Daybreak 分层、Astra 推迟:OpenAI 把\"网络安全模型\"做成一个独立产品线","openai-gpt-5-6-cyber-daybreak-astra-2026","2026-08-11T04:00:00+00:00"]