[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-gemini-3-8-flash-cyber-fairwind-launch":3,"topics-all":41,"news-related-e0a484e4-41e0-4f91-9b9b-a196bbdcf3ba":60},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":27,"news_slug":34,"published_at":35,"created_at":36,"modified_at":37,"is_published":38,"publish_type":39,"image_url":14,"view_count":40},"e0a484e4-41e0-4f91-9b9b-a196bbdcf3ba","Gemini 3.8 Flash 双发:同价升级 + Cyber 走可信项目 Fairwind","Google 9 月 2 日上线 Gemini 3.8 Flash 与 3.8 Flash Cyber 两个变体。前者主攻长链路编程与 Agent,价位沿用 3.7;后者通过可信项目向防御方开放,CyberGym、CWE-Bench 等基准跑到前沿水平。","## 技术背景\n\nGoogle 过去六周连发 Gemini 3.6 Flash Lite、3.7 Flash、3.5 Flash Cyber 等多个变体,密集程度让人没消化完基准对比就又来新版。9 月 2 日官方又甩出 3.8 Flash 与 3.8 Flash Cyber 两个变体,沿用「通用旗舰 + 防御专用」的「同一个底座 + 不同护栏」拆法——把网络安全等高敏感场景单独拉出来,搭配更宽松的护栏交给可信方。这与 Claude Fable\u002FMythos 5.1、DeepSeek 主线\u002F后训练变体的思路一致。\n\n## 核心内容\n\n3.8 Flash 继续是主力工作马。官方明确说「3.8 Flash works harder」——复杂任务上多花 token、多跑工具迭代,换更稳的结果。基准上 DeepSWE v1.1 接近更大参数量的前沿模型;Vals Finance Agent V2 与 Harvey 的 Legal Agent Benchmark 反超 3.7 Flash 与多数前沿模型;HLE-Verified 跑到 54.9%,能跨 STEM \u002F 人文 \u002F 专业领域做多步推理。价位沿用 3.7:输入 0.75 美元\u002F百万 token,输出 3.75 美元\u002F百万 token,促销到 2026 年 12 月 31 日,2027 年 1 月 1 日起涨到 1.50 和 7.50——与 3.7 一模一样的两档,等于「升级不涨价」。\n\n3.8 Flash Cyber 是这次最有意思的部分。Google 把同底座的 Cyber 变体放进新上线的 Fairwind 项目,首批 650+ 政府与基础设施合作伙伴优先拿到权限。Cyber 面向真实世界的漏洞发现与修复:跨 20 种编程语言的内部基准漏洞发现成功率超过 70%;CWE-Bench(Collinear 运营的外部补丁基准)Pass@1 跑到 47.2%,与第一梯队前沿模型的 47.8% 几乎并列,价格却低一档。Chrome 安全团队的内部测试里,Cyber 给出的正确补丁数量是市面上更强商业模型的 2.6 倍;Wiz 的内部渗透测试基准上召回率高 7.5-9.7%,成本低 2.3-5.2 倍;Google Cloud 漏洞研究团队用 Cyber 在不到 2 小时内找到一个关键基础漏洞——通常要数月研究。\n\nFairwind 本身的设计跟 Anthropic Mythos 5.1 Trusted Access、Life Sciences Verification Program 是同一类合规架构:同样权重挂「更宽松」护栏,访问能力锁定在「组织内部安全 \u002F 应急响应 \u002F 渗透测试」三类岗位,要求 MFA 等强安全规范。但 Google 措辞更强调「修复 vs 利用」——明确把 Cyber 的优化重点放在「写并验证补丁」而非「写漏洞利用」,把自动修复能力放进自家 CodeMender harness 跟模型一起打包。\n\n## 个人评论\n\n最值得关注的是定价。Gemini 3.8 Flash 把价位锁在 3.7 同档,等于「3.7 的价格、3.8 的能力」——对企业预算和现有 Gemini 工作流完全是零摩擦。对照同期 Claude Fable 5.1 cache read 砍 75%、Anthropic Mythos 5.1 走「高能力 + 严格门槛」、Google 选「同价升级 + 拆出 Cyber 走可信项目」,三家在护栏 + 价格组合上做出了不同取舍。\n\n对网络安全实际影响还要打折扣。Cyber 的基准表现亮眼(70%+ 跨语言漏洞发现、47.2% Pass@1 CWE-Bench、Chrome 内部 2.6 倍正确补丁),但这些都是在 Google 自身和合作伙伴的内部测试上的数字,Cyber 仅向可信方开放,目前无法外部独立验证。Google 也没披露 Cyber 在「良性代码审计」与「恶意利用生成」上的护栏差异细节,只笼统说「投入漏洞修复高于攻击能力」。这跟 Anthropic 在 Mythos 5.1 上对「提示注入与对齐」的更细颗粒披露相比,透明度低一档。\n\n对国内模型战的暗示:Google 用「六周连发三版 + 通用版同价不涨 + 防御版走可信项目」的组合拳,对各家都是压力。Qwen3.8、DeepSeek V4、Kimi K3、GLM-5.3 这一波国产旗舰正在抢「同等智能下的成本曲线」;Gemini 3.8 Flash 用 0.75 美元\u002F百万输入维持同价,把「Gemini 比开源贵但比前沿闭源便宜」的位置守得更紧。2026 下半年的模型选型不再是「最强 vs 最便宜」二选一,而是「最强 vs 最便宜 + 配套生态」的多维度博弈。\n\n参考链接:Google 官方博客《Introducing Gemini 3.8 Flash and 3.8 Flash Cyber》(blog.google,2026-09-02);Google 官方博客《Fairwind Program:Proactive cyber defense for governments and enterprises》(blog.google,2026-09-02);BenchLM model-updates(Gemini 3.8 Flash + 3.8 Flash Cyber,2026-09-02 收录)。","https:\u002F\u002Fblog.google\u002Finnovation-and-ai\u002Fmodels-and-research\u002Fgemini-models\u002F3-8-flash-and-3-8-flash-cyber\u002F","d884df39-706a-45db-b7ed-371f12e54f1f",[11,15,18,21,24],{"id":12,"name":13,"slug":13,"description":14,"color":14},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"a9524a82-a7c5-4daa-bb4b-a7ee77bb0b94","gemini",{"id":19,"name":20,"slug":20,"description":14,"color":14},"8cf7490f-2449-4ba7-be19-61befa0d92b4","google",{"id":22,"name":23,"slug":23,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":25,"name":26,"slug":26,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",[28],{"id":29,"lang":30,"title":31,"summary":32,"content":33},"3d1b4d2a-4efa-486d-babc-ec2e3d2c7014","en","Gemini 3.8 Flash dual launch: same-price upgrade + Cyber through Fairwind trusted program","Google released Gemini 3.8 Flash and Gemini 3.8 Flash Cyber on September 2. The former targets long-horizon coding and agents at the same price as 3.7; the latter ships through the Fairwind trusted program for cyber defenders and reaches frontier performance on CyberGym and CWE-Bench.","## Background\n\nGoogle has stacked up Gemini releases in the last six weeks: 3.6 Flash Lite, 3.7 Flash, 3.7 Flash Lite, 3.5 Flash-Lite, 3.5 Flash Cyber, and more — each announcement arriving before the benchmark chatter from the previous one dies down. On September 2 Google added Gemini 3.8 Flash and Gemini 3.8 Flash Cyber to the lineup. Both share the same underlying weights but split into a general-purpose flagship and a defense-only variant with looser safeguards — the same playbook Claude Fable 5.1 \u002F Mythos 5.1 and DeepSeek's main \u002F post-trained variants are running. The pattern is to peel out high-stakes domains like cybersecurity and life sciences, give vetted users more permissive guardrails, and keep the public surface narrow.\n\n## Core content\n\n3.8 Flash is the workhorse. Google explicitly says the model \"works harder\" — spending more tokens and running longer tool-iteration loops on complex tasks in exchange for steadier results. On DeepSWE v1.1 (long-horizon software engineering agents) it approaches many larger frontier models. On Vals Finance Agent V2 and Harvey's Legal Agent Benchmark it beats 3.7 Flash and most frontier models. HLE-Verified lands at 54.9%, demonstrating multi-step reasoning across STEM, humanities, and professional domains. Pricing stays at 3.7 levels: \u002Fbin\u002Fbash.75 per million input tokens and .75 per million output tokens during the introductory window through December 31, 2026, then .50 and .50 starting January 1, 2027 — exactly the same two-step ladder as 3.7, holding the price curve flat while the model climbs.\n\n3.8 Flash Cyber is the most interesting piece. Google bundled the same-base Cyber variant into the new Fairwind program, where 650+ government and infrastructure partners get priority access. Cyber targets real-world vulnerability discovery and remediation rather than the C\u002FC++-only CyberGym benchmark: on Google's internal 20-programming-language benchmark the vulnerability discovery success rate exceeds 70%; on the externally run CWE-Bench (operated by Collinear) Pass@1 hits 47.2%, almost tied with a leading frontier model at 47.8% but at a significantly lower cost. Internal Chrome Security testing found 3.8 Flash Cyber produced 2.6x more correct patches than stronger commercial models. Wiz's internal penetration testing benchmark showed +7.5-9.7% recall at 2.3-5.2x lower cost. Google Cloud Vulnerability Research used Cyber to find a critical foundational vulnerability in under two hours — work that usually takes months of research.\n\nFairwind itself is the same compliance architecture as Anthropic's Mythos 5.1 Trusted Access and Life Sciences Verification Program: the same weights ship with looser guardrails, but access is locked to three job categories — internal security, incident response, and penetration testing — and demands MFA and other strong operational standards. Google leans harder on the \"remediation vs exploitation\" framing, explicitly prioritizing Cyber for writing and validating patches rather than exploitation, and bundling it with the in-house CodeMender harness so defenders get model + auto-fix loop in one package.\n\n## Commentary\n\nThe pricing is the bigger story than the model. Pinning 3.8 Flash at 3.7 pricing means \"3.7's price, 3.8's capability\" — zero friction for enterprise budgets and existing Gemini workflows. Compare that to Claude Fable 5.1 cutting cache-read pricing 75% and Anthropic's Mythos 5.1 going \"frontier capability + strict access\" while Google chooses \"same-price upgrade + peel out Cyber into a trusted program\": three different takes on guardrails plus pricing from three frontier labs.\n\nThe cyber impact is real but discount what Google says. Cyber's benchmarks look strong — 70%+ cross-language vulnerability discovery, 47.2% Pass@1 CWE-Bench, 2.6x more correct patches than stronger commercial models in Chrome — but those numbers come from Google's own and partner-internal tests, and Cyber ships only to trusted defenders with no independent external verification available. Google also does not disclose the guardrail gap between benign code audit and malicious exploit generation, only the vague line about \"investing more in vulnerability remediation than offensive capability.\" That is thinner disclosure than Anthropic provides on Mythos 5.1's prompt-injection and alignment posture.\n\nThe pressure on Chinese model labs is real. Google's combo — three releases in six weeks, same-price upgrade on the general flagship, defense variant through a trusted program — sets the bar higher. The current Chinese flagship wave (Qwen3.8, DeepSeek V4, Kimi K3, GLM-5.3) is fighting for the \"intelligence-at-cost\" curve. Gemini 3.8 Flash holding the line at \u002Fbin\u002Fbash.75 per million input tokens keeps \"Gemini is pricier than open weights but cheaper than frontier closed\" locked in. The second half of 2026 is no longer \"strongest vs cheapest\"; it is \"strongest vs cheapest plus ecosystem\" — a multi-axis buyer problem.\n\nReferences: Google blog \"Introducing Gemini 3.8 Flash and 3.8 Flash Cyber\" (blog.google, 2026-09-02); Google blog \"Fairwind Program: Proactive cyber defense for governments and enterprises\" (blog.google, 2026-09-02); BenchLM model-updates (Gemini 3.8 Flash + 3.8 Flash Cyber, 2026-09-02).","gemini-3-8-flash-cyber-fairwind-launch","2026-09-03T03:00:00Z","2026-09-03T05:07:19.717446Z","2026-09-03T05:07:19.717454Z",true,"agent",200,[42,51],{"slug":43,"tag_slug":43,"title_zh":44,"title_en":45,"intro_zh":46,"intro_en":47,"id":48,"is_active":38,"created_at":49,"modified_at":50},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":52,"tag_slug":52,"title_zh":53,"title_en":54,"intro_zh":55,"intro_en":56,"id":57,"is_active":38,"created_at":58,"modified_at":59},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":61},[62,67,72,77,82,87],{"id":63,"title":64,"news_slug":65,"published_at":66},"7b9cdf6e-5ef0-4ece-ab6c-e8cec1b02397","Google 重组 DeepMind 领导层,Gemini 研发提速应对 Anthropic 与 OpenAI 竞争","google-deepmind-reshuffle-gemini-speed","2026-08-25T07:00:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"bcedeb8e-e5eb-4bbc-98b8-ea12f869055f","Google 收编 DeepMind：25 年最大 AI 重组","google-deepmind-centralization-gemini-catchup","2026-08-14T08:00:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"4bd8e8bd-7066-4ab7-bd97-e24ea3921395","Gemini 因编程落后推迟两月:Brin 4 月督促背后,Google 把研发「收回到一个人」手里的组织账本","google-gemini-coding-behind-deepmind-reshuffle","2026-08-14T03:30:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"16856034-439d-4915-aed4-80b42ae09c68","Gemini 3.7 Flash：FrontierCode 43.6%，价格腰斩","gemini-3-7-flash-coding-agent-fast","2026-08-13T09:00:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"8b6c20ec-7222-48cf-af2c-ac97466a2b0a","Gemini 月活破 10 亿:Google 第一次把 AI 助手做成自家「最快十亿用户产品」","gemini-app-1b-monthly-users","2026-08-12T03:00:00+00:00",{"id":88,"title":89,"news_slug":90,"published_at":91},"34edaffc-6b5c-4df1-9e2f-d864cada6063","Gemini 走进 K-12 课堂：Google 把「上下文」塞进每个作业","gemini-classroom-k12-contextualized-prompts","2026-08-07T02:00:00+00:00"]