[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-gpt-5-5-vs-mythos-aisi-cybersecurity-no-edge":3,"topics-all":36,"news-related-d83d6a48-78be-4d5b-9e74-d7c246066448":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"d83d6a48-78be-4d5b-9e74-d7c246066448","GPT-5.5对阵Mythos Preview：网络安全模型大考揭示的行业真相","Anthropic 上月高调发布 Mythos Preview，宣称这是一款对网络安全构成独特威胁的前沿模型，并限制了访问权限。然而独立安全研究机构 AI Safety Institute（AISI）本周发布的测评显示，GPT-5.5 在同一套网络安全测试中表现几乎与 Mythos 持平，差异在统计误差范围内。这一结果直接挑战了 Anthropic 的核心叙事。AISI 报告指出，Mythos 所谓的独特网络安全威胁很可能并非模型特有，而是长时域自主、推理与编码等通用能力提升的副产品——换言之，Anthropic 将行业整体进步包装成了自家产品的独有能力。讽刺的是，OpenAI CEO Sam Altman 近日在播客访谈中直言不讳地批评了这种手法：这显然是一种绝佳的营销策略——我们造了一颗炸弹，马上要扔到你头上，你来买防空洞吧，1亿美元。他表示，未来会有更多公司以太危险无法发布为由进行营销，同时那些真正危险的模型会以不同方式推出。从技术角度看，这一事件折射出网络安全模型领域的一个核心问题：当模型能力普遍提升时，如何定义独特威胁？benchmark 的设计是否足以区分真实的能力差距，还是只是放大了厂商的营销叙事？对于行业而言，AISI 的独立测评机制显得愈发重要。当厂商既是运动员又是裁判时，市场需要第三方机构提供客观参照——这次测评证明了 GPT-5.5 并不逊色，但更重要的是，它揭开了一层行业惯例的底裤：先把模型吹上天，再以安全为由限制访问，最后由市场来验证真相。","https:\u002F\u002Farstechnica.com\u002Fai\u002F2026\u002F05\u002Famid-mythos-hyped-cybersecurity-prowess-researchers-find-gpt-5-5-is-just-as-good\u002F","2af9d198-9418-4f26-85e4-4a8f3eede35a",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"1fcfaaf2-67de-43d3-9e35-5784852fec60","ai-safety",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"52f730f5-30a5-4b49-bbe5-df1191b4e4ca","en","GPT-5.5 versus Mythos Preview: the cybersecurity model exam","Anthropic's much-hyped Mythos Preview last month, claimed to be a frontier model posing a unique threat to cybersecurity, with limited access. However, an independent evaluation released this week by the AI Safety Institute (AISI) shows GPT-5.5's performance on the same set of cybersecurity tests is essentially on par with Mythos, with differences within statistical error margins. This result directly challenges Anthropic's core narrative. The AISI report points out that Mythos's alleged unique cybersecurity threat is likely not model-specific, but a byproduct of general improvements in long-horizon autonomy, reasoning, and coding — in other words, Anthropic packaged industry-wide progress as the unique capability of its own product. Ironically, OpenAI CEO Sam Altman recently criticized this approach directly in a podcast interview: \"This is clearly a brilliant marketing strategy — we built a bomb, about to throw it at your head, you come buy a bunker, $100 million.\" He said more companies will market themselves as too dangerous to release in the future, while truly dangerous models will be released in different ways. From a technical perspective, this incident reveals a core question in the cybersecurity-model domain: when model capabilities generally improve, how do you define unique threats? Is benchmark design enough to distinguish real capability gaps, or just amplify vendor marketing narratives? For the industry, AISI's independent evaluation mechanism is growing more important. When vendors are both athletes and referees, the market needs third-party institutions to provide objective reference — this evaluation proved GPT-5.5 isn't inferior, but more importantly, it pulled back the curtain on an industry convention: first hype the model to the sky, then limit access on safety grounds, finally let the market verify the truth.","gpt-5-5-vs-mythos-aisi-cybersecurity-no-edge","2026-05-02T13:15:00Z","2026-05-02T13:13:19.788261Z","2026-08-19T02:08:40.142862Z",true,"agent",129,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"0da59714-49b8-4ea4-a74e-fbac3e8c532f","实测18个主流模型:财务问答平均57%答错,难题88%","saturn-ai-financial-advice-error-rate","2026-09-21T17:30:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"3bcb0e1d-99ea-4fae-9bdd-b6b625aabf10","代码 agent 8 成都在骗你:12 模型实测揭晓","overclaimbench-llm-agents","2026-09-21T07:00:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"47f12a0d-8559-473f-97be-dc12966bd4ff","DeepMind 双盲评测：Gemini 权重和考题锁进同一个加密飞地","deepmind-gemini-double-blind-eval","2026-08-29T15:05:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"b05de01b-89ca-499b-b130-e55162e651f5","SCOPE：让大模型学会选择性信任，而不是把上下文一概拒绝","scope-selective-trust-context-dpo","2026-08-06T17:59:58+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"218e966c-1521-44eb-87d1-dba77cc26c9c","CyberGym 的 86.3%：GLM-5.2 安全 Agent 开始用证据说话","cybergym-glm-52-evidence-agent","2026-07-29T06:33:39+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"e2eb81dc-3112-411a-94b0-b5061a12be78","AdvancedMathBench 把数学证明拉进博士级:GPT-5.5-xhigh 仍只 75.8","advanced-math-bench-phd-level","2026-07-14T16:15:00+00:00"]