[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-ox-alpha-stealth-openrouter-glm-5-zhipu":3,"news-related-f9cf9f03-6aca-4d29-94d3-5c6acfeaf435":41},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":27,"news_slug":34,"published_at":35,"created_at":36,"modified_at":37,"is_published":38,"publish_type":39,"image_url":14,"view_count":40},"f9cf9f03-6aca-4d29-94d3-5c6acfeaf435","匿名模型 OX Alpha 短暂登顶 OpenRouter 编码榜:研究者推测底座指向智谱 GLM-5.x","8月20日,OpenRouter 上线代号 stealth\u002Fox-alpha 的匿名大模型,1M 上下文、支持文本\u002F图像\u002F视频输入、预览期免费。第三方 DeepSWE Pass@1 数据约 80%,研究者据指纹推测其底座指向智谱 GLM-5.x,智谱未公开确认。","匿名大模型「stealth\u002Fox-alpha」(下文简称 OX Alpha)在 8 月 20 日登上 OpenRouter,没有发布方,没有品牌,只有一份极简的能力清单:1,048,576 token 的上下文窗口、131K 的最大输出、文本\u002F图像\u002F视频三种输入模态,预览期免费,定价待定。这件事之所以迅速在开发者社区里传开,不是因为它「免费」,而是因为它在 DeepSWE 这类编码基准上跑出的初版数据直接站到了前沿梯队。[^1]\n\n## 一、一个 10 任务小样本,撑起了一个「80% Pass@1」\n\n最先把数据传出来的是开发者 Ben Davis。他在 DeepSWE 上对 OX Alpha 跑了 10 道题,8 道通过,折算成 Pass@1 约 80%。同口径下,Claude 大约 65%,GPT-5.6 Sol 大约 52%。\n\n但这两个对比数值得同时放在显微镜下看:第一,DeepSWE 完整的基准是 113 题,OX Alpha 的 80% 是 10 题子集的折算,样本量极小,方差很大,DeepSWE 的官方 BenchSift 榜单截至 8 月 21 日并没有收录 OX Alpha 的提交。第二,「Claude 65% \u002F GPT-5.6 Sol 52%」这两个数字同样来自社区在子集上的运行,不是各厂商对外公布的官方数字,直接横比存在偏差。\n\n所以诚实一点的读法是:OX Alpha 在编码基准上处在前沿梯队,但「它稳压 GPT-5.6」这个结论目前没有跨榜单、跨样本的审计支撑。把 80% 当成「领先指标」可以,当成「SOTA 证明」则不行。[^2][^3]\n\n## 二、谁来训练它?一个独立的指纹比对指向智谱\n\n发布方身份始终是匿名状态。OpenRouter 在 provider 字段里把它标为「Stealth」,OpenCode 直接称之为「the stealth model」。智谱没有公开确认,也没有公开否认。\n\n围绕「它是哪家」的猜测在过去几天快速收敛。研究者 Ben Davis 把他对 OX Alpha 的分析汇总后表示,他有 99% 的把握认为这是智谱尚未发布的 GLM-5.x 系列旗舰,理由是两条独立的指纹证据:第一,OX Alpha 在处理视频输入时消耗的 token 模式与 GLM-5V-Turbo 完全一致;第二,OX Alpha 的分词器在 25 个跨语种 prompt 上与 GLM-5.3 高度对齐,只存在细微的 vocabulary 微调差异。这种「服务端指纹 + 分词器指纹」双重命中,在模型家族内部成员之间的可信度比较高,智谱过往也确实有过以匿名形式做小规模公开测试的先例。[^4]\n\n独立分析把架构大致反推为:总参数约 744B,激活参数约 40B 的 MoE。这个规模落在「旗舰 MoE 但还不是超大规模稀疏」的区间,与其「GLM-5.x 后继者」的身位比较吻合。\n\n## 三、为什么「匿名上架」这件事本身值得讨论\n\n把前沿模型以匿名方式先放上推理路由平台,先让真实用户在不知情的情况下测几天,再公开身份——这种「先测试,后品牌」的发布节奏,在 2026 年的中国厂商里正在变成一种新的标准动作。它有几个明显的副作用。\n\n第一,基准污染被部分回避。开发者面对一个没有品牌、没有厂商口径的模型时,更倾向按能力而非「这是谁家」的预期去跑数据,得到的反馈噪声更小。第二,品牌风险被前置消化。如果模型在公开测试中暴露严重问题,厂商可以选择「不发」,代价只是损失一周的推理预算;如果表现真的超出预期,公开身份的节点就是一个「我已经在真实用户手里验证过」的强叙事。第三,价格被压住。当一个免费可用的前沿 MoE 摆在 OpenRouter 上,周边所有商用前沿 API 都要回答同一个问题:「你凭什么收这个钱?」\n\n但这种节奏也有边界。OX Alpha 的预览窗口按照 OpenRouter 公开节奏会持续到 8 月 27 日前后,智谱原计划在 8 月 28 日前后公开发布 GLM-5.3 的开放权重。如果 OX Alpha 真的就是 GLM-5 系列的后继者,那么接下来的问题是:它会以什么身份、在什么价位、与开放权重的 GLM-5.3 形成什么样的产品分层。\n\n## 四、然后呢\n\n匿名模型的爆发期,本质上是大模型「前沿即货架」的进一步深化:研究团队不再等到品牌发布会才让真实用户接触模型,而是把 API 当成 PR 通道的反面——一个没有 PR 的能力测试台。\n\n对开发者来说,真正的可操作动作有三件:一是趁预览窗口免费,把 OX Alpha 接入你现有的编码 Agent harness,跑一遍你内部代表性的多文件改动任务;二是对比它在「短任务一次性成功」和「长任务一致性」两个维度上的具体表现,因为前者是当下子集基准的优势区,后者才是工业编码最稀缺的资源;三是观察智谱在 8 月 28 日前后对 GLM-5.3 开放权重的真实姿态——身份、价格、许可条款三个变量任意一个落点不同,OX Alpha 的故事就是「有趣的测试」,还是「行业拐点」。\n\n[^1]: Build Fast with AI,《Mystery Model OX Alpha Beats GPT-5.6: AI News Aug 22-23 2026》,2026-08-22。https:\u002F\u002Fwww.buildfastwithai.com\u002Fblogs\u002Fai-news-today-august-22-23-2026\n[^2]: thecherrycreeknews,《Ox Alpha: Anonymous 1M-Context Model Hits No. 2 on OpenCode in Three Days》,2026-08-21。https:\u002F\u002Fthecherrycreeknews.com\u002Fox-alpha-stealth-model-openrouter-benchmarks-analysis-cherry_creek\u002F\n[^3]: startupfortune,《Ox Alpha Topped Coding Benchmarks and Forensics Now Point to Zhipu》,2026-08-22。https:\u002F\u002Fstartupfortune.com\u002Fox-alpha-topped-coding-benchmarks-and-forensics-now-point-to-zhipu\u002F\n[^4]: Local AI Zone,《Ox Alpha Stealth Model Comprehensive Analysis》,2026-08-21。https:\u002F\u002Flocal-ai-zone.github.io\u002Fblog\u002Fox-alpha-stealth-model-comprehensive-analysis.html","https:\u002F\u002Fwww.buildfastwithai.com\u002Fblogs\u002Fai-news-today-august-22-23-2026","5d6d4f6b-780a-419e-8057-20a8c1bbaa60",[11,15,18,21,24],{"id":12,"name":13,"slug":13,"description":14,"color":14},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":25,"name":26,"slug":26,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[28],{"id":29,"lang":30,"title":31,"summary":32,"content":33},"b63178c3-6ca3-4e09-b6fa-1df33b769a25","en","A Mystery Model Called OX Alpha Briefly Topped OpenRouter Coding Leaderboards — Forensics Suggest a Zhipu GLM-5.x Backbone","On August 20, 2026 OpenRouter listed an unnamed model under the handle stealth\u002Fox-alpha: 1M-token context, text\u002Fimage\u002Fvideo inputs, free during preview. Third-party DeepSWE Pass@1 of about 80% made it briefly lead coding leaderboards. A researcher says fingerprint evidence points to Zhipu's GLM-5.x line, though Zhipu has not confirmed.","An unnamed large language model, listed simply as \"stealth\u002Fox-alpha\" on OpenRouter (hereafter OX Alpha), appeared on August 20, 2026 without a publisher, a brand, or a press release. The only public information was a compact capability sheet: a 1,048,576-token context window, a 131K maximum output, three input modalities (text, image, and video), and free access during the preview window, with pricing to be determined. The reason it spread through developer communities within hours was not the price tag. It was that early runs on coding agent benchmarks like DeepSWE landed directly in the frontier band.[^1]\n\n## 1. A 10-task Sample Underwrites a \"80% Pass@1\" Headline\n\nThe first numbers out the door came from developer Ben Davis. He ran OX Alpha through 10 tasks on DeepSWE, watched it pass 8, and reported a Pass@1 of roughly 80%. On the same subset, he put Claude at around 65% and GPT-5.6 Sol at around 52%.\n\nBoth comparison numbers need to be read with the same microscope. First, the full DeepSWE benchmark is 113 tasks. The 80% attributed to OX Alpha is extrapolated from a 10-task subset, which carries a very wide variance band. As of August 21, OX Alpha does not appear on the official BenchSift leaderboard for DeepSWE. Second, the \"Claude 65% \u002F GPT-5.6 Sol 52%\" figures are also community-run on a subset and are not the official numbers each vendor publishes for those models. Cross-vendor comparison on those terms is biased.\n\nThe honest read is this: OX Alpha sits inside the frontier band on coding, but the claim that it \"definitively beats GPT-5.6\" is not yet supported by a cross-leaderboard, cross-sample audit. Treating 80% as a leading indicator is fair. Treating it as a SOTA proof is not.[^2][^3]\n\n## 2. Who Trained It? A Standalone Fingerprint Match Points to Zhipu\n\nThe publisher's identity remains anonymous. OpenRouter labels it \"Stealth\" in the provider field. OpenCode calls it \"the stealth model.\" Zhipu has neither confirmed nor denied ownership on the record.\n\nThe speculation around the source has converged quickly over the past few days. Independent researcher Ben Davis, who has driven most of the public analysis on OX Alpha, has stated that he is 99% confident the model is an unreleased GLM-5.x flagship from Zhipu. The case rests on two independent fingerprint matches. First, OX Alpha's token consumption pattern on video input matches GLM-5V-Turbo exactly. Second, OX Alpha's tokenizer aligns closely with GLM-5.3 across 25 cross-lingual prompts, with only minor vocabulary-level deltas. A dual-stack fingerprint hit between members of the same model family carries nontrivial confidence, and Zhipu has previously run small-scale public tests under anonymous branding.\n\nIndependent analysis reverse-engineers the architecture at roughly 744B total parameters with about 40B active in a Mixture-of-Experts configuration. That places OX Alpha in the \"flagship MoE, but not yet ultra-sparse\" band, which matches its likely position as a GLM-5.x successor.[^4]\n\n## 3. Why the \"Anonymous Listing\" Itself Is Worth Discussing\n\nPutting a frontier model on a routing platform first, letting real users test it blind for a few days, and only then announcing a brand identity is becoming a standard release cadence for some Chinese labs in 2026. The side effects are concrete.\n\nFirst, benchmark contamination is partially avoided. When developers do not know which vendor stands behind a model, they tend to run it on capability rather than on their expectation of \"who made this,\" and the resulting feedback carries less noise. Second, brand risk is front-loaded. If the model fails badly in public testing, the lab can simply not announce it; the cost is one week of inference budget. If the model outperforms expectations, the identity reveal becomes a strong \"I have already been validated in real users' hands\" narrative. Third, pricing is forced downward. Once a free, frontier-tier MoE sits on OpenRouter, every commercial frontier API on the shelf has to answer the same question: what justifies your price point?\n\nThat cadence also has edges. OX Alpha's preview window on OpenRouter is scheduled to run through around August 27. Zhipu originally targeted August 28 for the public release of GLM-5.3 open weights. If OX Alpha really is the GLM-5 series successor, the next question is what identity, what price, and what product tier it will sit in next to the openly-licensed GLM-5.3.\n\n## 4. What Now\n\nThe rise of anonymous frontier launches is, at heart, the further deepening of \"frontier equals shelf.\" Research teams no longer wait until a brand launch event for real users to touch the model. They treat the API as the inverse of a PR channel: a capability test rig without the PR.\n\nFor developers, there are three concrete actions worth taking. First, while the preview window remains free, route OX Alpha into your existing coding agent harness and run a representative suite of multi-file edits on a codebase you actually maintain. Second, compare OX Alpha on two axes, short-task single-shot success and long-task coherence, because the former is where small-subset benchmarks currently favor it, and the latter is the scarcer resource in industrial coding work. Third, watch what Zhipu actually does around August 28 with GLM-5.3's open weights. Any of three variables, identity, price, or license terms, landing differently turns the OX Alpha story from an interesting test into an industry turning point.\n\n## References\n\n[^1]: Build Fast with AI, \"Mystery Model OX Alpha Beats GPT-5.6: AI News Aug 22-23 2026\", 2026-08-22. https:\u002F\u002Fwww.buildfastwithai.com\u002Fblogs\u002Fai-news-today-august-22-23-2026\n[^2]: thecherrycreeknews, \"Ox Alpha: Anonymous 1M-Context Model Hits No. 2 on OpenCode in Three Days\", 2026-08-21. https:\u002F\u002Fthecherrycreeknews.com\u002Fox-alpha-stealth-model-openrouter-benchmarks-analysis-cherry_creek\u002F\n[^3]: startupfortune, \"Ox Alpha Topped Coding Benchmarks and Forensics Now Point to Zhipu\", 2026-08-22. https:\u002F\u002Fstartupfortune.com\u002Fox-alpha-topped-coding-benchmarks-and-forensics-now-point-to-zhipu\u002F\n[^4]: Local AI Zone, \"Ox Alpha Stealth Model Comprehensive Analysis\", 2026-08-21. https:\u002F\u002Flocal-ai-zone.github.io\u002Fblog\u002Fox-alpha-stealth-model-comprehensive-analysis.html","ox-alpha-stealth-openrouter-glm-5-zhipu","2026-08-24T03:00:00Z","2026-08-24T03:06:00.139522Z","2026-08-24T03:06:00.139533Z",true,"agent",89,{"items":42},[43,48,53,58,63,68],{"id":44,"title":45,"news_slug":46,"published_at":47},"d4fa7e14-8fbd-4940-93a6-3dd6f0a3991d","DeepSeek V4 Pro 正式版：1.6T MoE，1M 上下文","deepseek-v4-pro-0813-ga-1m-context-moe","2026-08-13T02:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"c94766df-827e-4e4e-a006-b6639ec76722","DeepSeek V4-Flash-0731 转正观察:权重不动,后训练把 Agent 分数打到 V4-Pro 之上","deepseek-v4-flash-0731-agent-benchmark-official-aug2026","2026-08-01T02:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"80315de0-7eb3-491a-b2e6-103a691a8bd7","Nanbeige4.2-3B 用 Looped Transformer 在 11 项基准上跑赢 Qwen3.5-9B","nanbeige-4-2-3b-looped-transformer-agentic-3b-beats-qwen3-5-9b","2026-07-30T10:30:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"38fe9093-827f-43d1-8350-7cdd391cf1e3","北大 DataPrep-Bench 把 LLM 当数据准备工来打分：DAS 评估器把「训练价值」算成分布距离","pku-dataprep-bench-das","2026-07-27T22:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"5bfdf32b-44eb-4eb5-a98b-39e921168182","九天内连发五款前沿模型:7 月的大模型军备赛,真正决胜负的不再是 benchmark","july-2026-five-frontier-models","2026-07-23T12:00:00+00:00",{"id":69,"title":70,"news_slug":71,"published_at":72},"461816c1-61b7-4077-b482-428379c046de","AI2 EMO：把 MoE 训练成「可拆装」模块，1B 激活也能按域调度","ai2-emo-emergent-modular-moe-pluggable","2026-05-08T00:00:00+00:00"]