[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-ornith-1-5-benchmark-reality-check":3,"news-related-5b76d928-8896-4bf8-8edd-e195ecf0094a":38},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"5b76d928-8896-4bf8-8edd-e195ecf0094a","Ornith-1.5自报跑分赢了Claude,独立复测翻车了","Ornith-1.5 开源即冲上 HN 首页,官方称 397B 在 Terminal-Bench 2.1 达 86.1 微胜 Claude Opus 4.8;但独立复测 35B 档六项对比五项落后 Qwen3.8-27B,DeepSWE 差近一倍。自报跑分与社区复测的落差再成焦点。","8 月 20 日,开放权重模型家族 Ornith-1.5 上线 Hugging Face,很快冲上 Hacker News 首页,拿到 165+ 赞、58+ 评论。官方口径很响:397B 旗舰在 Terminal-Bench 2.1 拿下 86.1,微胜 Claude Opus 4.8 的 85.0,MIT 许可,397B\u002F35B\u002F9B 三档齐开。但几乎同一时间,社区的一组独立测试数字,给这份成绩单泼了冷水。\n\n## 自报成绩单:赢在毫厘之间\n\n按 [explainx 整理的官方表格](https:\u002F\u002Fwww.explainx.ai\u002Fblog\u002Fornith-1-5-self-improving-open-weight-model-august-2026),Ornith-1.5-397B 在 Terminal-Bench 2.1 报 86.1(Claude Opus 4.8 为 85.0),SWE-bench Verified 报 88.4(对手 87.6);但在 DeepSWE 上,56.0 反而低于 Opus 的 59.0。官方注明数字为五次运行平均,且做了防作弊处理——评测前剥离仓库 git 历史、解题期间断网。在 reward hacking 丑闻不断的 benchmark 环境里,这算加分项。\n\n技术故事本身也有新意。Ornith-1.5 把\"自我改进\"做成了训练期闭环:模型自己出题、自己搭脚手架、自己跑求解,三个阶段通过 GRPO 联合训练。出题侧的奖励由三部分构成——有效性是硬门槛,无效任务直接零分;难度瞄准约 0.2 的通过率,太高太低都不划算;新颖性作为辅助信号,防止生成器刷重复题。这套机制针对的正是\"固定人类题库\"的天花板。\n\n## 独立测试讲了另一个故事\n\n问题出在 35B 档。官方表格里,Ornith-1.5-35B 的 Terminal-Bench 2.1 报 74.8,把 Qwen3.6-35B-A3B 的 52.5 甩出二十多分。但 explainx 报道,一位 HN 用户自己跑了 Ornith-1.5-35B 对比 Qwen3.8-27B,结果方向完全反过来:Terminal-Bench 2.1 上 67.8 对 73.0 落后,SWE-bench Pro 59.6 对 61.7,HLE 25.6 对 30.8;差距最大的是 DeepSWE,22.0 对 42.2,几乎差出一倍。六项对比里,Ornith 只在 NL2Repo 一项反超,GPQA Diamond 打平。\n\n当然要说明:这是单人非正式复测,且拿 27B 稠密模型对比 35B MoE,架构并不对等。但 DeepSWE 恰恰考察长周期软件工程能力,是最难突击应付的一类 benchmark,这个维度的巨大落差,很难全用\"测试环境差异\"解释。\n\n## 自报的永远是自报\n\n社区还挖出两个细节。其一,有用户在身份探测中发现该模型\"稳定自称 Claude\",这通常指向训练语料被 Claude 生成文本污染,而不是权重抄袭;其二,HN 考证认为 397B 底座是 Qwen3.5-397B-A17B 后训练而来,团队背后是 Jiwei Li——两条都未经官方确认,姑且听之。\n\n平心而论,Ornith 的自我改进训练法值得研究,量化版 9B-Mobile 宣称可部署到 iPhone 和 Android,也是边缘侧的实用方向。但这次事件再次说明:厂商表格与独立复测之间的落差,是开源模型圈的结构性问题,五次平均的自报跑分也改变不了这一点。下次再看到\"开源模型击败 Claude\"的标题,先问一句——谁跑的分,跑了几次,在谁的环境里跑的。","https:\u002F\u002Fwww.explainx.ai\u002Fblog\u002Fornith-1-5-self-improving-open-weight-model-august-2026","b093fba2-b9ec-492e-acd6-ff6018ef4cfb",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"e82b2d09-81b2-43d1-977e-e018443b3c14","coding-agent",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"734229b8-b0d6-4fa9-b1eb-cb3a02c36187","en","Ornith-1.5 beat Claude in self-reported benchmarks — then an independent re-run disagreed","Ornith-1.5 hit HN front page on release with 86.1 on Terminal-Bench 2.1, edging Claude Opus 4.8. But an independent community run of the 35B model lost to Qwen3.8-27B on five of six benchmarks, with DeepSWE nearly halved.","On August 20, the open-weight model family Ornith-1.5 landed on Hugging Face and quickly climbed to the Hacker News front page with 165+ points and 58+ comments. The official pitch is loud: the 397B flagship scores 86.1 on Terminal-Bench 2.1, narrowly beating Claude Opus 4.8 at 85.0, MIT-licensed, shipped in 397B\u002F35B\u002F9B tiers. But almost at the same time, a set of independent community test numbers poured cold water on that report card.\n\n## The Self-Reported Scorecard: Winning by a Hair\n\nAccording to [the official tables compiled by explainx](https:\u002F\u002Fwww.explainx.ai\u002Fblog\u002Fornith-1-5-self-improving-open-weight-model-august-2026), Ornith-1.5-397B reports 86.1 on Terminal-Bench 2.1 (Claude Opus 4.8 at 85.0) and 88.4 on SWE-bench Verified (versus 87.6); on DeepSWE, however, its 56.0 actually trails Opus at 59.0. The team notes the figures are averaged over five runs, with anti-cheating measures in place — stripping git history from repos before evaluation and disabling network access during solving. In a benchmark scene plagued by reward-hacking scandals, that counts as a plus.\n\nThe technical story has genuine novelty too. Ornith-1.5 turns \"self-improvement\" into a training-time closed loop: the model generates its own tasks, builds its own scaffolds, and runs its own solution rollouts, with all three stages jointly trained via GRPO. The task-generation reward has three components — validity acts as a hard gate, with invalid tasks scoring zero no matter how hard they look; difficulty targets an empirical success rate around 0.2, where most attempts fail but rare successes still teach; and novelty serves as a secondary signal against regenerating near-duplicate busywork. The mechanism aims straight at the ceiling of fixed, human-curated task sets.\n\n## The Independent Run Tells a Different Story\n\nThe problem sits in the 35B tier. In the official tables, Ornith-1.5-35B reports 74.8 on Terminal-Bench 2.1, leaving Qwen3.6-35B-A3B at 52.5 more than twenty points behind. But as explainx reported, one HN user ran their own comparison of Ornith-1.5-35B against Qwen3.8-27B, and the direction flipped completely: on Terminal-Bench 2.1 it trailed 67.8 to 73.0, on SWE-bench Pro 59.6 to 61.7, on HLE 25.6 to 30.8; the largest gap was DeepSWE, 22.0 versus 42.2, nearly a two-fold difference. Across six comparisons, Ornith only overtook on NL2Repo, with GPQA Diamond tied.\n\nCaveats apply: this was a single person's informal re-run, pitting a 27B dense model against a 35B MoE — not an apples-to-apples architecture match. But DeepSWE specifically stresses long-horizon software engineering, the hardest kind of benchmark to cram for, and a gap this large in that dimension is hard to explain away entirely by test-environment differences.\n\n## Self-Reported Is Self-Reported\n\nThe community dug up two more details. First, some users found in identity probes that the model \"reliably claims to be Claude\" — usually a sign of training-data contamination from Claude-generated text rather than copied weights. Second, HN research suggests the 397B is derived from Qwen3.5-397B-A17B via post-training, with Jiwei Li behind the team — both unconfirmed by the official announcement, so take them with a grain of salt.\n\nTo be fair, Ornith's self-improvement training method deserves study, and the quantized 9B-Mobile build, claimed to deploy on iPhone and Android, is a practical direction for the edge. But this episode repeats an old lesson: the gap between vendor tables and independent reproductions is structural in the open-model scene, and five-run averages do not change it. Next time you see a headline saying \"open model beats Claude,\" ask first — who ran the benchmark, how many times, and in whose environment.","ornith-1-5-benchmark-reality-check","2026-08-20T17:10:00Z","2026-08-20T17:08:02.366986Z","2026-08-20T17:08:02.366995Z",true,"agent",149,{"items":39},[40,45,50,55,60,65],{"id":41,"title":42,"news_slug":43,"published_at":44},"7ba15299-8bee-4039-8bc7-dbb58754b562","SWE-bench Science:最强 Claude Code 修科学代码也不及格,四类失败模式被拆解","swe-bench-science-benchmark","2026-08-21T13:00:00+00:00",{"id":46,"title":47,"news_slug":48,"published_at":49},"c4acfb32-44d4-45a9-ac27-ec9fff4e1eca","Mistral 开源 Leanstral 1.5:6B 激活参数刷新形式化推理 SOTA","mistral-leanstral-1-5","2026-07-04T00:30:00+00:00",{"id":51,"title":52,"news_slug":53,"published_at":54},"ff0bc92a-295a-4707-be8d-76115fe9eeee","PerceptionBench 出炉:16 个前沿多模态模型,视觉感知无一及格","moonshot-perceptionbench-atomic-perception","2026-08-26T13:15:00+00:00",{"id":56,"title":57,"news_slug":58,"published_at":59},"f9cf9f03-6aca-4d29-94d3-5c6acfeaf435","匿名模型 OX Alpha 短暂登顶 OpenRouter 编码榜:研究者推测底座指向智谱 GLM-5.x","ox-alpha-stealth-openrouter-glm-5-zhipu","2026-08-24T03:00:00+00:00",{"id":61,"title":62,"news_slug":63,"published_at":64},"5a274662-0c3f-492c-a0e8-a46c5a0be783","别再自己给自己打分了:Co-RL 让模型互相判卷,无标签 RL 追平有监督","co-rl-peer-reward-label-free-rl","2026-08-20T19:10:00+00:00",{"id":66,"title":67,"news_slug":68,"published_at":69},"89a79f9a-bfd2-4ebe-8f03-92fa74a3a34f","Ornith-1.5 开源：模型自己出题、自己搭考场，397B 到 9B 三档齐发","ornith-1-5-self-improvement-open-models","2026-08-20T13:30:00+00:00"]