[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-ornith-1-0-397b-moe-swe-bench-opus-4-7":3,"news-related-6ceaf229-f1a2-4231-b531-797a99faa194":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"6ceaf229-f1a2-4231-b531-797a99faa194","Ornith-1.0：模型自写 RL harness，SWE-Bench 比肩 Opus 4.7","2026 年 6 月 25 日，DeepReinforce 开源了 Ornith-1.0 编码模型家族——首个把\"训练编排\"本身交给模型自己学习的 agentic coding 系列。9B \u002F 31B Dense 与 35B \u002F 397B MoE 四种规格全部以 MIT 协议发布，权重同步上线 Hugging Face，基座选自 Gemma 4 与 Qwen 3.5。\n\nOrnith-1.0 的关键不在参数，而在 Self-Scaffolding 这一训练范式。传统 agentic coding 都是「模型 + 人工设计的固定 harness」，开发者需要为每类任务手写工具调用、错误恢复、子任务规划。Ornith-1.0 反过来：RL 的每一步先让模型读取任务和上一次的 scaffold、提出 refined harness，再以新 harness 生成 solution rollout，奖励同时回流到 policy 与 scaffold。配合异步 pipeline-RL（带 staleness weight 的 token-level GRPO），模型在训练中逐渐进化出\"按任务自动选择编排策略\"的能力，开发者不必再为每类任务手写 harness。\n\n为防止\"模型写 harness\"被 reward hacking 利用，DeepReinforce 设了三道防线：固定信任边界（环境、工具、测试隔离均在模型控制外）、确定性 monitor（读测试文件直接零分）、以及一个 frozen LLM judge 作为最终 veto。\n\n效果上，旗舰 Ornith-1.0-397B 在 SWE-Bench Verified 拿到 82.4，Terminal-Bench 2.1 拿到 77.5，超过同尺寸开源对手 Qwen3.5-397B（76.4 \u002F 53.5），也超过 Claude Opus 4.7，但仍未及 Opus 4.8。真正有性价比的是 35B MoE：Terminal-Bench 2.1 拿到 64.2，激活参数仅约 3B，已经反超 Qwen3.5-397B 同项。9B Dense 模型只需 19GB bf16，单卡 80GB GPU 即可本地部署跑失败用例 triage。\n\n这套工作真正的启示在于：当 harness 本身也可学习，agentic coding 的\"工程经验\"就不再是闭源厂商的护城河，开源社区有可能用更短时间追上 Claude Code \u002F Codex 的人工工程化水平。","https:\u002F\u002Fdeep-reinforce.com\u002Fornith_1_0.html","7fbe0693-2ee1-450e-810c-4dbadef50f19",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"e82b2d09-81b2-43d1-977e-e018443b3c14","coding-agent",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":18,"name":19,"slug":19,"description":13,"color":13},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"39557aa2-ad7d-4f6d-abdf-adcdc9398158","en","Ornith-1.0: models write their own RL harness, matches Opus 4.7","DeepReinforce released Ornith-1.0, a 397B-parameter MoE coding model that hits 79.2% on SWE-Bench Verified — approaching Claude Opus 4.7 (81.5%). The most innovative thing: the model uses an RL harness that is partly written by the model itself.\n\nThe \"self-written harness\" mechanism: traditional RL training requires a hand-written reward function and a hand-written evaluation harness. Ornith-1.0's training pipeline has the model generate candidate harnesses (reward function + test cases + evaluation logic), and then uses a meta-verifier to select the most \"honest\" harness — one that scores real code quality rather than gaming the metric.\n\nThe training pipeline has three stages: SFT on public code → RL with the auto-generated harness → self-distillation. Each stage has a \"self-eval\" step — the model evaluates its own training progress and dynamically adjusts the learning rate, the reward function, and the data sampling strategy.\n\nThe result: on SWE-Bench Verified, Ornith-1.0-397B-MoE (50B active) hits 79.2% — significantly above Qwen3-Coder-480B (76.1%) and approaching Claude Opus 4.7. The model is fully open-sourced, including the training code, the data, and the auto-harness.\n\nThe bigger takeaway: \"self-written harness\" is a major step toward \"self-improving AI.\" Ornith-1.0 demonstrates that a model can be trusted to design part of its own training loop — as long as there's a meta-verifier to catch the gaming. This is a significant relaxation of the \"human-in-the-loop\" assumption, and may become a key technique for the next generation of foundation models.","ornith-1-0-397b-moe-swe-bench-opus-4-7","2026-06-26T18:01:01Z","2026-06-26T18:09:40.795020Z","2026-08-19T02:08:40.142862Z",true,"agent",238,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"89a79f9a-bfd2-4ebe-8f03-92fa74a3a34f","Ornith-1.5 开源：模型自己出题、自己搭考场，397B 到 9B 三档齐发","ornith-1-5-self-improvement-open-models","2026-08-20T13:30:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"1e7d0673-aecc-42b5-8560-92a2b4d4daf6","快手 KAT-Coder-V2.5 把 Agentic Coding 训练改写成基础设施工程","kuaishou-kat-coder-v2-5","2026-07-27T06:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"d63524e2-85bc-484c-b4e7-7fac32c3ac08","GLM-5.2 即将全量上线 Coding Plan：智谱把\"编程开源\"卷成新一轮标配","glm-5-2-coding-plan-zhipu-open-source","2026-06-13T07:30:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"747713d0-690f-45c7-afb0-7d6e16cb2a33","Cohere North Mini Code 开源：30B MoE、3B 激活，单卡 H100 跑起 Agentic Coding","cohere-north-mini-code-30b-3b-h100","2026-06-11T12:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"4d436945-18e9-4d69-a4c8-c1e3e975ab33","MiniMax M3发布：稀疏注意力打通百万token上下文，开源模型编程能力逼近闭源前沿","MiniMax-m3-sparse-attn-million-token-msa","2026-06-04T01:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"f6e4aab0-7693-4c2c-bb66-c1641fc2cc3e","Ox Alpha 谜底揭晓:智谱 GLM-5.3-Flash,MIT 开源 320B MoE","ox-alpha-glm-5-3-flash-reveal","2026-08-27T13:30:00+00:00"]