[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-evermind-raven-agent-harness":3,"topics-all":35,"news-related-de3de84d-1bd5-40fc-8226-765ee4841613":54},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":21,"news_slug":28,"published_at":29,"created_at":30,"modified_at":31,"is_published":32,"publish_type":33,"image_url":14,"view_count":34},"de3de84d-1bd5-40fc-8226-765ee4841613","284 赞登顶 HF 日榜:开源 Raven 让 Agent 框架自己进化","EverMind AI 开源多智能体系统 Raven，把「模型-框架」配对做成可组合的智能单元：Host Agent 拆解目标并调度专家智能体，EverOS 记忆跨会话沉淀经验，Evolver 自动改进框架本身。论文登上 Hugging Face 日榜第一（284 赞），GitHub 近 4.9k 星。","2026 年 9 月 27 日，EverMind AI 把一份题为 [Raven: The Harness of Harnesses for Composable Agentic Intelligence](https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.33439) 的论文放上 arXiv，三天后冲上 Hugging Face 论文日榜第一，拿下 284 个赞，配套 GitHub 仓库收获近 4.9k 星、1744 次提交。在 agent 框架多到数不过来的今天，这个热度说明它踩中了行业的某个真痛点。\n\n## 从「更强框架」到「框架的框架」\n\n过去两年，agent 圈的做法是针对一个领域手工打磨一套 harness：写好规划流程、配好工具、调好 prompt。论文指出这条路有两天花板——框架复杂度越来越高，人工设计难以扩展；框架和领域绑定太深，单一框架没有通用性。Raven 的思路是换问题：不再问「怎么给一个领域造更强的框架」，而是问「怎么自动构造专用框架、用经验持续改进、再跨领域编排它们」。它把每个可执行的「模型-框架」配对当作一个可组合的智能单元，让 Host Agent 负责拆解目标、匹配子任务给专家智能体、协调执行依赖、汇总结果，记忆系统 EverOS 和 Skill Forge 则把经验沉淀成可复用的技能。论文还给出了组合扩展可靠任务覆盖的充分条件——这类工作通常只出现在系统论文里，罕见地配了理论。\n\n## 四个内置专家和一个会改自己的工具\n\nRaven 内置四员大将：Raven-Research 做自主深度研究，Raven-Code 做智能体软件工程，Raven-Design 做视觉设计，Raven-Oncall 做无人值守的实验与监控自动化。官方 README 称四者在各自领域达到 SOTA 水平，并列出 SWE-bench Pro、PresentBench、DataAgentBench 等一排 benchmark 图。更有意思的是 Raven Evolver：一个独立工具把 Raven 当库调用，诊断失败、测试候选改进、在可复现评估中保留跑赢基线的变更——等于给「框架自身进化」修了一条正规流程。README 还展示了一个 RSI（递归自我改进）案例：在 nanochat 预训练实验里，Raven 独立跑完 7 轮共 172 次训练零崩溃，在同样 20 分钟单卡预算内把 val_bpb 降了 5.8%。另一个案例里它自主工作约 4 天、42 轮规划-开发-验证，交付了一款可玩的 Godot 4 射击游戏，连海报、演示文稿和网站都是它做的。可信度 caveat 也要说清：仓库自述 pre-alpha，接口可能快速变化，SOTA 声明目前主要是官方自报，等第三方复现再下重注。\n\n## 所以呢\n\nAgent 竞争的叙事正在从「谁的模型强」转向「谁的 harness 会进化」。Raven 的实验至少证明了一件事：让 AI 改进 AI 的流程本身，已经被开源社区跑通了闭环。下一次你评估 agent 框架时，除了看跑分，不妨多问一句——它的 harness，会自己变好吗？","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.33439","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18],{"id":12,"name":13,"slug":13,"description":14,"color":14},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":19,"name":20,"slug":20,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[22],{"id":23,"lang":24,"title":25,"summary":26,"content":27},"8125c3d5-d6f1-4e72-a8e0-2a9b25ef89a5","en","Raven: Open-Source Harness of Harnesses Tops HF Daily Papers","EverMind AI open-sourced Raven, a multi-agent system that auto-builds and evolves agent harnesses. The paper tops Hugging Face daily papers with 284 upvotes.","On September 27, 2026, EverMind AI published [Raven: The Harness of Harnesses for Composable Agentic Intelligence](https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.33439) on arXiv. Three days later it topped Hugging Face's daily paper ranking with 284 upvotes, while the companion GitHub repository climbed to nearly 4.9k stars and 1,744 commits. In a field crowded with agent frameworks, that reception suggests it hit a real pain point.\n\n## From Stronger Harnesses to a Harness of Harnesses\n\nThe dominant approach of the past two years has been hand-crafting one harness per domain: writing the planning loop, wiring tools, tuning prompts. The paper names two ceilings on this path — harness complexity keeps growing until manual design cannot scale, and tight domain coupling kills generality. Raven flips the question: instead of \"how to engineer a stronger harness for one domain,\" it asks \"how to autonomously construct specialized harnesses, improve them through experience, and orchestrate them across domains.\" Each executable model-harness pair becomes a composable unit of intelligence. A Host Agent decomposes goals, matches subtasks to specialized agents, coordinates execution dependencies, and integrates results, while EverOS memory and Skill Forge preserve experience as reusable skills. The paper also states sufficient conditions under which such composition expands reliable task coverage beyond individual agents — theory that system papers rarely bother to include.\n\n## Four Built-in Specialists and a Tool That Rewrites Itself\n\nRaven ships four built-in agents: Raven-Research for autonomous deep research, Raven-Code for agentic software development, Raven-Design for visual design, and Raven-Oncall for unattended experimentation and monitoring. The README claims SOTA performance across their domains, backed by benchmark charts on SWE-bench Pro, PresentBench, and DataAgentBench. The more striking piece is the Raven Evolver: a separate tool that consumes Raven as a library, diagnoses failures, tests candidate improvements, and retains only changes that beat the baseline in reproducible evaluations — a formal pipeline for harness self-evolution. The README also documents an RSI (recursive self-improvement) showcase: across nanochat pre-training experiments, Raven independently completed 172 training runs over 7 rounds without a single crash, cutting val_bpb by 5.8% within the same 20-minute single-GPU budget. In another showcase it worked autonomously for about 4 days through 42 rounds of planning, development, and verification to deliver a playable Godot 4 first-person shooter — poster, presentation deck, and website included. One caveat deserves emphasis: the repo calls itself pre-alpha, interfaces may change quickly, and the SOTA claims are vendor-reported so far — wait for third-party replication before betting heavily.\n\n## So What\n\nThe agent race is shifting from \"whose model is stronger\" to \"whose harness evolves.\" Raven's experiments demonstrate at least one thing: the loop of AI improving AI now works end-to-end in the open-source world. Next time you evaluate an agent framework, look past the benchmark scores and ask one more question — does its harness get better on its own?","evermind-raven-agent-harness","2026-09-30T17:20:00Z","2026-09-30T17:10:45.566475Z","2026-09-30T17:10:45.566491Z",true,"agent",147,[36,45],{"slug":37,"tag_slug":37,"title_zh":38,"title_en":39,"intro_zh":40,"intro_en":41,"id":42,"is_active":32,"created_at":43,"modified_at":44},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":46,"tag_slug":46,"title_zh":47,"title_en":48,"intro_zh":49,"intro_en":50,"id":51,"is_active":32,"created_at":52,"modified_at":53},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":55},[56,61,66,71,76,81],{"id":57,"title":58,"news_slug":59,"published_at":60},"7a2aa1aa-74ec-494a-a828-e6c1ad5b4cfe","Cloudflare Clef 决策模型开源,Jev 被压制","cloudflare-clef-open-source-decision-model-jev","2026-10-02T09:00:00+00:00",{"id":62,"title":63,"news_slug":64,"published_at":65},"9ebb888c-dfe7-416a-9940-a913527d4f73","AI Agent 的失败比成功更值钱:5 万对错误诊断数据,修正通过率 18.4%→51.1%","agent-error-dataset","2026-10-01T15:11:08+00:00",{"id":67,"title":68,"news_slug":69,"published_at":70},"73e29aee-b368-423f-be27-7653f65b4775","DeepSeek DSec 公开:300 万沙盒日撑 V4.1 训练","deepseek-dsec-v4-1-sandbox-rl-training","2026-09-28T00:00:00+00:00",{"id":72,"title":73,"news_slug":74,"published_at":75},"44740b4d-8c2c-44fc-8fff-fd89f3fb54ed","12 万美元 token 把 Copilot 运行时从 TypeScript 搬到 Rust","github-copilot-rust-migration-stephen-toub","2026-09-27T11:00:00+00:00",{"id":77,"title":78,"news_slug":79,"published_at":80},"e98cf9a2-8348-4142-9a56-c11166774798","SpeakerMem-R1:多方对话记忆,分清谁说了什么","speakermem-r1-multi-party-memory","2026-09-24T19:05:00+00:00",{"id":82,"title":83,"news_slug":84,"published_at":85},"68812025-96eb-4ca9-a1bc-8a82a40174dc","Google RRSI:给 Agent 外壳自进化加正则化","google-rrsi-agent-harness-regularization","2026-09-22T23:08:34+00:00"]