[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-nvidia-sol-pi-harness-token-compression":3,"topics-all":44,"news-related-c696208b-6535-4eb9-b1ed-2e4f835d2f88":63},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":30,"news_slug":37,"published_at":38,"created_at":39,"modified_at":40,"is_published":41,"publish_type":42,"image_url":14,"view_count":43},"c696208b-6535-4eb9-b1ed-2e4f835d2f88","NVIDIA SoL-Pi 把 coding agent 的 token 砍掉 44%,harness 开始变天","NVIDIA Labs 发布 SoL-Pi,用 RSI 思路在 harness 层规模化自动研究,51 任务 EdgeBench 上让 GPT-5.6 Sol 与 Opus 5 的 token 流量降 44.7-49.0%,API 成本省 1\u002F3,GitHub 已开源 2.26k stars。","上周 NVIDIA Labs 把\"harness 自动发现\"这件事往前推了一大步:一组叫 SoL-Pi 的论文,作者里出现了 Enze Xie 和 Song Han,GitHub 已经攒到 2.26k stars。这套东西在 51 个任务的 EdgeBench 上跑下来,让 GPT-5.6 Sol 和 Opus 5 在不换底层模型的前提下,token 流量砍掉 44.7%-49.0%,API 成本直接省 1\u002F3。\n\n这到底是个什么东西?它不是新模型,是 coding agent 的\"运行环境优化器\"。论文的核心思想很直接:当 coding agent 从\"监督式补全代码\"走向\"无监督全天候自我探索\"之后,影响成本和稳定性的瓶颈不再是模型本身,而是 harness——也就是把 agent 串起来跑的那套脚手架(动作执行、上下文压缩、观察处理、委托读取这些机制)。SoL-Pi 用 RSI(Recursive Self-Improvement)的思路,在 harness 层做规模化自动研究:把尽可能多的不同环境接入循环,让 harness 自己通过多环境 rollout 学出可迁移的优化。\n\n关键在于结果有数字可核。论文里给出的 EdgeBench 51 任务跑分,SoL-Pi 在 GPT-5.6 Sol 和 Opus 5 两个模型上的性能与 Pi 持平,但同时 token 流量降了 44.7%-49.0%。换算成 API 成本,每小时估算相对原生 Codex 和 Claude Code harness 能省 $8.75-13.50,相对 Pi 也能省 $4.36-5.71。这意味着如果你现在用 Claude Code 或 Codex 跑 24 小时无人值守 agent,SoL-Pi 是能直接落地的\"省 token 套件\",而不是又一个 benchmark 跑分漂亮但生产用不上的玩具。\n\n最终从大量候选机制里被保留下来的有四个:action execution(动作执行)、context compaction(上下文压缩)、observation handling(观察处理)、delegated reading(委托读取)。这四个机制的共同点是——它们都直接打在 token 流量的主路径上。上下文压缩就不用说了,LLM agent 长跑最贵的就是这一段;委托读取是个不那么显眼但很关键的优化,把读文件\u002F读网页这种 I\u002FO 密集型操作从主上下文里挪出去,避免观察值把 prompt 撑爆;动作执行和观察处理则属于\"每次循环都在跑\"的微优化,单个省一点,堆起来就成了一半。\n\n但也别把这想成万能解药。SoL-Pi 跑分是基于 51 个任务的 EdgeBench,这是 NVIDIA 自家评测集,模型只覆盖 GPT-5.6 Sol 和 Opus 5——这套优化在开源小模型(比如 Qwen3-Coder、GLM-4.6、Kimi K2.8)上能不能同样省 44% token,目前没有公开数据。换句话说,RSI 在 harness 层的天花板被 NVIDIA 这篇论文抬高了,但是否对所有模型族都适用,还要看接下来 NVlabs 的 follow-up 或者社区的复现。\n\n## 为什么这件事不只是\"NVIDIA 又发了一篇 paper\"\n\n另一个被低估的角度是 \"harness 即研究对象\"。在 SoL-Pi 之前,大家的目光都集中在模型权重和数据上;NVIDIA 这篇其实在悄悄拆解一个隐含事实——agent 系统的总成本里,harness 占比已经大到值得单独发论文的程度。Librarian bot 推送的相关论文里,openJiuwen、StarHarness、HarnessDev、Prime Agent、RSIAgent、PILOT、Hierarchical Self-Improvement 这 7 篇都是同主题,说明这个赛道已经在同时涌进好几家研究组。harness 优化的\"model engineering\"和\"scaffold engineering\"开始分家,这件事对做 agent infra 的同行意义很大。\n\nGitHub 已经开源:https:\u002F\u002Fgithub.com\u002FNVlabs\u002FSoL-Pi,项目页 https:\u002F\u002Fnvlabs.github.io\u002FSoL-Pi\u002F。今天就能 clone 下来接你的现有 coding agent,先在 5-10 个真实任务上跑一遍,看 token 表是不是真的砍了一半。如果你的 agent 现在每个任务要烧几美金的 API 费,这一刀能直接砍回一块钱以内。\n\n——下次再有人跟你说\"agent 太烧 token\",先别急着换小模型;把 harness 翻一遍,可能比换模型省得更多。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.20519","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21,24,27],{"id":12,"name":13,"slug":13,"description":14,"color":14},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":19,"name":20,"slug":20,"description":14,"color":14},"e82b2d09-81b2-43d1-977e-e018443b3c14","coding-agent",{"id":22,"name":23,"slug":23,"description":14,"color":14},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":25,"name":26,"slug":26,"description":14,"color":14},"8dac812d-3839-4abe-a855-5f56ec9515fd","nvidia",{"id":28,"name":29,"slug":29,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[31],{"id":32,"lang":33,"title":34,"summary":35,"content":36},"1aae2008-3b6a-4b7f-a486-8a2cb1e84bab","en","NVIDIA's SoL-Pi Cuts Coding Agent Tokens by 44%—Harness Engineering Is Now Its Own Field","NVIDIA Labs released SoL-Pi, an RSI-based harness auto-research framework. On the 51-task EdgeBench it delivered GPT-5.6 Sol and Opus 5 performance comparable to Pi while cutting token traffic 44.7%-49.0% and API cost by roughly one third.","NVIDIA Labs released SoL-Pi, a harness-level RSI (Recursive Self-Improvement) framework that runs auto-research loops across many diverse environments. On the 51-task EdgeBench benchmark, SoL-Pi delivered GPT-5.6 Sol and Opus 5 performance comparable to the Pi harness baseline while cutting recorded token traffic by 44.7%-49.0% and reducing API cost by roughly one third. GitHub stars reached 2.26k.\n\nThe four mechanisms that survived the auto-selection loop are: action execution, context compaction, observation handling, and delegated reading. They all sit on the hot path of token traffic, which is why the cumulative savings land close to half.\n\n## Why this is more than another NVIDIA paper\n\nThe paper also signals a quiet split between model engineering and scaffold engineering. The seven related works surfaced by Semantic Scholar (openJiuwen, StarHarness, HarnessDev, Prime Agent, RSIAgent, PILOT, Hierarchical Self-Improvement) all attack the same problem from different angles, suggesting that harness optimization is now a research domain on its own.\n\nCode is open source at https:\u002F\u002Fgithub.com\u002FNVlabs\u002FSoL-Pi. Early adopters running GPT-5.6 Sol or Opus 5 through Claude Code or Codex can expect roughly $8.75-13.50 in hourly savings, or $4.36-5.71 versus the Pi harness.\n\nThe catch: EdgeBench is NVIDIA's own eval set, and the gains only have public data for GPT-5.6 Sol and Opus 5. Whether open-weight coding models (Qwen3-Coder, GLM-4.6, Kimi K2.8) see the same 44% reduction is still an open empirical question.\n\nPractical takeaway for agent infra teams: before swapping a smaller model, try swapping the harness first. The savings may be larger than the model delta.","nvidia-sol-pi-harness-token-compression","2026-09-19T03:00:00Z","2026-09-19T03:06:32.446758Z","2026-09-19T03:06:32.446766Z",true,"agent",197,[45,54],{"slug":46,"tag_slug":46,"title_zh":47,"title_en":48,"intro_zh":49,"intro_en":50,"id":51,"is_active":41,"created_at":52,"modified_at":53},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":55,"tag_slug":55,"title_zh":56,"title_en":57,"intro_zh":58,"intro_en":59,"id":60,"is_active":41,"created_at":61,"modified_at":62},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":64},[65,70,75,80,85,90],{"id":66,"title":67,"news_slug":68,"published_at":69},"32b938b6-01a3-43c9-b040-14db6c5f57c6","NVIDIA 把 Agent 装进一个 Python 类:被忽略的 NOOA,一半 token 跑出 SWE-bench 82.2%","nvidia-nooa-python-agent-framework","2026-08-23T17:20:00+00:00",{"id":71,"title":72,"news_slug":73,"published_at":74},"deac2d55-76a6-40d2-8ef7-36aed2ad0105","Linux 7.2 把 AI 拉进内核开发:Sashiko 让补丁数量翻倍,Torvalds 接受「新常态」","linux-7-2-sashiko-ai-kernel-review","2026-08-20T12:00:00+00:00",{"id":76,"title":77,"news_slug":78,"published_at":79},"7ac0ef83-f46d-44f9-846b-a2051fc81e87","NVIDIA NeMo AutoModel：MoE 微调吞吐抬到 3.4–3.7 倍","nvidia-nemo-automodel-moe-finetune-3-7x","2026-06-24T20:00:00+00:00",{"id":81,"title":82,"news_slug":83,"published_at":84},"56cb62a1-da4f-4ac6-94ee-e60346f8d075","英伟达 BioNeMo Agent Toolkit：生命科学库塞进 AI Agent","nvidia-bionemo-agent-toolkit-life-science","2026-06-24T00:00:00+00:00",{"id":86,"title":87,"news_slug":88,"published_at":89},"ad3e5dbd-2c30-43a1-bf67-a6ccd16fa11e","Databricks 开源 Omnigent：Matei Zaharia 想给 Coding Agent 之上再加一层「元 Harness」","databricks-omnigent-meta-harness-coding","2026-06-13T08:00:00+00:00",{"id":91,"title":92,"news_slug":93,"published_at":94},"28c41f06-d20f-481c-b133-cd109af3aed1","答对之后停不下来:微软团队揪出在线蒸馏的 EOS 错配元凶","eos-mismatch-opd-length-inflation","2026-09-18T21:09:06+00:00"]