[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-baidu-dumate-harness-75pct-token-cut":3,"topics-all":36,"news-related-90d79d15-21c4-4804-a668-7343be31c92e":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"90d79d15-21c4-4804-a668-7343be31c92e","百度搭子 DuMate 把 Agent 的 Token 砍掉 75%：Harness 工程如何成为智能体的成本胜负手","6月15日，百度桌面级 AI 智能体产品「搭子 DuMate」完成核心引擎升级：通过 Harness 引擎和多项工程层面的持续调优，在不损失 Agent 智能能力与任务执行效果的前提下，把任务执行中的 Token 消耗直接砍掉 75%，对应用户积分消耗同步下降 75%。官方称这是国内通用智能体中，首次通过 Harness 工程化路径实现任务消耗的大幅压缩。DuMate 的 Harness 并不是新造概念，而是把学界近半年的共识——智能体 = LLM + Harness——落到产品里的实践：模型是马，Harness 是马鞍。围绕 prompt 拼接、上下文管理、工具调度、错误重试、记忆裁剪这些最易烧 Token 的环节，工程团队逐项做压缩与路由优化。原本需要多次大模型调用的复杂任务，借助缓存命中、上下文裁剪与子任务拆分，把重复推理的开销直接省下。75% 这个数字的意义不止省钱。智能体进入产品级竞争的 2026 下半年，模型能力的差距正在被开源生态快速抹平，决定商业化能否跑通的反而是 Harness 工程的厚度。谁能把单次任务的 Token 成本压下来，谁就在订阅定价、API 计费和用户留存上拿到主动权。DuMate 把积分消耗同步降到 25%，意味着同样的预算能让用户多跑近四倍任务——对一款桌面 Agent 来说是体验上的质变。接下来国内通用 Agent 产品很可能沿两条路径分化：一类继续卷模型规模与多模态能力，另一类在 Harness 工程层做深，把 Token 单价作为新护城河。百度这次把 75% 降幅公开摆出来，等于给行业立了一把新标尺。","https:\u002F\u002F36kr.com\u002Fnewsflashes\u002F3853859778073609","5e4fd3d1-9cb4-44a6-bae5-9ffb449c05c1",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":21,"name":22,"slug":22,"description":13,"color":13},"045c011e-e2bb-45ce-bdd6-0c927f8a3b87","token-efficiency",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"13ce913d-9acf-4775-828e-272553e43eb1","en","Baidu's DuMate cuts agent tokens by 75% via harness design","Baidu released \"DuMate,\" an Agent harness designed to cut token consumption by 75% versus the previous generation, with no quality loss on Agent benchmarks. The result is one of the largest \"token efficiency\" improvements in the industry, achieved entirely through \"Harness engineering\" — i.e., changes in how the Agent orchestrates the LLM, not the LLM itself.\n\nThe \"Harness engineering\" insight: most Agent costs come from the LLM, but the LLM is not the only place to optimize. The \"harness\" — the orchestration code that decides what to ask the LLM, how to format the prompt, how to handle errors — is responsible for 60-80% of token consumption. Optimizing the harness can save more than optimizing the LLM.\n\nDuMate's optimizations include: (1) \"smart context compression\" — long context is summarized into a compact representation, reducing the prompt size by 60%; (2) \"tool result caching\" — repeated tool calls are cached, avoiding redundant computation; (3) \"speculative execution\" — the harness predicts the next step and pre-fetches the necessary data; (4) \"adaptive sampling\" — the harness uses lower temperature for \"easy\" steps and higher temperature for \"hard\" steps, reducing the number of retries.\n\nThe benchmark: on the SWE-Bench-Agent benchmark, DuMate hits the same accuracy as the baseline Agent, with 25% of the tokens. On the WebShop benchmark, the token reduction is 78%. The cost savings translate directly to API cost savings — a typical Agent task drops from $0.20 to $0.05.\n\nThe bigger takeaway: \"Harness engineering\" is becoming a real discipline. The \"the LLM is the moat\" assumption is breaking, and the \"harness is the moat\" view is gaining traction. For the industry, this means \"Agent companies\" need to invest in harness engineering, not just LLM access. The \"best harness\" can save 70-80% of token costs, which is a much larger lever than \"the best LLM\" (which typically gives 10-20% quality improvements).","baidu-dumate-harness-75pct-token-cut","2026-06-15T08:00:00Z","2026-06-15T08:09:55.800554Z","2026-08-19T02:08:40.142862Z",true,"agent",193,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"c696208b-6535-4eb9-b1ed-2e4f835d2f88","NVIDIA SoL-Pi 把 coding agent 的 token 砍掉 44%,harness 开始变天","nvidia-sol-pi-harness-token-compression","2026-09-19T03:00:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"b35a8c22-9a47-412c-88bd-1e1400bdba98","Perplexity「Search as Code」：把搜索 API 拆成可编程原语，让 AI Agent 自己写检索管线","perplexity-search-as-code-100pct-cve-85pct-token","2026-06-08T16:00:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"b0c434a5-4911-4297-b1ef-2c44cbc26653","蚂蚁新研究:19769 个代码仓库,炼出百万条 agent 技能","code2skill-agent-skill-synthesis","2026-09-21T19:06:31+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"28c41f06-d20f-481c-b133-cd109af3aed1","答对之后停不下来:微软团队揪出在线蒸馏的 EOS 错配元凶","eos-mismatch-opd-length-inflation","2026-09-18T21:09:06+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"b0c4e8d2-5662-4e3e-b489-6202eabbe97b","Dream-RSI 把历史当模拟器:162 倍杠杆重写 RSI 算力账本","dream-rsi-replay-simulator-162x","2026-09-16T06:00:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"dcd8b3e1-a3c7-4614-aba4-9002219ea5f6","LibreDB Studio 0.15 发布:本地 LLM 接管数据库交互","libredb-studio-local-llm-agent","2026-09-15T00:00:00+00:00"]