[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-holo4-open-weight-computer-use":3,"topics-all":38,"news-related-bbe8d55a-1069-42fa-a342-d944d53fdb4b":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"bbe8d55a-1069-42fa-a342-d944d53fdb4b","操作电脑的 27B 开源权重模型 Holo4:最强版禁商用","H Company 发布开源 Agent 模型 Holo4:27B 稠密与 35B-A3B MoE 双版本,可操作屏幕、写代码、调工具。OSWorld 2.0 上 27B 官方自报 61.7 分、单任务 1.22 美元,约为 Opus 5.5 的七分之一;但 27B 禁商用,可商用版仅 30.9 分。","9 月 28 日,法国 AI 公司 H Company 发布新一代 Agent 模型 Holo4:27B 稠密与 35B-A3B MoE 双版本(总参 35B、激活约 3B),基座 Qwen3.8-27B;同场的 Holotron4 Nano 是 Nemotron 3 Nano Omni 的同配方 Agent 改造版。\n\n## 一个模型,四种接口\n\nHolo4 的核心设计是「全接口通用」:同一个模型既能点击、输入、操控屏幕,也能写代码并运行,还能调 MCP 和 API 工具。官方指出,多数 Agent 模型只为单一接口训练:专注 GUI 的没屏幕就失明,偏好工具调用的碰上无 API 的应用就卡死,而真实任务常需混用。桌面、网页、Android、沙箱和业务 API 上都是同一个模型、同一种调用方式。\n\n## 跑分逼近 frontier,成本砍到七分之一\n\n按 H 公司自报数据,最难的桌面操控基准 OSWorld 2.0 上,Holo4 27B 拿 61.7 分(平均部分得分)、成功率 41.5%、单任务约 1.22 美元;对照 Claude Opus 5.5 的 81.8 分、48.7%、8.48 美元,GPT-6 Astra 的 73.5 分、9.07 美元。基座 Qwen3.8 27B 只有 48.0 分、3.49 美元,后训练增益肉眼可见。注意原版 OSWorld 是另一回事:短任务上 27B 拿 85.2、基座 84.3,差距小,真正拉开的是长流程。API 自动化基准 AutomationBench 上它拿 45.4 分、每任务 0.05 美元。\n\n训练上先用 127B token 做监督微调,再训两个 RL 专家合并;任务来自「Agentic Task Factory」——仅凭文档自动构建可验证任务,已产出约 1 万个;harness 也重建:跨数百步的可靠记忆加桌面机上的 shell。官方 Pac-Man 演示可见 token 效率差距:Holo4 用 68 次调用、2.4M token 完成,基座要 197 次、11.4M。\n\n## 泼冷水的三处\n\n**许可证分裂。** 两个尺寸都开放权重下载(BF16、FP8、NVFP4、4-bit GGUF),但最强的 27B 是 CC BY-NC 4.0 禁商用;只有 35B-A3B 是 Apache 2.0,而它 OSWorld 2.0 只有 30.9 分。想自托管做商业产品,只能用分数减半的版本。\n\n**自报数字与训练集重叠。** 所有分数出自自家 harness(完整轨迹开源在 trajectories.hcompany.ai 供回放,值得肯定);但 MarkTechPost 指出,AutomationBench 600 个公开任务里 480 个落在训练数据划分内,120 个保留任务上 27B 得 49.3。canberk.me 也提醒,85.2 与 61.7 是两种度量,不能混着宣传。\n\n**长任务差距仍是结构性的。** 61.7 对 81.8 差 20 个点,成功率 41.5% 对 48.7% 也是硬差距——开源阵营第一次把成本打进 frontier 七分之一区间,但「便宜能用」和「可靠交付」之间还隔着这些数字。\n\n## 所以呢\n\n选型先看许可证再看跑分,这是 Holo4 最实际的教训。对行业,它把「轨迹全开源」立成 Agent 透明度新标准,给出 27B 级后训练逼近 frontier 的可复现样本;DSpark 加速 checkpoint 几天内跟进,推理优化可盯一眼。\n\n参考:huggingface.co\u002Fblog\u002FHcompany\u002Fholo4;marktechpost.com;canberk.me;techaiwire.com","https:\u002F\u002Fhuggingface.co\u002Fblog\u002FHcompany\u002Fholo4","d48b2c3e-bb69-4483-afb6-3ca22fc6c06f",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":19,"name":20,"slug":20,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"0366ec51-6042-457f-9869-d72c598505e0","en","Holo4: Open-Weight 27B Computer-Use Agent With a License Catch","H Company's Holo4 open-weight agents reach 61.7% on OSWorld 2.0 at $1.22 per task, but the strongest 27B is non-commercial.","On September 28, French AI startup H Company released Holo4, its new series of agentic models: a 27B dense version and a 35B-A3B mixture-of-experts variant (35B total parameters, roughly 3B active per token), both built on the Qwen3.8-27B base. The release also includes Holotron4 Nano, an agentic rework of NVIDIA's Nemotron 3 Nano Omni using the same post-training recipe.\n\n## One model, four interfaces\n\nHolo4's core design is interface generality: the same model clicks and types on screens, writes and runs its own code, and calls MCP or API tools. As the company notes, most agentic models are trained for a single interface — GUI-focused models go blind without a screen, tool-calling models stall in front of apps with no API — while real business tasks routinely mix these modes. Across desktops, the web, Android, code sandboxes and business APIs, it is the same model called the same way.\n\n## Closing on the frontier at one-seventh the cost\n\nBy H Company's own reporting, on OSWorld 2.0, the hardest academic benchmark for desktop control, Holo4 27B scores 61.7 (average partial score) with a 41.5% success rate at roughly $1.22 per task. Claude Opus 5.5 sits at 81.8 \u002F 48.7% \u002F $8.48, and GPT-6 Astra at 73.5 \u002F $9.07. The Qwen3.8 27B base manages only 48.0 at $3.49 per task — the post-training gain is plain to see. Note that the original OSWorld is a different measurement: on short tasks Holo4 27B scores 85.2 against the base's 84.3, a modest gain; the real separation shows up on long workflows. On AutomationBench, a benchmark for API automation, it scores 45.4 at $0.05 per task.\n\nFor training, the company describes supervised fine-tuning on 127B tokens, followed by two RL experts that are merged. Environments and tasks come from an \"Agentic Task Factory\" that builds verifiable tasks from documentation alone — screenshots of real websites, open-source software docs — and has produced about 10,000 tasks so far. On the engineering side, the team rebuilt its harness around OSWorld 2.0 failure analysis: agents get reliable memory across hundreds of steps and a shell on the desktop machine itself. The company's Pac-Man demo shows the token-efficiency gap: Holo4 27B finishes with 68 calls and 2.4M tokens; the base needs 197 calls and 11.4M.\n\n## Three cold showers\n\n**License split.** Both sizes ship downloadable weights (BF16, FP8, NVFP4, 4-bit GGUF), but the strongest 27B is CC BY-NC 4.0 — no commercial use. Only the 35B-A3B is Apache 2.0, and it scores just 30.9 on OSWorld 2.0. Self-hosting inside a commercial product means living with the half-score variant.\n\n**Self-reported numbers and training overlap.** Every score comes from H Company's own harness (the company open-sources every trajectory behind its scores at trajectories.hcompany.ai for step-by-step replay, which deserves credit). But MarkTechPost notes that 480 of AutomationBench's 600 public tasks fall in the split H Company collected training data from; on the 120 held-out tasks, the 27B scores 49.3. canberk.me adds that the 85.2 on original OSWorld and the 61.7 on OSWorld 2.0 are different measurements and should not be mixed in marketing.\n\n**The long-task gap is structural.** 61.7 versus 81.8 is a 20-point gap, and 41.5% versus 48.7% success is just as hard. Open-weight models have pushed cost into one-seventh of frontier territory for the first time, but between \"cheap and usable\" and \"reliable\" sit exactly these numbers.\n\n## So what\n\nRead the license before the benchmark — that is Holo4's most practical lesson. For the industry, it sets trajectory-level openness as a new transparency standard for agents, and offers a reproducible sample of 27B-scale post-training closing on the frontier. The company says DSpark drafter checkpoints for faster inference arrive within days — worth watching for anyone doing inference optimization.\n\nReferences: huggingface.co\u002Fblog\u002FHcompany\u002Fholo4; marktechpost.com; canberk.me; techaiwire.com","holo4-open-weight-computer-use","2026-10-02T13:12:01Z","2026-10-02T13:13:16.791721Z","2026-10-02T13:13:16.791732Z",true,"agent",43,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"dc61debd-1d77-4d5a-9d59-5b23c3da07de","蚂蚁开源Realtime-Venus：9B全双工模型边说边干活，三项续聊指标超GPT-4o","ant-realtime-venus-full-duplex-delegation","2026-09-30T23:10:53+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"983fb4d1-6c63-4828-adbf-a59c57e02b64","Perplexity开源决策模型:总分微胜Jev","perplexity-decider-v1-27b-open-source","2026-10-02T21:20:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"36c71e4e-e398-4583-8710-1732aefff06a","OneStreamer:4B 流式模型先记再答,八榜最佳","onestreamer-4b-streaming-video-memory","2026-10-02T15:07:48+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"7a2aa1aa-74ec-494a-a828-e6c1ad5b4cfe","Cloudflare Clef 决策模型开源,Jev 被压制","cloudflare-clef-open-source-decision-model-jev","2026-10-02T09:00:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"0d357a0b-42da-40af-8038-e8c035cb9810","Apple 开源 LensVLM-9B:先扫压缩图,再读原页","apple-lensvlm-9b-weights-huggingface","2026-09-26T15:20:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"d055ddb8-4d82-4523-99b7-39c5f77e2ff7","PhysBrain 1.5 开源：8B 具身基座 28 项评测均分 72.5，官方称追平 GPT-6-Astra","physbrain-1-5-open-embodied-base","2026-09-16T21:07:24+00:00"]