[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-holo3-1-h-company-computer-use-quant":3,"news-related-c2e18ecb-00b3-47ee-935a-f0e3cd0dee4a":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"c2e18ecb-00b3-47ee-935a-f0e3cd0dee4a","Holo3.1 把 Computer Use Agent 拉进本地：FP8\u002FNVFP4\u002FQ4 三种量化让消费级 GPU 跑得动","H Company 6 月 2 日开源 Holo3.1——面向 Computer Use Agent 的视觉-动作模型，0.8B\u002F4B\u002F9B\u002F35B-A3B 四种规格，基于 Qwen 微调。这是该领域首次提供 FP8、Q4 GGUF、NVFP4 三种量化权重，让电脑操作 Agent 第一次在消费级硬件上本地运行。\n\nAndroidWorld 移动端 35B-A3B 从 67% 升到 79.3%，4B\u002F9B 从 58% 升到 72%；跨框架新增 function-calling 支持；Holotab harness 相对 Holo3 提升 25% 以上。NVFP4 用 NVIDIA Model Optimizer W4A16 生成，DGX Spark 端到端步时从 6.8s 压到 3.3s；Q4 GGUF 瞄准 Apple Silicon。\n\nComputer Use Agent 最大障碍是延迟和隐私。Holo3.1 用「小模型+激进量化」把能力下放到 4B\u002F9B 档，35B-A3B 留给云端。它走的是和 frontier 模型相反的路：开源、量化、本地优先。这才是企业落地的真正起点。","https:\u002F\u002Fhuggingface.co\u002Fblog\u002FHcompany\u002Fholo31","d48b2c3e-bb69-4483-afb6-3ca22fc6c06f",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":18,"name":19,"slug":19,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b49648f9-963e-4082-8684-3d085b7358fe","quantization",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"3dbeec31-5bb0-4b99-b4af-1bebbd9b5266","en","Holo3.1 runs Computer Use agents locally on consumer GPUs","H Company open-sourced Holo3.1 on June 2 — a vision-action model for Computer Use Agent, with four specs of 0.8B \u002F 4B \u002F 9B \u002F 35B-A3B, fine-tuned based on Qwen. This is the first time the field has provided FP8, Q4 GGUF, and NVFP4 three quantization weights, letting computer-operation Agents run locally on consumer-grade hardware for the first time.\n\nAndroidWorld mobile 35B-A3B rose from 67% to 79.3%, 4B\u002F9B from 58% to 72%; cross-framework adds function-calling support; Holotab harness improves over Holo3 by more than 25%. NVFP4 is generated with NVIDIA Model Optimizer W4A16, DGX Spark end-to-end step time compressed from 6.8s to 3.3s; Q4 GGUF targets Apple Silicon.\n\nThe biggest obstacle for Computer Use Agent is latency and privacy. Holo3.1 uses \"small model + aggressive quantization\" to bring capability down to the 4B\u002F9B tier, with 35B-A3B reserved for the cloud. It takes the path opposite to frontier models: open source, quantized, local-first. This is the true starting point for enterprise deployment.","holo3-1-h-company-computer-use-quant","2026-06-07T06:00:00Z","2026-06-07T06:14:28.477955Z","2026-08-19T02:08:40.142862Z",true,"agent",144,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"988bbfb8-672a-4c6c-98f7-3a170b6bd8b3","Macaw 把 LFM2.5 装进 1.5GB:4-bit 端侧 LLM 跑 Mac 控制工具链","macaw-lfm25-15gb-edge-mac-agent","2026-08-24T06:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"d5203c93-2022-4769-a6d2-c7765ded2b40","腾讯混元 Hunyuan-A13B 开源实测:80B 总参 \u002F 13B 激活,GQA + FP8\u002FINT4 把 MoE 推理门槛打到消费卡","tencent-hunyuan-a13b-80b-13b-gqa-angelslim-moe","2026-07-30T06:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"1abe59b8-844d-4bb1-bc27-cfe28099d101","Anemll\u002FFlash-iOS：把 400B MoE 大模型塞进 iPhone 的端侧实验","anemll-flash-ios-400b-moe-iphone-edge","2026-06-07T12:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"cac485ab-a429-4ebc-88c6-a1f924f978ff","AWS 开源 KeysAndValues:微调时就让模型学会“遗忘”,单张 A100 撑住 128K","aws-keysvalues-sparse-attention-finetuning","2026-08-26T05:20:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"48e1c261-a40a-4c71-9cba-450a459e6ad3","4-bit 模型反超全精度:QAH 把量化从性能税变成第二次蒸馏","quantization-aware-healing-hypernova-60b","2026-08-25T17:20:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"c94bdf86-5de9-49fe-8c98-0f5c47611bfe","SGLang v0.5.18 发布:大模型冷启动提速 2.38 倍,710 个 PR 都改了什么","sglang-v0-5-18-cold-start-2-38x","2026-08-24T23:15:00+00:00"]