H Company open-sourced Holo3.1 on June 2 — a vision-action model for Computer Use Agent, with four specs of 0.8B / 4B / 9B / 35B-A3B, fine-tuned based on Qwen. This is the first time the field has provided FP8, Q4 GGUF, and NVFP4 three quantization weights, letting computer-operation Agents run locally on consumer-grade hardware for the first time.

AndroidWorld mobile 35B-A3B rose from 67% to 79.3%, 4B/9B from 58% to 72%; cross-framework adds function-calling support; Holotab harness improves over Holo3 by more than 25%. NVFP4 is generated with NVIDIA Model Optimizer W4A16, DGX Spark end-to-end step time compressed from 6.8s to 3.3s; Q4 GGUF targets Apple Silicon.

The biggest obstacle for Computer Use Agent is latency and privacy. Holo3.1 uses "small model + aggressive quantization" to bring capability down to the 4B/9B tier, with 35B-A3B reserved for the cloud. It takes the path opposite to frontier models: open source, quantized, local-first. This is the true starting point for enterprise deployment.