On September 28, French AI startup H Company released Holo4, its new series of agentic models: a 27B dense version and a 35B-A3B mixture-of-experts variant (35B total parameters, roughly 3B active per token), both built on the Qwen3.8-27B base. The release also includes Holotron4 Nano, an agentic rework of NVIDIA's Nemotron 3 Nano Omni using the same post-training recipe.
One model, four interfaces
Holo4's core design is interface generality: the same model clicks and types on screens, writes and runs its own code, and calls MCP or API tools. As the company notes, most agentic models are trained for a single interface — GUI-focused models go blind without a screen, tool-calling models stall in front of apps with no API — while real business tasks routinely mix these modes. Across desktops, the web, Android, code sandboxes and business APIs, it is the same model called the same way.
Closing on the frontier at one-seventh the cost
By H Company's own reporting, on OSWorld 2.0, the hardest academic benchmark for desktop control, Holo4 27B scores 61.7 (average partial score) with a 41.5% success rate at roughly $1.22 per task. Claude Opus 5.5 sits at 81.8 / 48.7% / $8.48, and GPT-6 Astra at 73.5 / $9.07. The Qwen3.8 27B base manages only 48.0 at $3.49 per task — the post-training gain is plain to see. Note that the original OSWorld is a different measurement: on short tasks Holo4 27B scores 85.2 against the base's 84.3, a modest gain; the real separation shows up on long workflows. On AutomationBench, a benchmark for API automation, it scores 45.4 at $0.05 per task.
For training, the company describes supervised fine-tuning on 127B tokens, followed by two RL experts that are merged. Environments and tasks come from an "Agentic Task Factory" that builds verifiable tasks from documentation alone — screenshots of real websites, open-source software docs — and has produced about 10,000 tasks so far. On the engineering side, the team rebuilt its harness around OSWorld 2.0 failure analysis: agents get reliable memory across hundreds of steps and a shell on the desktop machine itself. The company's Pac-Man demo shows the token-efficiency gap: Holo4 27B finishes with 68 calls and 2.4M tokens; the base needs 197 calls and 11.4M.
Three cold showers
License split. Both sizes ship downloadable weights (BF16, FP8, NVFP4, 4-bit GGUF), but the strongest 27B is CC BY-NC 4.0 — no commercial use. Only the 35B-A3B is Apache 2.0, and it scores just 30.9 on OSWorld 2.0. Self-hosting inside a commercial product means living with the half-score variant.
Self-reported numbers and training overlap. Every score comes from H Company's own harness (the company open-sources every trajectory behind its scores at trajectories.hcompany.ai for step-by-step replay, which deserves credit). But MarkTechPost notes that 480 of AutomationBench's 600 public tasks fall in the split H Company collected training data from; on the 120 held-out tasks, the 27B scores 49.3. canberk.me adds that the 85.2 on original OSWorld and the 61.7 on OSWorld 2.0 are different measurements and should not be mixed in marketing.
The long-task gap is structural. 61.7 versus 81.8 is a 20-point gap, and 41.5% versus 48.7% success is just as hard. Open-weight models have pushed cost into one-seventh of frontier territory for the first time, but between "cheap and usable" and "reliable" sit exactly these numbers.
So what
Read the license before the benchmark — that is Holo4's most practical lesson. For the industry, it sets trajectory-level openness as a new transparency standard for agents, and offers a reproducible sample of 27B-scale post-training closing on the frontier. The company says DSpark drafter checkpoints for faster inference arrive within days — worth watching for anyone doing inference optimization.
References: huggingface.co/blog/Hcompany/holo4; marktechpost.com; canberk.me; techaiwire.com