Holo4: 27B and 35B-A3B agentic models for cross-interface computer use

A new series of agentic models called Holo4 is available in two sizes: a 27B dense model and a 35B-A3B Mixture of Experts variant. An updated Nano agent, Holotron4 Nano, is also released as the team’s follow-up to Holotron 3. The team reports both Holo4 sizes can act through multiple interfaces — GUIs, code sandboxes, MCP and APIs — using whichever interface a task requires.

Models and capabilities

The authors describe Holo4 as a single model that runs on desktop, web, Android, inside code sandboxes and against business APIs, accessed via the H Models API. Holo4 is said to type and click on screens, generate and execute its own code, and call MCP or API tools when appropriate. Training combined supervised and reinforcement learning on a large collection of environments and tasks, including those produced by an internal Agentic Task Factory.

The team reports the Agentic Task Factory has produced about 10,000 verifiable tasks spanning web apps, MCP servers and desktop environments, including hybrid setups that expose the same state via GUI and MCP.

Benchmarks, training and releases

On the desktop-control benchmark OSWorld 2.0, the authors report Holo4-27B scored 61.7% compared with 81.8% for Opus 5.5; Holo4 35B-A3B reached 30.9% on the same benchmark. The team says Holo4 competes with frontier closed models on OSWorld 2.0 and AutomationBench while achieving lower cost per task and using far fewer parameters. Cost estimates in their charts are derived from input/output tokens, H Models API pricing and other cloud list prices as described by the authors.

Examples included by the authors compare Holo4-27B and its base model Qwen3.8 27B on interactive tasks. For a FreeCAD Eiffel Tower task Holo4-27B used 84 calls and 1.3M tokens versus Qwen3.8’s 60 calls and 1.0M tokens. For a Pac-Man-style game Holo4-27B used 68 calls, 2.4M tokens and produced 268 lines of code; Qwen3.8 required 197 calls and 11.4M tokens.

The authors say the training harness was rebuilt to provide reliable long-term memory over hundreds of steps and to expose a desktop shell, and that the post-training stack adapts to different foundation models. Weights are released in BF16, FP8, NVFP4 and 4-bit GGUF formats, both sizes are accessible via the H Models API, and optimized DSpark drafter checkpoints are expected to follow.


Original source: Hugging Face Blog

Leave a Comment