Skip to main content
Back to timeline
Hugging FaceSource publication:

H company releases Holo4 generalist computer-use agents: 27B scores 61.7% and 35B-A3B 30.9% on OSWorld 2.0, with all trajectories open-sourced

Synopsis

H company released the Holo4 series of generalist computer-use agents in two sizes (27B dense and 35B-A3B Mixture of Experts), where the same model operates desktops, the web, Android, a code sandbox and business APIs through whichever interface fits (GUI, code, MCP, APIs), trained with supervised and reinforcement learning on a large set of environments and tasks including those from its Agentic Task Factory, scoring 61.7% for 27B and 30.9% for 35B-A3B on OSWorld 2.0 against 81.8% for Opus 5.5, open-sourcing every trajectory behind its public-benchmark scores, and turning Nemotron 3 Nano Omni into Holotron4 Nano with the same recipe.

AI-generated editorial illustration: Holo4: powering generalist computer-use agents

Interpretation

Holo4 is a generalist computer-use agent that works across interfaces with the same model weights and the same way of being called, covering desktops, the web, Android, a code sandbox and business APIs. The text states most agentic models are trained for one interface only: GUI-focused models are blind without a screen, while models that prefer tool calling are stuck in front of an application with no API; Holo4 unifies clicking and typing, writing and running its own code, and calling MCP or API tools, choosing whichever fits the task. Supported by model specifications (27B dense, 35B-A3B MoE), the list of available interfaces, and the description of running across platforms; this is a product and capability statement without an ablation on cross-interface generalization.

On long workflows Holo4 trails the strongest closed models but with far fewer parameters and much lower cost: 61.7% for 27B and 30.9% for 35B-A3B on OSWorld 2.0, against 81.8% for Opus 5.5. The text reports significant improvement over the Qwen base and says Holo4 competes with frontier models at much lower cost per task on the hardest academic benchmarks for desktop control (OSWorld 2.0) and API use (AutomationBench). Scores come from public benchmarks, but the text explicitly notes that releases, harnesses and task subsets differ across points, and that costs are estimated from input and output tokens of each agentic run with Holo4 priced at H Models API rates, so cross-point comparison needs care.

Training data comes from an internal Agentic Task Factory that builds interactive environments and verifiable tasks from documentation alone, such as screenshots of real websites or open-source software, producing about 10,000 tasks across web apps, MCP servers and desktop environments, including hybrid environments exposing the same state through a GUI and MCP. It describes environment and task generation as a scalable data pipeline and highlights the hybrid environment form that exposes two interfaces over the same state, distinct from collecting trajectories for a single interface. Gives the scale figure of about 10,000 and the task-type breakdown as a self-report of an internal pipeline, without an independent assessment of task quality or difficulty distribution.

The post-training recipe transfers: as a member of the NVIDIA Nemotron Coalition, the team applied the same recipe to Nemotron 3 Nano Omni to produce Holotron4 Nano, which improves over the base model on GUI workflows and in environments exposing MCP, APIs or coding sandboxes. The text says these gains show the recipe transfers well and can turn a generalist model into an agentic expert, and that nothing in it is size-specific, i.e. it does not depend on one parameter scale. Expressed as absolute percentage-point gains over Nemotron 3 Nano Omni, but the body does not list the specific numbers, stating only that gains are absolute percentage-point improvements.

Perspective

The result is aimed at teams that need to execute long workflows across interfaces in real business settings, for example scenarios involving desktop software, web apps, MCP servers and business APIs at once; the same model and the same way of being called means deployment does not require switching models per platform. For researchers, the public trajectory viewer and dataset make it possible to replay behavior step by step on public benchmarks, and the Agentic Task Factory idea of generating environments and verifiable tasks from documentation can be used to build training and evaluation tasks for other software ecosystems. The text says nothing in the recipe is size-specific, so the post-training stack is meant for teams wanting to turn a new foundation model into an agentic expert rather than for one parameter scale only.

The text is an incomplete release note missing figures and tables, so several key details cannot be confirmed from it: the cost-performance curve values for OSWorld 2.0 and AutomationBench, the specific absolute percentage-point gains of Holotron4 Nano over Nemotron 3 Nano Omni, and the outcome comparison between the two models in the example tasks (only call counts, token volumes and lines of code are given). The text also notes comparison points come from different releases, harnesses and task subsets, and that Holo4 on the AutomationBench private set will be reported once evaluated, so cross-model scores should be read as references under differing conditions. Readers can watch for the DSpark drafter checkpoints, the private-set evaluation, and how hybrid GUI-and-MCP environments perform on real business tasks.

Sources