🔍 Read the full analysis: Why Holo4 Is Designed For Generalist Computer-Use Agents on ThorstenMeyerAI.com
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
H Company has released Holo4, an open-weight model series designed to handle software through screens, code and tools with one model. The company reports a 61.7% score for its 27B model on OSWorld 2.0, but the results have not been independently verified.
H Company has released Holo4, an open-weight series of agentic models designed to operate software through graphical interfaces, code, MCP tools and APIs, as detailed in the original analysis. The company reports that its 27-billion-parameter model scored 61.7% on OSWorld 2.0; independent evaluations have not yet established how the models compare with other systems.
Holo4 is available in two configurations: a 27B dense model and a 35B-A3B mixture-of-experts model. H Company says both can click and type on screens, write and run code, and call MCP or API tools, selecting an interface suited to the task. The company describes the models as usable across desktop software, the web, Android, code sandboxes and business APIs. They are offered through the H Models API and as downloads on Hugging Face in FP16, FP8 and GGUF formats.
H Company reports that Holo4 27B scored 61.7% on OSWorld 2.0, while Holo4 35B-A3B scored 30.9%. The company compares the 27B result with 81.8% for Opus 5.5, which it identifies as the strongest closed model in its comparison. These are company-reported figures; the announcement says model releases, evaluation harnesses and task subsets differ across comparisons.
The company says the models were trained with supervised and reinforcement learning across a large collection of environments and tasks, including tasks generated by its Agentic Task Factory. It has also published the trajectories behind its public benchmark scores through trajectories.hcompany.ai and Hugging Face, allowing outside reviewers to inspect the recorded steps. H Company presents examples in FreeCAD and Godot as evidence of improvement over its Qwen base models, but those demonstrations do not establish performance across typical business workflows.
One Model Across Work Interfaces
Many software tasks move between screens, scripts and services. An agent that can use only a graphical interface may stall when a direct API is available, while a tool-calling agent may be unable to work in software without an exposed API. H Company is pitching Holo4’s ability to use several interfaces as a way to handle those handoffs within one model.
If independent testing confirms the reported OSWorld result and the models prove reliable on routine work, open weights could give developers more options to run or adapt computer-use agents. H Company also claims lower costs than large closed models, but the figures depend on its pricing assumptions and evaluation setup. The release therefore offers a testable performance claim, while the practical cost and reliability benefits remain to be established.
From Holo1 to Holo4
Holo4 follows H Company’s earlier Holo1 agentic model and arrives alongside an updated model called Holotron4 Nano. According to the company’s benchmark notes, Holo4 27B is based on Qwen3.8 27B and the 35B-A3B version on Qwen3.6 35B-A3B. The release combines open-weight downloads with API access, giving developers a route to try the models without relying solely on a hosted service.
H Company’s cost comparisons use H Models API rates for Holo4 and Alibaba Cloud list prices for Qwen, with cache hits priced at 20% of input cost for the MoE model. The company says its GPT and Opus effort sweeps draw on OpenAI launch data. It also cautions that compared releases, harnesses and task subsets differ. Those qualifications limit what can be concluded from a direct ranking.
“Real work is not siloed that way, and a single business task can require combining these different approaches.”
— H Company
Benchmark Comparisons Need Review
Holo4’s benchmark results are self-reported and have not yet been independently reproduced in the supplied material. H Company says its evaluation setup differs from those used for other models. On AutomationBench, the company’s internal harness version 1.0.6 was used for Holo4, while comparison scores for other models came from the public set; the cost figures draw on a leaderboard using a private set. Holo4 has not yet been evaluated on that private set.
The announcement does not explain why the 35B-A3B model scored 30.9% on OSWorld 2.0, well below the 27B model’s 61.7%. It also provides no independent results showing how often the agents complete real business workflows reliably, or how the reported cost advantage holds across different usage patterns. Published trajectories may help reviewers examine benchmark runs, but their release alone does not settle those questions.
Private Set Results Pending
H Company says it plans to report Holo4’s results on the AutomationBench private set after that evaluation is complete. The public weights and benchmark trajectories also make it possible for external researchers and developers to reproduce runs or evaluate the models on other tasks. No independent replication or adoption results are included in the announcement.
Developers can access Holo4 through the H Models API or download model files from Hugging Face. Further comparisons will be needed to establish whether the benchmark scores carry over to varied software and business settings, and whether the models’ performance and operating costs meet users’ needs.
Key Questions
What is Holo4?
Holo4 is H Company’s series of open-weight agentic models designed to use graphical interfaces, code, MCP tools and APIs to operate software.
What benchmark result did H Company report?
The company reports that Holo4 27B scored 61.7% on OSWorld 2.0 and Holo4 35B-A3B scored 30.9%. These results have not been independently verified in the supplied material.
Where can developers get the models?
H Company offers Holo4 through its H Models API and as downloads on Hugging Face in FP16, FP8 and GGUF formats.
Have Holo4’s AutomationBench results been independently compared?
Not in the supplied announcement. H Company used its internal harness for Holo4, while comparison scores came from AutomationBench’s public set; the company says Holo4 has not yet been evaluated on the private set.
Primary source: Hugging Face · via ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
