Holo4: Powering Generalist Computer-use Agents
AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

H Company has released Holo4, a series of models designed to use software through graphical interfaces, code, MCP tools and APIs. The company reports OSWorld 2.0 scores of 61.7% for Holo4 27B and 30.9% for Holo4 35B-A3B, while noting that benchmark setups and costs vary across models. Independent comparisons and Holo4’s performance on AutomationBench’s private set remain outstanding.

H Company has released Holo4, a new series of agentic models that can interact with software through graphical user interfaces, code, MCP tools and APIs. The release includes a 27B dense model and a 35B-A3B Mixture of Experts model, both available through the H Models API; H Company also announced Holotron4 Nano, an updated version of Holotron 3.

H Company describes Holo4 as a single model series for tasks that may span different software interfaces. The models can click and type on a screen, write and run code, and call MCP or API tools. The company says the same model can operate on desktops, the web, Android, a code sandbox and business APIs, without requiring users to select a separate model for each platform.

The company says Holo4 was trained with supervised and reinforcement learning across a large collection of environments and tasks, including tasks generated by its Agentic Task Factory. H Company presents the models as improvements over their Qwen base models and says it developed them for business workflows as well as academic benchmarks. The supplied report does not provide a detailed account of the training data or an independent evaluation of that claim.

On OSWorld 2.0, a desktop-control benchmark, H Company reports a score of 61.7% for Holo4 27B and 30.9% for Holo4 35B-A3B. The report compares the 27B score with 81.8% for Opus 5.5. H Company says the smaller models can compete with leading models at lower cost per task, but the comparison draws on different releases, harnesses and task subsets. The company also publishes trajectories for its public benchmark runs so readers can inspect or download the recorded steps.

At a glance
announcementWhen: Announced in the H Company report; the…
The developmentH Company announced and released Holo4, a pair of agentic models built to combine several ways of interacting with software.

One Model Across Software Interfaces

Many workplace tasks require moving between applications, not just operating one interface. A task might involve entering information through a website, transforming it with code and sending it through an API. H Company’s central proposition is that one model can handle these different steps, which could simplify how developers build software agents for work that crosses systems.

The benchmark results offer an early, measurable reference point, but they do not establish how reliably Holo4 will perform in routine business use. H Company says the models cost less per task than leading alternatives, yet its published comparisons rely on varying evaluation methods and pricing assumptions. The models’ practical value will depend on performance, cost and reliability across comparable tasks outside the company’s own demonstrations.

From Benchmarks to Business Tasks

H Company frames Holo4 as a continuation of its earlier model work, with added emphasis on using several interfaces in the same agent. Its report contrasts this approach with models trained for a single mode of interaction, such as screen control or tool calls. The company says real business tasks can require both.

For desktop control, the report focuses on OSWorld 2.0. It says Holo4 costs are estimated from the input and output tokens used in each agent run, with Holo4 priced at H Models API rates. Other models’ costs draw on different sources and assumptions, including public leaderboard results, model cards, cloud list prices and cache pricing. For AutomationBench, H Company reports results from its internal harness for Holo4 and two Qwen models; other comparisons use public scores and costs from a leaderboard running a private set. The company says it will report Holo4 on that private set after evaluation.

The report also provides examples of tasks involving FreeCAD 3D modeling and building a Pac-Man-style game in Godot. H Company says the examples compare Holo4 27B with its Qwen 3.8 27B base model using the same prompt and harness. Such demonstrations show the kinds of work the company is targeting, but they do not by themselves establish how consistently the models succeed across users, applications or repeated runs.

“We built it for real business workflows.”

— H Company, describing its intended use for Holo4

Independent Comparisons Remain Pending

The supplied report does not specify its publication date, so the precise timing of the announcement cannot be established from this material. It also does not provide independent confirmation of Holo4’s benchmark scores or detailed evidence about performance on business workflows in production.

Cost and score comparisons have limits: H Company says the benchmark releases, harnesses and task subsets differ, and its AutomationBench comparison mixes internal results with public-set scores and private-set costs. Holo4’s score on AutomationBench’s private set is not yet available. The report also does not give a complete, directly comparable account of latency, reliability across repeated runs or the conditions under which the models’ cost estimates apply.

Private-Set Evaluation and Adoption

H Company says it plans to report Holo4’s AutomationBench result on the private set once the model has been evaluated there. Readers can also inspect the company’s released benchmark trajectories and access the models through the H Models API. Those results and artifacts may help developers examine how the models act step by step, while broader independent evaluations will be needed to clarify how their reported performance compares under consistent conditions.

For now, the release establishes what H Company is offering and the scores it reports. How well Holo4 handles varied business tasks, and whether its claimed cost advantages hold across comparable workloads, remain open questions.

Key Questions

What models did H Company release?

H Company released Holo4 27B dense and Holo4 35B-A3B Mixture of Experts. It also announced Holotron4 Nano, an updated version of Holotron 3.

What kinds of software can Holo4 use?

According to H Company, Holo4 can use graphical interfaces, write and run code, and call MCP tools or APIs. The company says it is intended to run across desktops, the web, Android, code sandboxes and business APIs.

What OSWorld 2.0 scores did H Company report?

H Company reports 61.7% for Holo4 27B and 30.9% for Holo4 35B-A3B. It compares the 27B score with 81.8% for Opus 5.5, while cautioning that benchmark setups and task subsets differ across models.

Has Holo4 been evaluated on AutomationBench’s private set?

The supplied report says H Company will report a private-set result once Holo4 has been evaluated there. It currently presents Holo4 results from the company’s internal harness on AutomationBench v1.0.6.

Source: rss

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Pick The Best AI Workflow Automation Tools Of 2026: Top 15

A new 2026 review ranks the top 15 AI workflow automation tools, naming AutoFlow Pro best overall and Zapier AI best for integrations.

Forecasting AI Growth: SenseTime’s Guidance For The First Half Of 2026

SenseTime has issued earnings guidance for the first half of 2026, but no figures or specific financial expectations are disclosed, leaving the outlook unclear.

A Look At SenseTime’s SenseNova U1 Pro Image Model

SenseTime’s new SenseNova U1 Pro image generation model claims output up to 8K, targeting professional workflows; details and independent tests remain sparse.

Breakthrough In AI: SpaceXAI’s Grok Bot Offers Continuous Digital Assistance For $120/Month

SpaceXAI’s Grok Bot reportedly provides continuous AI-powered digital coworkers for $120/month, capable of operating user applications. Details are still emerging.