~/satyajit

Holo4: open computer-use weights, graded on the lab's own benchmarks

mdjsonmcp

2026-10-02 · 16 min · explainer · agents · computer-use · open-weights · mixture-of-experts · reinforcement-learning · benchmarks

On 2026-10-02 H Company released Holo4, a family of generalist computer-use models that, in their words, "click, code, and call tools across desktop, web, Android, code sandboxes, MCP servers, and business APIs." Three checkpoints ship with open weights, alongside quantized variants, a rebuilt agent harness, a described training recipe, and — the part worth the most — every evaluation trajectory behind the scores.

There is no arXiv paper and no technical report. What exists is a blog post, three Hugging Face model cards, and a 7,366-run trajectory dataset. Every benchmark number is H's own, run in H's own harness. That is not a disqualifier — most model launches are self-graded — but it changes what you can conclude from a number, and it is the reason the trajectory release matters more here than the headline percentages do. This piece takes the mechanism apart and then checks three specific claims: how much the post-training actually adds over the Qwen base, how far Holo4 sits behind Opus 5.5 on the hard benchmark, and what licence each checkpoint actually carries.

What a computer-use agent is, and how OSWorld scores one

Start from the interface, because it is the whole game. A computer-use agent does not get a special API into an application. It gets a screenshot. The harness captures the screen, the model looks at the pixels and emits one action — a click at a coordinate, some typed text, a line of code to run in a sandbox, or a tool call — the harness executes it, and the new screen becomes the next observation. The loop repeats, for hundreds of steps on a long task, until the model decides it is finished.

The computer-use loopobserve → decide → act → verify
next observation · up to hundreds of stepsScreenshot+ tool resultsHolo4pick one actionExecuteclick · type · code · toolVerifierpass / fail
1 · observe

The harness sends a screenshot. The agent does not get a DOM or an API by default. It gets a picture of the screen, plus any results from the last action. That is the whole observation.

This is the same design point as Qwen-CUA, which controls a machine from screenshots alone, and it is the opposite instinct from a coding agent's tidy tool table — the subject of Agent harnesses. A shell command can rename 400 files in one call; a GUI agent has to click and type its way through the same job one primitive at a time. The payoff is that the interface never goes stale: it works on software that has no API, which is most software. Holo4 widens the vocabulary past pure pixels — it can also run code and call MCP tools and business APIs directly — but the observe-decide-act spine is unchanged.

The grading is the part people skip. OSWorld is a set of tasks on a real Ubuntu desktop, each with a programmatic verifier that inspects the machine's final state — a file that should exist, a setting that should have changed, an app in the right state — and returns pass or fail. The model never sees the verifier. Averaged over the task set, that pass rate is the OSWorld score. OSWorld 2.0 raises the difficulty to long, multi-step workflows and reports an average partial score as well as strict success, so its numbers run lower and its costs run higher. AndroidWorld does the same for phone apps; AutomationBench does it for business automation over MCP tools and APIs; ALE-CLI (the 105-task Linux split of Agents' Last Exam) does it for expert terminal workflows. All of these are external benchmark suites. What is self-reported is the act of running them.

Three checkpoints, three licences

The family is one recipe applied to three different base models, and the licences are not the same across them — which is the first thing to get straight, because "open weights" is doing a lot of work in the announcement.

CheckpointBase modelArchitectureWeights licence
Holo4-27BQwen3.8-27BdenseCC BY-NC 4.0 (non-commercial)
Holo4-35B-A3BQwen3.6-35B-A3BMoE, 3B activeApache 2.0
Holotron4-30B-A3BNVIDIA Nemotron 3 Nano OmniNemotronH Nano OmniNVIDIA Open Model Agreement

Each ships in BF16, FP8, NVFP4 and 4-bit GGUF (Holotron4 in BF16 and FP8), with a 262,144-token context. The licences are read straight off the model-card metadata, and the spread has a sharp edge in it.

The announcement says Holo4 is "available today on the H Models API for commercial use." True — but that is the hosted API. If you want to download weights and self-host commercially, your only option in this family is Holo4-35B-A3B, the Apache-2.0 MoE. The dense Holo4-27B — which, as the numbers below show, is the stronger model on every benchmark here — is CC BY-NC 4.0: non-commercial. So the strongest model you can run yourself is closed to commercial self-hosting, and the strongest model you can self-host commercially is the weaker MoE. The base models invert this, incidentally: both Qwen3.8-27B and Qwen3.6-35B-A3B are Apache-2.0 upstream, so the non-commercial clause is H's addition on the 27B, not Alibaba's.

Hcompany/Holo4-35B-A3B@5084587 · snapshot 2026-10-02
parameters
35.11B
repo size
70.23 GB
architecture
Qwen3_5MoeForConditionalGeneration
task
image-text-to-text
library
transformers
license
apache-2.0
safetensors
26 shards
largest file
4.00 GB
files
39
downloads
243
likes
22
parameters by dtype
BF1635.11B
multimodalcomputer-useagent

repo last modified 2026-09-28

Holotron4-30B-A3B is the odd one out: built on NVIDIA's Nemotron rather than Qwen, governed by the NVIDIA Open Model Agreement, and interesting mostly as a data point about how far the recipe transports. More on that at the end.

The recipe: fine-tune, then two RL experts, then merge

Three-stage training pipeline. Stage 01, Supervised fine-tuning on 127B tokens: about three quarters are successful agentic trajectories from the Agentic Task Factory — desktop 45 percent, web 14 percent, MCP and API 12 percent, mobile 3 percent — with the remaining 26 percent covering multimodal reasoning, GUI grounding, and text-only tool use and coding. Stage 02, Two RL experts: asynchronous online reinforcement learning on long-horizon tasks trains two specialized LoRA experts on the fine-tuned model, one for desktop and web, one for terminal, MCP and API. Stage 03, One merged model: both experts merge back into the fine-tuned model with equal weight and no further training. A flow graph on the right shows token sources converging into the fine-tuned model, splitting into the two experts under 'Online RL', and merging into Holo4.
The Holo4 training pipeline: a 127B-token SFT stage, two asynchronously-trained LoRA experts, and an equal-weight merge with no further training (Holo4 blog).

The recipe has three stages, and the interesting choice is in how the last two fit together.

Supervised fine-tuning on 127B tokens. About three quarters of that is successful agentic trajectories from H's Agentic Task Factory — desktop (45%), web (14%), MCP and API (12%), and mobile (3%), which sums to the "three quarters" the post claims. The remaining 26% is multimodal reasoning, GUI grounding, and text-only tool use and coding. The word "successful" is load- bearing: this is rejection-sampled imitation. The factory generates tasks, an agent attempts them, and only the runs that the verifier marks as solved become training data. You are teaching the model from its own wins.

The Agentic Task Factory itself is the part worth stealing. It is a set of pipelines that "builds interactive environments and verifiable tasks from documentation alone," and it has produced about 10,000 tasks — roughly 4,000 web-app tasks, 3,000 MCP-server tasks, and 3,000 desktop and OS tasks. The gate on a task is strict: it is kept only if its verifier fails on the untouched seed state, passes on the golden state, rejects every near miss, and an agent actually solved it through the real interface. That last clause is what keeps the task set from filling up with puzzles no agent can reach — the same "keep only tasks with a mix of successes and failures" discipline that scaling agentic RL treats as the real engineering problem behind an RL fleet.

Two asynchronous online-RL experts. On top of the fine-tuned model, H trains two specialized LoRA experts with asynchronous online RL on long-horizon tasks: one for desktop and web, one for terminal, MCP and API. Splitting by workflow class is a bet that the two skill families — pixel-grounded GUI control versus structured tool and code use — pull the weights in different enough directions that a single RL run would have them fighting.

An equal-weight merge, no further training. Then the twist: "both experts merge back into the fine-tuned model with equal weight and no further training." This is weight-space model merging, not a mixture-of-experts router and not a second distillation pass. Two LoRA deltas, each the product of its own RL run, are averaged back onto the shared SFT base. It costs nothing beyond the two RL runs — no joint fine-tune, no router to train — and it is the kind of result that only works because both adapters started from the same initialization, so their updates live in comparable coordinates. Whether equal weight is optimal is exactly the ablation a technical report would carry and this release does not.

The numbers, and what they actually say

Here is the headline table, transcribed from the blog (percentages are scores; dollar figures are cost per task, at H Models API rates for Holo4 and Alibaba Cloud list prices for the Qwen base):

BenchmarkHolo4-27BHolo4-35B-A3BQwen3.8-27B (base)Best frontier shown
OSWorld85.2% · $0.0880.8% · $0.0584.3% · $0.22Qwen3.8 Max 86.1%
OSWorld 2.061.7% · $1.2230.9% · $0.6148.0% · $3.49Opus 5.5 81.8% · $8.48
ALE-CLI44.1% · $0.8230.9% · $0.2943.5%Opus 5.5 63.7% · $8.22
AutomationBench45.4% · $0.0534.5% · $0.0240.3% · $0.09Opus 5 50.3% · $3.05
AndroidWorld85.1% · $0.0877.6% · $0.0781.9% · $0.13Fable 5 88.8%

Before reading anything into the cross-model gaps, read H's own footnotes, because they are honest and they matter. Holo4's scores are the mean of two to four runs in H's harness (a single run on OSWorld 2.0 and ALE-CLI). The frontier numbers are public scores from other providers, "across different harnesses and effort levels" — Opus 5.5 at max effort in Anthropic's harness, costs read off providers' charts. This is not a matched comparison, and H says so. Treat the Holo4-vs-frontier columns as "roughly where each lands," not a controlled result.

Claim 1: how much does the post-training add over the base?

This is the claim I most wanted to check, because "improves significantly over their Qwen base models" is the model card's own phrasing, and on the most-cited benchmark it is barely true.

On OSWorld, Holo4-27B scores 85.2% against the Qwen3.8-27B base's 84.3%. That is +0.9 points (reasoned, from the two reported figures). OSWorld is close to saturated at this point — three systems in the table sit between 84% and 86% — and the post-training moves the needle almost not at all on it. What it moves is the price: $0.08 a task against the base's $0.22, so the base model costs about 2.75x as much to reach a slightly lower score (reasoned). The gain on OSWorld is efficiency, not accuracy.

The accuracy gains show up where the tasks are hard. On OSWorld 2.0, Holo4-27B is at 61.7% against the base's 48.0% — +13.7 points (reasoned) — and on AndroidWorld it is +3.2. The sharpest gain is on tool use: on the Agentic Task Factory's held-out MCP tasks (14 tool servers), Holo4-27B reaches 89.4% against the base's 74.2%, about +15 points. The two RL experts earn their keep on long-horizon control and structured tool calling — exactly the two classes they were specialized on — and contribute almost nothing on the saturated single-step benchmark. That is a coherent story, and a more precise one than "improves significantly."

Claim 2: the price gap to Opus 5.5 on OSWorld 2.0

Score vs cost per task · H Company’s reported numbersleft is cheaper · up is better
0%20%40%60%80%$0.01$0.05$0.1$0.5$1$5$10USD per task · log scaleHolo4-27B61.7% · $1.22Holo4-35B-A3B30.9% · $0.61Qwen3.8-27B base48% · $3.49Opus 5.581.8% · $8.48
Holo4Qwen baseclosed frontier

The hard one. Holo4-27B is 13.7 points over its base here, but 20.1 points under Opus 5.5 — which costs $8.48 a task against Holo4-27B's $1.22.

Holo4 scores are the mean of two to four runs in H’s own harness (a single run on OSWorld 2.0); frontier scores are public numbers from other harnesses at other effort settings, so the across-model comparison is H’s own, not a matched one.

OSWorld 2.0 is where the frontier is clearly ahead. Opus 5.5 scores 81.8% against Holo4-27B's 61.7% — 20.1 points of average partial score (reasoned) — and on strict success it is 48.7% against 41.5%. Holo4 does not close that gap; the honest read is that on the hardest long-horizon benchmark, the best closed model is meaningfully better.

The lever Holo4 pulls instead is cost. Opus 5.5's OSWorld 2.0 run is reported at $8.48 a task; Holo4-27B's at $1.22 — roughly 7x cheaper (reasoned, 8.48 ÷ 1.22 ≈ 6.95). So the trade on offer is explicit: give up about 20 points of score on the hard benchmark to run at a seventh of the price, with the weights on your own machine and the screenshots never leaving it. On the saturated OSWorld that trade is nearly free — Holo4-27B matches the frontier within a point at a fraction of the cost. On OSWorld 2.0 it is a real concession. The scatter above lets you switch benchmarks and watch the trade change shape; it is the blog's cost-vs-score framing, redrawn from the same reported numbers.

Scatter plot titled OSWorld 2.0, average partial score in percent on the vertical axis against USD per task on a log-scale horizontal axis from one cent to one hundred dollars. Holo4-27B is highlighted in violet at about 62 percent and just over one dollar, sitting on the closed-frontier cost-performance line; Holo4-35B-A3B is highlighted lower at about 31 percent near sixty cents. Its Qwen base models sit to the right at higher cost: Qwen3.8 27B near 48 percent at several dollars. Closed frontier models — GPT-6 Astra, Opus 5, GPT-6 Luna, GPT-5.5 — trace a rising line toward the upper right at ten dollars and above.
OSWorld 2.0 average partial score against cost per task. Holo4-27B sits on the closed-frontier cost-performance line while its base model sits to the right of it at higher cost; the strongest closed models are up and to the right, at several times the price (Holo4 blog).

One caveat the scatter cannot show: the AutomationBench headline is measured on 600 public tasks, but 480 of those overlap the split H collected training data from. On the 120 genuinely held-out tasks, Holo4-27B actually scores higher — 49.3% — and the 35B-A3B scores 31.7% against its own Qwen3.6-35B-A3B base's 13.1%, a +18.6-point jump that is the clearest "the training worked" signal in the whole release (reasoned, from the footnote's figures). And on ALE-CLI, H discloses that it excluded a handful of attempts that reached reference answers someone had left on the test machine — 2 of 105 for the 27B, 1 for the 35B-A3B. Small, but the kind of thing you only learn because they wrote it down.

Claim 3: does releasing the trajectories matter?

This is the part I would keep if I had to throw the rest away. H published Hcompany/trajectories: 7,366 agent runs — every run behind the benchmark scores, from both Holo4 models — with each step's reasoning, actions, tool results and screenshots, plus the verifier's result and token usage per trajectory. The per-benchmark counts (1,096 and 1,102 on OSWorld, 106 each on OSWorld 2.0, 1,600 and 1,598 on AutomationBench, and so on) sum to exactly 7,366 (measured). The layout is plain: an index with one summary row per run, a JSON file per trajectory, and WebP screenshots the steps reference. Credentials and personal data are masked.

Why this is the real contribution: when every number is self-reported and there is no protocol document, a score is a claim you cannot check. A trajectory is a claim you can replay. You can open the OSWorld 2.0 runs, watch where the agent succeeds and where it loops, confirm the verifier fired on the state it says it did, and recompute the pass rate yourself. It is the same move cua-s1-forms rewarded in the small-model world — the releases that committed per-row predictions were the ones whose claims survived scrutiny, and the one instability that mattered was only visible because someone published the raw predictions and the provenance to join them. A leaderboard number you can only trust; a trajectory you can audit. For a self-graded release with no report, the 7,366 runs are the report.

Holotron4: the recipe on someone else's base

The third checkpoint is the stress test. Holotron4-30B-A3B applies the same pipeline to NVIDIA's Nemotron 3 Nano Omni instead of a Qwen model, and the gains over that base are large where the base was weak:

BenchmarkInterfaceNemotron 3 Nano OmniHolotron4-30B-A3BGain
OSWorldGUI21.076.3+55.3
OSWorld 2.0GUI and code0.27.9+7.7
AutomationBenchMCP19.435.6+16.2
PinchBenchTerminal84.788.6+3.9
ALE (Linux, code)Terminal0.68.5+7.9

The OSWorld jump from 21.0 to 76.3 is the eye-catching one, but read it the right way: a base model at 21% on OSWorld was barely a computer-use agent at all, so a +55.3-point gain is partly a measure of how much headroom there was, not only how good the recipe is. The absolute Holotron4 OSWorld 2.0 number is 7.9% — the recipe transports, but a weaker base stays a weaker agent on the hard benchmark. It is a useful demonstration that the Agentic Task Factory plus the two-expert merge is not Qwen-specific, and an honest one, because the small absolute numbers are printed right next to the large deltas.

The scorecard

Holo4 is a real release with three things in it that are genuinely useful: an Apache-2.0 MoE you can actually self-host commercially, a described and reproducible-in-principle recipe (SFT on rejection-sampled trajectories, two asynchronous RL LoRA experts, an equal-weight merge), and 7,366 evaluation trajectories that turn self-reported scores into auditable ones. The price story is real: on the benchmarks that are near saturated, Holo4-27B matches the frontier within a point at a fraction of the cost, and that is a meaningful shift in what an open model buys you.

The honest counterweights are three. The strongest checkpoint, Holo4-27B, is non-commercial — the Apache model is the weaker one, so "open and commercial" and "strongest" do not point at the same file. The accuracy gain over the Qwen base is small on the headline benchmark (+0.9 on OSWorld) and concentrated on the harder long-horizon and tool-calling tasks, which is a narrower and more defensible claim than the one the model card makes. And on the hardest benchmark, OSWorld 2.0, Opus 5.5 is still about 20 points ahead — bought, admittedly, at roughly seven times the cost. All of these are visible in the release because H wrote the footnotes and shipped the trajectories, which is the behaviour you want from a lab grading its own exam, and the reason this release is worth taking at close to face value even without a report.


Built on H Company's Holo4 announcement (2026-10-02) and the Hugging Face model cards for Holo4-27B, Holo4-35B-A3B and Holotron4-30B-A3B, plus the trajectories dataset. There is no technical report. Figures are reproduced from the blog for commentary, flattened onto white. Every benchmark number is H's own, self-reported; the agent-loop diagram and the cost-vs-score scatter are my own illustrations of the mechanism and the reported numbers, not measured traces.

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "Holo4: open computer-use weights, graded on the lab's own benchmarks", ai.thesatyajit.com, October 2026.

bibtex
@misc{ghana2026holo4,
  author = {Satyajit Ghana},
  title  = {Holo4: open computer-use weights, graded on the lab's own benchmarks},
  url    = {https://ai.thesatyajit.com/articles/holo4},
  year   = {2026}
}
share