# Holo4: open computer-use weights, graded on the lab's own benchmarks

> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/holo4
> date: 2026-10-02
> tags: explainer, agents, computer-use, open-weights, mixture-of-experts, reinforcement-learning, benchmarks

On 2026-10-02 [H Company](https://hcompany.ai/) released **Holo4**, a family of
generalist computer-use models that, in their words, "click, code, and call
tools across desktop, web, Android, code sandboxes, MCP servers, and business
APIs." Three checkpoints ship with open weights, alongside quantized variants, a
rebuilt agent harness, a described training recipe, and — the part worth the
most — every evaluation trajectory behind the scores.

There is no arXiv paper and no technical report. What exists is a
[blog post](https://hcompany.ai/newsroom/holo4), three Hugging Face model cards,
and a 7,366-run [trajectory dataset](https://huggingface.co/datasets/Hcompany/trajectories).
Every benchmark number is H's own, run in H's own harness. That is not a
disqualifier — most model launches are self-graded — but it changes what you
can conclude from a number, and it is the reason the trajectory release matters
more here than the headline percentages do. This piece takes the mechanism apart
and then checks three specific claims: how much the post-training actually adds
over the Qwen base, how far Holo4 sits behind Opus 5.5 on the hard benchmark, and
what licence each checkpoint actually carries.

## What a computer-use agent is, and how OSWorld scores one

Start from the interface, because it is the whole game. A computer-use agent
does not get a special API into an application. It gets a screenshot. The
harness captures the screen, the model looks at the pixels and emits one action —
a click at a coordinate, some typed text, a line of code to run in a sandbox, or
a tool call — the harness executes it, and the new screen becomes the next
observation. The loop repeats, for hundreds of steps on a long task, until the
model decides it is finished.

<AgentLoop />

This is the same design point as [Qwen-CUA](/articles/qwen-cua), which controls a
machine from screenshots alone, and it is the opposite instinct from a coding
agent's tidy tool table — the subject of [Agent harnesses](/articles/agent-harness).
A shell command can rename 400 files in one call; a GUI agent has to click and
type its way through the same job one primitive at a time. The payoff is that the
interface never goes stale: it works on software that has no API, which is most
software. Holo4 widens the vocabulary past pure pixels — it can also run code and
call MCP tools and business APIs directly — but the observe-decide-act spine is
unchanged.

The grading is the part people skip. [OSWorld](https://github.com/xlang-ai/OSWorld)
is a set of tasks on a real Ubuntu desktop, each with a **programmatic verifier**
that inspects the machine's final state — a file that should exist, a setting
that should have changed, an app in the right state — and returns pass or fail.
The model never sees the verifier. Averaged over the task set, that pass rate is
the OSWorld score. OSWorld 2.0 raises the difficulty to long, multi-step
workflows and reports an average **partial** score as well as strict success, so
its numbers run lower and its costs run higher. AndroidWorld does the same for
phone apps; AutomationBench does it for business automation over MCP tools and
APIs; ALE-CLI (the 105-task Linux split of Agents' Last Exam) does it for expert
terminal workflows. All of these are external benchmark suites. What is
self-reported is the act of running them.

## Three checkpoints, three licences

The family is one recipe applied to three different base models, and the licences
are not the same across them — which is the first thing to get straight, because
"open weights" is doing a lot of work in the announcement.

| Checkpoint | Base model | Architecture | Weights licence |
|---|---|---|---|
| **Holo4-27B** | Qwen3.8-27B | dense | **CC BY-NC 4.0** (non-commercial) |
| **Holo4-35B-A3B** | Qwen3.6-35B-A3B | MoE, 3B active | **Apache 2.0** |
| **Holotron4-30B-A3B** | NVIDIA Nemotron 3 Nano Omni | NemotronH Nano Omni | **NVIDIA Open Model Agreement** |

Each ships in BF16, FP8, NVFP4 and 4-bit GGUF (Holotron4 in BF16 and FP8), with a
262,144-token context. The licences are read straight off the model-card
metadata, and the spread has a sharp edge in it.

The announcement says Holo4 is "available today on the H Models API for
commercial use." True — but that is the hosted API. If you want to **download**
weights and self-host commercially, your only option in this family is
**Holo4-35B-A3B**, the Apache-2.0 MoE. The dense **Holo4-27B** — which, as the
numbers below show, is the stronger model on every benchmark here — is
**CC BY-NC 4.0**: non-commercial. So the strongest model you can run yourself is
closed to commercial self-hosting, and the strongest model you can self-host
commercially is the weaker MoE. The base models invert this, incidentally: both
Qwen3.8-27B and Qwen3.6-35B-A3B are Apache-2.0 upstream, so the non-commercial
clause is H's addition on the 27B, not Alibaba's.

<ModelCard repo="Hcompany/Holo4-35B-A3B" />

Holotron4-30B-A3B is the odd one out: built on NVIDIA's Nemotron rather than Qwen,
governed by the NVIDIA Open Model Agreement, and interesting mostly as a data
point about how far the recipe transports. More on that at the end.

## The recipe: fine-tune, then two RL experts, then merge

<Figure
  src="https://ai.thesatyajit.com/articles/holo4/fig1.png"
  alt="Three-stage training pipeline. Stage 01, Supervised fine-tuning on 127B tokens: about three quarters are successful agentic trajectories from the Agentic Task Factory — desktop 45 percent, web 14 percent, MCP and API 12 percent, mobile 3 percent — with the remaining 26 percent covering multimodal reasoning, GUI grounding, and text-only tool use and coding. Stage 02, Two RL experts: asynchronous online reinforcement learning on long-horizon tasks trains two specialized LoRA experts on the fine-tuned model, one for desktop and web, one for terminal, MCP and API. Stage 03, One merged model: both experts merge back into the fine-tuned model with equal weight and no further training. A flow graph on the right shows token sources converging into the fine-tuned model, splitting into the two experts under 'Online RL', and merging into Holo4."
  caption="The Holo4 training pipeline: a 127B-token SFT stage, two asynchronously-trained LoRA experts, and an equal-weight merge with no further training (Holo4 blog)."
/>

The recipe has three stages, and the interesting choice is in how the last two
fit together.

**Supervised fine-tuning on 127B tokens.** About three quarters of that is
successful agentic trajectories from H's **Agentic Task Factory** — desktop
(45%), web (14%), MCP and API (12%), and mobile (3%), which sums to the "three
quarters" the post claims. The remaining 26% is multimodal reasoning, GUI
grounding, and text-only tool use and coding. The word "successful" is load-
bearing: this is rejection-sampled imitation. The factory generates tasks, an
agent attempts them, and only the runs that the verifier marks as solved become
training data. You are teaching the model from its own wins.

The Agentic Task Factory itself is the part worth stealing. It is a set of
pipelines that "builds interactive environments and verifiable tasks from
documentation alone," and it has produced about **10,000 tasks** — roughly 4,000
web-app tasks, 3,000 MCP-server tasks, and 3,000 desktop and OS tasks. The gate
on a task is strict: it is kept only if its verifier fails on the untouched seed
state, passes on the golden state, rejects every near miss, and an agent actually
solved it through the real interface. That last clause is what keeps the task set
from filling up with puzzles no agent can reach — the same "keep only tasks with
a mix of successes and failures" discipline that [scaling agentic
RL](/articles/scaling-agentic-rl) treats as the real engineering problem behind
an RL fleet.

**Two asynchronous online-RL experts.** On top of the fine-tuned model, H trains
**two specialized LoRA experts** with asynchronous online RL on long-horizon
tasks: one for **desktop and web**, one for **terminal, MCP and API**. Splitting
by workflow class is a bet that the two skill families — pixel-grounded GUI
control versus structured tool and code use — pull the weights in different
enough directions that a single RL run would have them fighting.

**An equal-weight merge, no further training.** Then the twist: "both experts
merge back into the fine-tuned model with equal weight and no further training."
This is weight-space model merging, not a mixture-of-experts router and not a
second distillation pass. Two LoRA deltas, each the product of its own RL run,
are averaged back onto the shared SFT base. It costs nothing beyond the two RL
runs — no joint fine-tune, no router to train — and it is the kind of result that
only works because both adapters started from the same initialization, so their
updates live in comparable coordinates. Whether equal weight is optimal is
exactly the ablation a technical report would carry and this release does not.

## The numbers, and what they actually say

Here is the headline table, transcribed from the blog (percentages are scores;
dollar figures are cost per task, at H Models API rates for Holo4 and Alibaba
Cloud list prices for the Qwen base):

| Benchmark | Holo4-27B | Holo4-35B-A3B | Qwen3.8-27B (base) | Best frontier shown |
|---|---|---|---|---|
| OSWorld | **85.2%** · \$0.08 | 80.8% · \$0.05 | 84.3% · \$0.22 | Qwen3.8 Max 86.1% |
| OSWorld 2.0 | **61.7%** · \$1.22 | 30.9% · \$0.61 | 48.0% · \$3.49 | Opus 5.5 81.8% · \$8.48 |
| ALE-CLI | **44.1%** · \$0.82 | 30.9% · \$0.29 | 43.5% | Opus 5.5 63.7% · \$8.22 |
| AutomationBench | **45.4%** · \$0.05 | 34.5% · \$0.02 | 40.3% · \$0.09 | Opus 5 50.3% · \$3.05 |
| AndroidWorld | **85.1%** · \$0.08 | 77.6% · \$0.07 | 81.9% · \$0.13 | Fable 5 88.8% |

Before reading anything into the cross-model gaps, read H's own footnotes, because
they are honest and they matter. Holo4's scores are the **mean of two to four
runs** in H's harness (a single run on OSWorld 2.0 and ALE-CLI). The frontier
numbers are public scores from other providers, "across different harnesses and
effort levels" — Opus 5.5 at max effort in Anthropic's harness, costs read off
providers' charts. This is not a matched comparison, and H says so. Treat the
Holo4-vs-frontier columns as "roughly where each lands," not a controlled result.

### Claim 1: how much does the post-training add over the base?

This is the claim I most wanted to check, because "improves significantly over
their Qwen base models" is the model card's own phrasing, and on the most-cited
benchmark it is barely true.

On **OSWorld**, Holo4-27B scores 85.2% against the Qwen3.8-27B base's 84.3%. That
is **+0.9 points** (reasoned, from the two reported figures). OSWorld is close to
saturated at this point — three systems in the table sit between 84% and 86% —
and the post-training moves the needle almost not at all on it. What it moves is
the **price**: \$0.08 a task against the base's \$0.22, so the base model costs
about 2.75x as much to reach a slightly lower score (reasoned). The gain on
OSWorld is efficiency, not accuracy.

The accuracy gains show up where the tasks are hard. On **OSWorld 2.0**,
Holo4-27B is at 61.7% against the base's 48.0% — **+13.7 points** (reasoned) — and
on **AndroidWorld** it is +3.2. The sharpest gain is on tool use: on the Agentic
Task Factory's held-out MCP tasks (14 tool servers), Holo4-27B reaches 89.4%
against the base's 74.2%, about +15 points. The two RL experts earn their keep on
long-horizon control and structured tool calling — exactly the two classes they
were specialized on — and contribute almost nothing on the saturated
single-step benchmark. That is a coherent story, and a more precise one than
"improves significantly."

### Claim 2: the price gap to Opus 5.5 on OSWorld 2.0

<CostAccuracy />

OSWorld 2.0 is where the frontier is clearly ahead. Opus 5.5 scores 81.8% against
Holo4-27B's 61.7% — **20.1 points** of average partial score (reasoned) — and on
strict success it is 48.7% against 41.5%. Holo4 does not close that gap; the honest
read is that on the hardest long-horizon benchmark, the best closed model is
meaningfully better.

The lever Holo4 pulls instead is cost. Opus 5.5's OSWorld 2.0 run is reported at
\$8.48 a task; Holo4-27B's at \$1.22 — roughly **7x cheaper** (reasoned,
8.48 ÷ 1.22 ≈ 6.95). So the trade on offer is explicit: give up about 20 points
of score on the hard benchmark to run at a seventh of the price, with the weights
on your own machine and the screenshots never leaving it. On the saturated
OSWorld that trade is nearly free — Holo4-27B matches the frontier within a point
at a fraction of the cost. On OSWorld 2.0 it is a real concession. The scatter
above lets you switch benchmarks and watch the trade change shape; it is the
blog's cost-vs-score framing, redrawn from the same reported numbers.

<Figure
  src="https://ai.thesatyajit.com/articles/holo4/fig2.png"
  alt="Scatter plot titled OSWorld 2.0, average partial score in percent on the vertical axis against USD per task on a log-scale horizontal axis from one cent to one hundred dollars. Holo4-27B is highlighted in violet at about 62 percent and just over one dollar, sitting on the closed-frontier cost-performance line; Holo4-35B-A3B is highlighted lower at about 31 percent near sixty cents. Its Qwen base models sit to the right at higher cost: Qwen3.8 27B near 48 percent at several dollars. Closed frontier models — GPT-6 Astra, Opus 5, GPT-6 Luna, GPT-5.5 — trace a rising line toward the upper right at ten dollars and above."
  caption="OSWorld 2.0 average partial score against cost per task. Holo4-27B sits on the closed-frontier cost-performance line while its base model sits to the right of it at higher cost; the strongest closed models are up and to the right, at several times the price (Holo4 blog)."
/>

One caveat the scatter cannot show: the AutomationBench headline is measured on
600 public tasks, but 480 of those overlap the split H collected training data
from. On the **120 genuinely held-out** tasks, Holo4-27B actually scores *higher*
— 49.3% — and the 35B-A3B scores 31.7% against its own Qwen3.6-35B-A3B base's
13.1%, a +18.6-point jump that is the clearest "the training worked" signal in the
whole release (reasoned, from the footnote's figures). And on ALE-CLI, H discloses
that it excluded a handful of attempts that reached reference answers someone had
left on the test machine — 2 of 105 for the 27B, 1 for the 35B-A3B. Small, but the
kind of thing you only learn because they wrote it down.

### Claim 3: does releasing the trajectories matter?

This is the part I would keep if I had to throw the rest away. H published
[Hcompany/trajectories](https://huggingface.co/datasets/Hcompany/trajectories):
**7,366 agent runs** — every run behind the benchmark scores, from both Holo4
models — with each step's reasoning, actions, tool results and screenshots, plus
the verifier's result and token usage per trajectory. The per-benchmark counts
(1,096 and 1,102 on OSWorld, 106 each on OSWorld 2.0, 1,600 and 1,598 on
AutomationBench, and so on) sum to exactly 7,366 (measured). The layout is plain:
an index with one summary row per run, a JSON file per trajectory, and WebP
screenshots the steps reference. Credentials and personal data are masked.

Why this is the real contribution: when every number is self-reported and there
is no protocol document, a score is a claim you cannot check. A trajectory is a
claim you can **replay**. You can open the OSWorld 2.0 runs, watch where the agent
succeeds and where it loops, confirm the verifier fired on the state it says it
did, and recompute the pass rate yourself. It is the same move
[cua-s1-forms](/articles/cua-s1-forms) rewarded in the small-model world — the
releases that committed per-row predictions were the ones whose claims survived
scrutiny, and the one instability that mattered was only visible because someone
published the raw predictions and the provenance to join them. A leaderboard
number you can only trust; a trajectory you can audit. For a self-graded release
with no report, the 7,366 runs are the report.

## Holotron4: the recipe on someone else's base

The third checkpoint is the stress test. **Holotron4-30B-A3B** applies the same
pipeline to NVIDIA's Nemotron 3 Nano Omni instead of a Qwen model, and the gains
over that base are large where the base was weak:

| Benchmark | Interface | Nemotron 3 Nano Omni | Holotron4-30B-A3B | Gain |
|---|---|---|---|---|
| OSWorld | GUI | 21.0 | 76.3 | +55.3 |
| OSWorld 2.0 | GUI and code | 0.2 | 7.9 | +7.7 |
| AutomationBench | MCP | 19.4 | 35.6 | +16.2 |
| PinchBench | Terminal | 84.7 | 88.6 | +3.9 |
| ALE (Linux, code) | Terminal | 0.6 | 8.5 | +7.9 |

The OSWorld jump from 21.0 to 76.3 is the eye-catching one, but read it the right
way: a base model at 21% on OSWorld was barely a computer-use agent at all, so a
+55.3-point gain is partly a measure of how much headroom there was, not only how
good the recipe is. The absolute Holotron4 OSWorld 2.0 number is 7.9% — the recipe
transports, but a weaker base stays a weaker agent on the hard benchmark. It is a
useful demonstration that the Agentic Task Factory plus the two-expert merge is
not Qwen-specific, and an honest one, because the small absolute numbers are
printed right next to the large deltas.

## The scorecard

Holo4 is a real release with three things in it that are genuinely useful: an
Apache-2.0 MoE you can actually self-host commercially, a described and
reproducible-in-principle recipe (SFT on rejection-sampled trajectories, two
asynchronous RL LoRA experts, an equal-weight merge), and 7,366 evaluation
trajectories that turn self-reported scores into auditable ones. The price story
is real: on the benchmarks that are near saturated, Holo4-27B matches the frontier
within a point at a fraction of the cost, and that is a meaningful shift in what an
open model buys you.

The honest counterweights are three. The strongest checkpoint, Holo4-27B, is
non-commercial — the Apache model is the weaker one, so "open and commercial" and
"strongest" do not point at the same file. The accuracy gain over the Qwen base is
small on the headline benchmark (+0.9 on OSWorld) and concentrated on the harder
long-horizon and tool-calling tasks, which is a narrower and more defensible claim
than the one the model card makes. And on the hardest benchmark, OSWorld 2.0,
Opus 5.5 is still about 20 points ahead — bought, admittedly, at roughly seven
times the cost. All of these are visible in the release because H wrote the
footnotes and shipped the trajectories, which is the behaviour you want from a lab
grading its own exam, and the reason this release is worth taking at close to face
value even without a report.

---

*Built on H Company's [Holo4 announcement](https://hcompany.ai/newsroom/holo4)
(2026-10-02) and the Hugging Face model cards for
[Holo4-27B](https://huggingface.co/Hcompany/Holo4-27B),
[Holo4-35B-A3B](https://huggingface.co/Hcompany/Holo4-35B-A3B) and
[Holotron4-30B-A3B](https://huggingface.co/Hcompany/Holotron4-30B-A3B), plus the
[trajectories dataset](https://huggingface.co/datasets/Hcompany/trajectories).
There is no technical report. Figures are reproduced from the blog for commentary,
flattened onto white. Every benchmark number is H's own, self-reported; the
agent-loop diagram and the cost-vs-score scatter are my own illustrations of the
mechanism and the reported numbers, not measured traces.*
