# Malena: how little harness a strong ML-engineering agent needs

> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/minimal-ml-engineering-agent
> date: 2026-10-07
> tags: agents, harness, benchmarks, mle-bench, multi-agent, agentic-coding

A summary of this paper went round with five crisp claims attached. One coding
agent with a shell matched or beat four multi-agent ML systems. Every wrapper
was re-run on one codebase, same model, same hardware, same 24 hours. The shell
was the only change that clearly mattered. With GLM 5.2 the minimal agent
medalled on 62.5% of Kaggle tasks against 47.1% for the best published harness.
Adding parallel agents and a message channel dropped it from 55.7% to 33.3%.

I wanted it to be true. This site has written up a lot of harnesses, from
[AIRA's operator graphs](/articles/aira-research-agents) to
[self-rewriting ones](/articles/rrsi-harness-evolution), and a week ago I argued
about [which half of a scaffold gets eaten](/articles/scaffolding-gets-eaten) by
a better base model. A clean, matched experiment saying "most of it" would be
useful. So I read the whole paper, appendices included, and checked each of the
five claims against the table it came from.

The result mostly holds. The summary does not. The four systems are not the
ones people assume, they were not rebuilt in one codebase, the two headline
medal rates come from different task splits, and the multi-agent collapse is
something more specific and more interesting than "coordination hurts". The
error bars also measure less than they appear to.

## The paper

[*How Much of a Harness Does a Strong Agent Need for Autonomous ML
Engineering?*](https://arxiv.org/abs/2609.40303) by Kirill Brilliantov, Alejandro
Hernández-Cano and Emmanuel Abbé (EPFL and Apple), posted 30 September 2026.
No code is released. I looked on the arXiv page and through every link in the
HTML; the only repositories cited are [OpenCode](https://github.com/anomalyco/opencode),
which the agent runs on, and MLE-bench's issue tracker. So unlike most pieces
here, there is no file and line to quote. Everything below comes from the paper's
text and tables.

The question is narrow and well posed. MLE harnesses grew up when the unit of
work was one chat completion: the model writes a script, the harness runs it,
reads the error, decides what to try next, keeps a tree of candidates, and so on.
Modern models are post-trained to do that loop themselves inside a coding agent.
So which parts of the harness still do anything?

Their method is to hold everything fixed that usually moves between papers. One
backbone at a time, served locally with vLLM. One A100 80GB, 12 CPU cores and
144GB of RAM per agent. A 24-hour budget. The same tasks. Then change one thing.

## What Malena actually is

The name is short for "machine learning engineer agent", and the paper calls it
a baseline, which undersells how specific it is. It is an OpenCode v1.15.6
session that is allowed to run for the whole 24 hours. When the model ends a
turn early, the harness re-prompts it with a short message giving the time left.
That loop is the whole orchestration.

The tools are OpenCode's defaults (`read`, `write`, `edit`, `task`, `todowrite`)
with three switched off (`doom_loop`, `plan_exit`, `question`), and the native
`bash` replaced by their own. Their `bash` matters more than it sounds. It only
runs commands from a short whitelist (`pwd`, `ls*`, `python*`, `nvidia-smi` and
some others), reports duration and average and peak CPU, GPU, RAM and VRAM after
every command, and runs a watchdog that kills a command shortly before the
machine runs out of memory, so a runaway training job does not get the pod
evicted.

On top of that sit three tools, and these are what the paper means by a minimal
harness:

| Tool | What it does |
|---|---|
| `submissions_register` / `submissions_list` | Adds a submission file to a registry with its validation score and free-form notes. No test score ever comes back. |
| `system` | Hardware, utilisation, and time remaining. |
| `job_submit`, `job_wait`, `job_queue`, `job_info` | Run commands in the background, wait on them, inspect them. With these on, every foreground `bash` call is capped at 10 minutes. |

The prompt is built from a template shared with the other iterations: a persona,
some "scientific-approach questions", the resource and time budget, a few general
recommendations, and then Malena's own instructions, which say it may design any
number of pipelines and register any number of submissions. Unlike Chat, it is
not handed a data summary, only the paths. The prompt text itself is not
published, so I can't tell you what "well-prompted" amounts to in words.

The registry is the quiet load-bearing part. At the end of the run, the
self-selected answer is the registered submission with the best validation score
the agent reported. That makes Malena's choice of final answer depend on its own
validation split, and that one design fact explains most of what happens when
you add more agents.

## The ladder

Section 4 is the controlled half of the paper. Everything here is built inside
one codebase, and each rung adds one class of intervention. The authors build
"the simplest faithful version" of each idea rather than a copy of a particular
published system, which is the right call for attribution and the wrong one if
you wanted to know how AIDE itself does.

The bottom rung is **Chat**: one completion, given the task and a harness-computed
overview of the data, which must return a whole Python script. If it fails, the
error goes back and the model must rewrite the whole script. The next rung is
**Oneshot**: one coding-agent session, six hours, free to explore the data and run
cheap tests, but allowed exactly one production run, after which it must write
`main.py`, `submission.csv`, `design.md` and `score.txt`.

That step is the big one, and Figure 2 shows it across nine backbones.

<Figure
  src="https://ai.thesatyajit.com/articles/minimal-ml-engineering-agent/fig2.png"
  alt="Two dot plots side by side, medal rate on the left and percentile on the right, with nine backbones on the x-axis from GPT OSS 120B to Kimi K3. Each backbone has a blue Chat pair and an orange Oneshot pair, circles for one iteration and squares for best-of-N at six hours. For the weak models the pairs overlap near ten percent medal rate. For DeepSeek V4 Flash, DeepSeek V4 Pro, GLM 5.2 and Kimi K3 the orange Oneshot points sit clearly above the blue Chat points."
  caption="Chat against Oneshot on the default29 split. The gap opens only for models with heavy agentic post-training; the DeepSeek V4 Preview releases show almost none, their final releases a clear one (paper, Figure 2)."
/>

For GLM 5.2, a single Oneshot iteration reaches a 55.15 mean percentile and a 36.8%
medal rate (Table 13). The Chat point for the same model sits at about 21% medal
rate and about 40 percentile, which I read off the chart because it is not
tabulated. The interesting detail is the DeepSeek V4 pair. The Preview releases
show nearly no Chat-to-Oneshot gap; the final releases, which the authors say got
much more coding-agent post-training on comparable base models, open a clear one.
It is the paper's best evidence for its broader claim that post-training ate
the harness, and they flag it as suggestive rather than shown.

Above Oneshot, the harness chains sessions at a matched 24 hours under four
search rules: **Chain** (always refine the newest), **Greedy** (refine the best),
**UCB1** (an upper-confidence rule over a rank-normalised score with
$c = 0.75$) and **Best-of-N** (independent sessions, no sharing). Then comes
Malena, which removes the search loop entirely. Then come additions to Malena,
then more Malenas.

<HarnessLadder />

The widget holds the four tables that matter. Read each tab on its own: the
splits differ, so a number in one tab is not comparable with a number in another.

On the search tab, no rule beats any other. The biggest pairwise gap is UCB1
ahead of Best-of-N by 2.71 percentile points, with a 95% interval of −1.44 to
+6.60. Malena then beats all four on medal rate in the paired test, by 6.6 to 11.9
points. Its percentile edge is significant against Best-of-N and Chain only.

This is my first disagreement with the summary. "Giving the model a shell was the
only change that clearly mattered" is close to the paper's own framing (the
coding-agent environment is "by far the largest effect we measure"), but it is
not what the tables say. Letting the session run for the whole budget instead of
chaining six-hour sessions also mattered, significantly, on medal rate. And
removing things mattered: Malena without its `system` and jobs tools loses 8.9
medal points, significant, and adding time nudges at 25/50/75/90% of the run
costs 10.6, also significant. What the paper actually shows is narrower and still
useful: once the agent has a shell and the whole budget, nothing they *added*
helped. That includes subagent delegation, a Hugging Face search MCP, strict
validation guidelines, eight hand-written ML-engineering skills, and a nudge when
the GPU sits idle. The idle-GPU nudge was the worst of them: 30.4% against 40.3%
in the 14-task, 12-hour runs.

The selection gap column is the other thing to watch on that tab. Best-of-N
reaches an oracle percentile of 71.04, as high as Malena's 71.18, but its
self-selected number is 64.21. Its best submission was in the pile. It could not
tell which one it was. Malena's gap is the smallest of the five at 1.876 points.

## Four harnesses, run as they ship

Section 5 is the uncontrolled half, on purpose: it takes released MLE systems
and runs them as released, under the same backbone, hardware and budget. The four
are not AIDE, ML-Master or R&D-Agent. They are:

- **MLEvolve** (Du et al., 2026): a UCT tree of complete training scripts with
  six operators (draft, improve, evolution, debug, fusion, fusion_draft), a
  hand-designed reward, a schedule that shifts to exploitation after half the
  budget, a BM25 plus FAISS memory of past attempts, and three candidates running
  at once on the one GPU. Its unit of work is a chat completion.
- **AiScientist** (Chen et al., 2026): one conversation stepping through tool
  calls, with a git repository and two append-only markdown logs as its only
  memory, and blocking subagents. It is closer to Malena than anything else here.
- **Arbor** (Jin et al., 2026): a Coordinator that keeps an "Idea Tree" and an
  Executor that implements each idea in its own git worktree, three executors in
  parallel, run in its native CLI mode with the timeout and early-stopping
  defaults from Arbor's own Kaggle plugin.
- **ScienceFlow** (Zhao et al., 2026): a manifest-driven research agent with
  parallel exploratory workers sharing one GPU.

So "four multi-agent ML systems" is loose. One is tree search over completions,
one is a single conversation, two are genuinely multi-worker. And "re-ran every
wrapper on one codebase" mixes up the two halves of the paper: the one-codebase
part is the ladder above; these four ran on their own code.

Fidelity is where a comparison like this usually dies, and the authors spend an
appendix on it. They fixed two bugs that broke the baselines on GLM 5.2: MLEvolve
forced a named `tool_choice`, which threw the server's grammar-constrained
decoder off when the model wanted to reason first, and ScienceFlow sent an empty
`tools: []` array, which the deployment rejected. They ran MLEvolve on DeepSeek
V4 Flash at medium reasoning effort instead of max, because at max each call
reasoned about five times longer (37,124 against 7,069 characters), the tree only
grew 18.5 nodes a run instead of 62.8, and the medal rate on a 10-task check went
from 0% to 25.0%. They tried one and two ScienceFlow workers, and longer Arbor
timeouts and no early stopping. None of those moved beyond noise. It is a
reasonable amount of effort spent on someone else's system, and it is more than
most papers report. It is still less than the effort that went into Malena, and
the authors list that asymmetry as a limitation themselves.

<Figure
  src="https://ai.thesatyajit.com/articles/minimal-ml-engineering-agent/fig1.png"
  alt="Two dot-and-whisker panels, medal rate above and percentile below, grouped by backbone: Kimi K3, GLM 5.2, DeepSeek V4 Flash and Gemma 4 31B. Colours are Malena blue, Arbor orange, AiScientist green, MLEvolve purple, ScienceFlow red; filled circles are self-select and open squares oracle. At GLM 5.2 the blue Malena points sit near 62 to 65 percent medal rate, well above every other harness. At DeepSeek V4 Flash, Arbor is slightly above Malena. At Gemma 4 31B every harness is low and MLEvolve is highest on percentile."
  caption="Malena against the released harnesses on fixed30 at 24 hours, four backbones. Brackets are 95% bootstrap intervals that resample seeds with the tasks held fixed (paper, Figure 1)."
/>

The headline numbers are in Table 15, and the summary's pair checks out exactly.
With GLM 5.2 on the 30-task split, Malena's self-selected medal rate is 62.5% [58.3,
66.7]. The best external harness by the same measure is AiScientist at 47.1%
[42.3, 52.2], with Arbor at 45.8%, MLEvolve at 39.2% and ScienceFlow at 24.4%.

The other backbones are less tidy than "matched or beat" suggests.

- Kimi K3: Malena and Arbor tie on medal rate at 60.0% each. Malena is ahead
  on percentile, 72.75 against 68.46.
- DeepSeek V4 Flash: Arbor is nominally ahead, 50.0% against 44.4%, and so is
  AiScientist on percentile. Neither gap is significant, which is what "matches"
  means here.
- Gemma 4 31B: MLEvolve wins on percentile, 42.88 against 32.35, by a
  significant 10.53 points. Its medal rate is not significantly different.

That last one is the pattern the authors build their bitter-lesson reading on:
hand-built search helps the weak model and stops helping as the model gets
better. It is a fair reading of four backbones. It is also four points on a
curve, and the ordering of "capability" is the paper's.

NatureBench, a newer set of 40 tasks built from Nature-family papers, is the
contamination hedge. There the score is how often the run's best result beats
the published state of the art, at 8 hours. With GLM 5.2, Malena gets 21.7%
against AiScientist's 17.5% and MLEvolve's 10.8%. With Kimi K3, AiScientist is
nominally ahead, 27.1% to 26.7%. Arbor and ScienceFlow were not run on it.

## Where 62.5% and 55.7% come from

The summary puts the two numbers in adjacent sentences, which invites reading the
second as a drop from the first. They are different experiments.

| Number | Table | Split | Budget | What it is |
|---|---|---|---|---|
| 62.5% | 15 | fixed30, 30 tasks | 24h | Malena, GLM 5.2, self-select |
| 47.1% | 15 | fixed30 | 24h | AiScientist, GLM 5.2, self-select |
| 55.7% | 1 | fixed14, 14 tasks | 24h | Malena base, GLM 5.2, self-select |
| 33.3% | 1 | fixed14 | 24h | +B+P3, GLM 5.2, self-select |

The task sets need a word. MLE-bench has 75 Kaggle competitions. The paper uses 30
of them, following AIRA₂'s 30-task set from the GPT-5 system card with one swap
for licensing (5 Lite, 21 Medium, 4 High). The ladder runs on **default29**, the
original MLE-bench tasks minus the swapped pair. The comparison with released
harnesses runs on **fixed30**, where four tasks have their preparation scripts
repaired (more on those below). And **fixed14** is a "challenging" subset chosen
by dropping every task Gemma 4 31B solved and every task Kimi K3 did not.

There is one more 62.5% in the paper, and it is a coincidence worth knowing about
if you search the PDF: Malena's *oracle* medal rate on default29 in Table 9 is
also 62.5%. Same agent, same backbone, different split, different selection rule.

## Why three agents with a message channel did worse

The multi-agent experiment has three arms. **+D** turns on
OpenCode's `task` tool for background subagents and adds a prompt nudge to use it:
the planner and executor pattern. **+P3** launches three independent Malena
sessions on the same machine, sharing the one A100. **+B+P3** adds a `broadcast`
tool on top: any agent can post a message, and it lands in each peer's queue at
the end of their current turn.

Switch the widget above to the multi-agent tab, set it to medal rate, and look at
the squares rather than the dots.

The three sessions working independently (+P3) are no worse than one: 52.4% against
55.7%, and their oracle, the best submission any of them produced, is 60.7%,
level with base. With broadcast, the self-selected rate falls to 33.3%. The oracle
only falls to 53.0%. Most of the collapse is in the gap between them. The runs
still produced medal-worthy submissions about as often as a lone session. What
went wrong is the choice of which one to hand in.

The paper does not dig into the mechanism, and says so ("how to properly coordinate
strong agents towards a common goal remains an open question"). But its own
appendix on selection has the most plausible explanation. A single session scores
all its candidates on one validation split it designed, so its scores are
comparable with each other. Independent sessions each build their own split,
and the harness then ranks numbers that were never on a common scale. The authors
offer that as a hypothesis for Best-of-N's large gap and say they have not tested
it. Three sessions writing into one registry inherit the same problem. I would
guess broadcast makes it worse by letting a peer's optimistic score pull the
others toward the same over-fitted idea, but that is my guess, and nothing in
the paper checks it. Shared hardware is the other suspect: three training jobs on
one GPU, which the authors note as a contention risk.

Two more things temper the headline drop. The paired test calls it significant,
+22.4 medal points with an interval of 10.8 to 33.7. But fixed14 has 14 tasks, and
the per-task tally the paper prints is 7 wins for base, 2 for +B+P3 and 5 ties. A
sign test on 7 against 2 gives p = 0.18. And the same intervention at 12 hours,
with time nudges switched on in both arms, gives +5.2 [−3.7, +14.2], not
significant. I believe broadcast hurt. I would not quote 55.7 to 33.3 as the size
of the effect.

## What the error bars measure

Every interval in the paper comes from the same recipe (Appendix A.5): keep the
set of tasks fixed, resample seeds within each task, recompute the macro average,
take quantiles. At 24 hours they could afford 3 to 4 seeds per task.

That interval answers "how much would this number move if I re-ran the same 30
tasks with new seeds?" It does not answer "how much would it move on a different
30 tasks of the same kind?", which is what most readers mean when they see
"62.5% of Kaggle tasks". Both are legitimate. The paper reports the first and
calls the second out only indirectly, in its statistical-power limitation.

Medal outcomes are lumpy. Most tasks are nearly always medalled or nearly never,
so seed-to-seed noise is small and task-to-task variation is large. You can
estimate how lumpy from the paper's own bracket. With 30 tasks and 4 seeds (the
granularity of 62.5% and 65.0% fits 120 runs exactly), [58.3, 66.7] implies that
about 76% of a single run's variance is between tasks. Plug that back in with the
tasks resampled as well and the interval for 62.5% becomes roughly 47% to 78%.

<IntervalExplorer />

Try the presets. AiScientist's 47.1% widens to roughly 32% to 63% the same way, so
the two marginal intervals overlap heavily once tasks move. The fixed14 numbers
widen far more: one task changing its mind moves a 14-task rate by 7.1 points.
And on all 75 MLE-bench tasks the same agent would still carry about ±10 points,
however many seeds you ran, because only more tasks shrink that part.

None of this sinks the main comparison. Marginal
intervals are the wrong tool for comparing two harnesses on the same tasks;
pairing by task cancels the difficulty. The paper does pair, and the per-task win
counts it prints let you run a test that does let the tasks vary. At GLM 5.2,
Malena beats Arbor and AiScientist 12 tasks to 1 each (p = 0.003 on a sign test),
MLEvolve 14 to 2 (p = 0.004) and ScienceFlow 19 to 0. Those hold. At Kimi K3,
Malena against Arbor is 5 to 5 and against AiScientist and MLEvolve 7 to 2
(p = 0.18 each). The GLM 5.2 result survives a change of tasks. The Kimi K3 medal
result is a tie.

My model behind the widget is a simplification. It treats each task as having one
medal probability, uses a normal approximation, and assumes 4 seeds when some
rows have 3. It gets the order of magnitude right. Treat the exact edges as
approximate.

## What it does with a day

If Malena is not running a search tree, what is it doing for 24 hours? The trace
analysis is the most interesting reading in the paper, and the least tested.

<Figure
  src="https://ai.thesatyajit.com/articles/minimal-ml-engineering-agent/fig3.png"
  alt="Line chart of oracle best-so-far percentile against wall-clock time from 0 to 24 hours for four harnesses with shaded bands. All rise steeply in the first four hours. Malena, in blue, plateaus near 60 by hour 5 and creeps up to about 72 by hour 24. MLEvolve in purple ends near 65, Arbor in orange and AiScientist in green near 62 to 63, with AiScientist the slowest early."
  caption="Best submission so far against wall-clock time, GLM 5.2, fixed30. Every harness does most of its work in the first five to ten hours; later gains arrive as occasional steps (paper, Figure 3)."
/>

The authors snapshot the code repository at checkpoints (before each long job or
submission), have GLM 5.2 label the ML techniques in each snapshot against a
13-way taxonomy, cluster the labels per task with Qwen3-Embedding-8B, and track
the mix over the run. Malena builds one general pipeline early and then spends
the rest of the day patching it. On the whale task it settles on a two-stage
metric-learning embedder plus nearest-neighbour retrieval within its first 17
checkpoints and never changes that architecture over the remaining 63: backbone
swaps, resolution bumps, threshold retuning, ensembling. The authors call it
hyperparameter search over code patches, which is a good phrase.

<Figure
  src="https://ai.thesatyajit.com/articles/minimal-ml-engineering-agent/fig5.png"
  alt="Three stacked area charts, AiScientist, Malena and MLEvolve, showing the share of technique categories from the start of a run to the end. In Malena the pink ensembling and brown model-selection bands grow steadily and take about 60 percent of the share by the end. AiScientist shows a milder version of the same drift. MLEvolve's mix stays roughly flat after the first tenth of the run, dominated by optimisation and architecture."
  caption="Technique mix over a run, pooled over 10 tasks. Malena drifts toward ensembling, model selection and post-processing; MLEvolve's mix settles early and stays put (paper, Figure 5)."
/>

The drift toward ensembling late in the run is what a Kaggle competitor does,
and nobody told Malena to do it. On the aptos2019 task, its GLM 5.2 and Kimi K3
runs both landed on soft-label pseudo-labelling, a technique no AiScientist,
MLEvolve or Oneshot run on that task used. I like this section. I also notice that
the labeller and the clusterer's reviewer are GLM 5.2, one of the two models
being studied, and that "rarity" is relative to the other runs in the study, not
to Kaggle. It is a good description of behaviour, not a measurement of quality.

The contamination check is more convincing than most. Malena's code was compared
with 1,350 real Kaggle solutions across the 30 tasks, by token containment and
by chunk embeddings. Zero of 924,750 pairs crossed both thresholds, and Malena's
upper tail sits below the real-against-real baseline. Combined with the Chat
result (a model that memorised the answer would not need the shell), I don't
think memorisation explains Malena.

## What a single session costs

A single long session is not free. Malena never resets, so every turn resends the
whole context, and its cache-read tokens run about an order of magnitude above
the others.

<Figure
  src="https://ai.thesatyajit.com/articles/minimal-ml-engineering-agent/fig4.png"
  alt="Box plots on a log scale of tokens per run for input, output and cache-read tokens, plus inference cost in dollars on a right-hand axis, for Malena, Arbor, AiScientist and MLEvolve. Input and output tokens are similar across the agentic harnesses, with MLEvolve highest. Malena's cache-read tokens are around ten to the eighth, roughly ten times the others. Malena's median cost is a few dollars with a long tail to about 100 dollars."
  caption="Tokens and modelled cost per 24-hour run at GLM 5.2 prices. Malena's cost is driven by cache reads from one ever-growing session; MLEvolve has none because each tree node is an independent call (paper, Figure 3, right panel)."
/>

At GLM 5.2 prices, Malena's modelled cost is \$12.12 a run against AiScientist's
\$1.95, about 6.2 times more. That figure prices cache reads at a flat tenth of
the input price, because none of the harnesses report real cache hits, so treat
it as a model. The authors' point stands anyway: 24 hours of an A100 costs
between \$24 and \$85 depending on the provider, so the GPU dominates. If you run
this on a pay-per-token API instead of your own vLLM server, the ratio is the
number to watch.

Context length is the other worry with one long session. GLM 5.2 runs at its
native 1M tokens here. When they cap it at 64k and let OpenCode compact the
session whenever it fills (9.5 compactions per run on average, each discarding
about 76% of the context), the self-selected medal rate on a 10-task subset falls
from 57.5% to 50.0%, with overlapping intervals. The 256k window on Gemma 4 31B
is one of their guesses for why the small model prefers Best-of-N.

## The four MLE-bench tasks they had to fix

This is a side note, and it is the part I would act on first if I ran MLE-bench.
Four of the 30 tasks have preparation-script bugs that the maintainers have
documented and chosen not to fix, so as not to invalidate the leaderboard:

- **hubmap-kidney-segmentation** copies the test images' polygon annotations,
  which are the target, into the public test folder.
- **smartphone-decimeter-2022** deletes `ground_truth.csv` but leaves the
  reference receiver's NMEA log that the ground truth is interpolated from.
- **multi-modal-gesture-recognition** builds its test archive from a fully
  labelled training archive, labels included.
- **champs-scalar-coupling** drops the 3D structures for test molecules that
  Kaggle actually provides, which makes the task harder than the real one.

The paper prints the diffs. Three of the four are answer keys sitting in the
agent's working directory. A strong coding agent that reads every file in its
folder, which is the behaviour this paper rewards, will find them. Any MLE-bench
number from an agent with a shell should say which version of those tasks it ran.
[AIRA₂'s own audit](/articles/aira-research-agents) found the same class of
problem on a different benchmark.

## What I'd take from it

I'd start an ML-engineering agent the way Malena is built: one long coding-agent
session, a submission registry that records the agent's own validation score, a
tool that tells it the time left, a watchdog on memory, and nothing else. On
GLM 5.2 that beat four released systems on the same tasks, and the margin
survives a test that lets the tasks vary. The [harness-effect
experiment](/articles/harness-effect) swapped only the orchestration layer at a
fixed model and got cheaper, faster runs at the same quality, and
[Context Language Models](/articles/context-language-models) found that letting
the model manage its own context beat hand-designed compaction. Three papers is
not a law, but it is a direction.

I'd be careful with two readings. The paper does not show "no harness". It shows a
small, well-chosen one, and removing two of its three tools cost 8.9 medal points.
And it does not show that coordination is useless. It shows that letting three
agents broadcast into one registry, without a shared validation protocol, breaks
selection. The paper's own data say where the next effort should go: Best-of-N
and the broadcast run both made good submissions and then could not pick them.
Picking is a validation-design problem, and the one intervention that shrank the gap
(a shared baseline workspace with a fixed split) also cut the ceiling. Nobody has
solved it.

The result applies to strong models. With Gemma 4 31B the old tree search still
wins on percentile, and the [scaffolding-gets-eaten
argument](/articles/scaffolding-gets-eaten) applies with the same caveat: it is
eaten above a capability threshold, and you have to know which side of it you are on.

## How I checked

I read the arXiv HTML of 2609.40303v1 end to end, including Appendices A to G, and
took every number above from its tables (1, 2, 5, 6, 9, 10, 12, 13, 15, 16, 17, 21,
22) or its text. The Chat medal rate for GLM 5.2 is the one figure I read off a
chart, and I say so where it appears. The figures are the paper's own SVGs,
rendered to PNG and flattened onto white. There is no released code to read, and
neither the X post's thread nor its replies pointed to any. The sign-test
p-values are mine, computed exactly from the per-task win counts the paper prints,
with ties dropped. The task-resampled intervals are mine too, from a simple model
whose between-task share I backed out of the paper's own printed brackets; the
widget's comments state the formula. I did not re-run anything.
