2026-10-07 · 24 min · agents · harness · benchmarks · mle-bench · multi-agent · agentic-coding
Why read this
Notabletop 60%Traces each viral number to its table and split, shows the multi-agent drop is a selection failure, and widens the paper's seed-only error bars.
- Original analysis
- Widely used
- A lasting reference
Agents & harnessesNeeds datacenter GPUsPractitioner paper
How this was scored
- Is it new?
- 2 of 3: A real new idea, method or capability
- Can I trust it?
- 2 of 3: Measures key facts from files, code or configs
- Can I run it?
- 1 of 3: API-only, gated or restrictive licence
- Will I understand it?
- 2 of 3: Mechanism from first principles with figures
- Can I act on it?
- 2 of 3: A concrete recipe, numbers or comparison
- Will it last?
- 2 of 3: A reference for a year or more
- Does it affect many?
- 2 of 3: A widely used model, tool or lab release
- Only here?
- 2 of 3: A teardown or measurement few others did
Score 65 of 100, ranked 167 of 454 rated articles. Each question is answered 0–3 by hand, and a 3 is rare. How articles are scored
A summary of this paper went round with five crisp claims attached. One coding agent with a shell matched or beat four multi-agent ML systems. Every wrapper was re-run on one codebase, same model, same hardware, same 24 hours. The shell was the only change that clearly mattered. With GLM 5.2 the minimal agent medalled on 62.5% of Kaggle tasks against 47.1% for the best published harness. Adding parallel agents and a message channel dropped it from 55.7% to 33.3%.
I wanted it to be true. This site has written up a lot of harnesses, from AIRA's operator graphs to self-rewriting ones, and a week ago I argued about which half of a scaffold gets eaten by a better base model. A clean, matched experiment saying "most of it" would be useful. So I read the whole paper, appendices included, and checked each of the five claims against the table it came from.
The result mostly holds. The summary does not. The four systems are not the ones people assume, they were not rebuilt in one codebase, the two headline medal rates come from different task splits, and the multi-agent collapse is something more specific and more interesting than "coordination hurts". The error bars also measure less than they appear to.
The paper
How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering? by Kirill Brilliantov, Alejandro Hernández-Cano and Emmanuel Abbé (EPFL and Apple), posted 30 September 2026. No code is released. I looked on the arXiv page and through every link in the HTML; the only repositories cited are OpenCode, which the agent runs on, and MLE-bench's issue tracker. So unlike most pieces here, there is no file and line to quote. Everything below comes from the paper's text and tables.
The question is narrow and well posed. MLE harnesses grew up when the unit of work was one chat completion: the model writes a script, the harness runs it, reads the error, decides what to try next, keeps a tree of candidates, and so on. Modern models are post-trained to do that loop themselves inside a coding agent. So which parts of the harness still do anything?
Their method is to hold everything fixed that usually moves between papers. One backbone at a time, served locally with vLLM. One A100 80GB, 12 CPU cores and 144GB of RAM per agent. A 24-hour budget. The same tasks. Then change one thing.
What Malena actually is
The name is short for "machine learning engineer agent", and the paper calls it a baseline, which undersells how specific it is. It is an OpenCode v1.15.6 session that is allowed to run for the whole 24 hours. When the model ends a turn early, the harness re-prompts it with a short message giving the time left. That loop is the whole orchestration.
The tools are OpenCode's defaults (read, write, edit, task, todowrite)
with three switched off (doom_loop, plan_exit, question), and the native
bash replaced by their own. Their bash matters more than it sounds. It only
runs commands from a short whitelist (pwd, ls*, python*, nvidia-smi and
some others), reports duration and average and peak CPU, GPU, RAM and VRAM after
every command, and runs a watchdog that kills a command shortly before the
machine runs out of memory, so a runaway training job does not get the pod
evicted.
On top of that sit three tools, and these are what the paper means by a minimal harness:
| Tool | What it does |
|---|---|
submissions_register / submissions_list | Adds a submission file to a registry with its validation score and free-form notes. No test score ever comes back. |
system | Hardware, utilisation, and time remaining. |
job_submit, job_wait, job_queue, job_info | Run commands in the background, wait on them, inspect them. With these on, every foreground bash call is capped at 10 minutes. |
The prompt is built from a template shared with the other iterations: a persona, some "scientific-approach questions", the resource and time budget, a few general recommendations, and then Malena's own instructions, which say it may design any number of pipelines and register any number of submissions. Unlike Chat, it is not handed a data summary, only the paths. The prompt text itself is not published, so I can't tell you what "well-prompted" amounts to in words.
The registry is the quiet load-bearing part. At the end of the run, the self-selected answer is the registered submission with the best validation score the agent reported. That makes Malena's choice of final answer depend on its own validation split, and that one design fact explains most of what happens when you add more agents.
The ladder
Section 4 is the controlled half of the paper. Everything here is built inside one codebase, and each rung adds one class of intervention. The authors build "the simplest faithful version" of each idea rather than a copy of a particular published system, which is the right call for attribution and the wrong one if you wanted to know how AIDE itself does.
The bottom rung is Chat: one completion, given the task and a harness-computed
overview of the data, which must return a whole Python script. If it fails, the
error goes back and the model must rewrite the whole script. The next rung is
Oneshot: one coding-agent session, six hours, free to explore the data and run
cheap tests, but allowed exactly one production run, after which it must write
main.py, submission.csv, design.md and score.txt.
That step is the big one, and Figure 2 shows it across nine backbones.

For GLM 5.2, a single Oneshot iteration reaches a 55.15 mean percentile and a 36.8% medal rate (Table 13). The Chat point for the same model sits at about 21% medal rate and about 40 percentile, which I read off the chart because it is not tabulated. The interesting detail is the DeepSeek V4 pair. The Preview releases show nearly no Chat-to-Oneshot gap; the final releases, which the authors say got much more coding-agent post-training on comparable base models, open a clear one. It is the paper's best evidence for its broader claim that post-training ate the harness, and they flag it as suggestive rather than shown.
Above Oneshot, the harness chains sessions at a matched 24 hours under four search rules: Chain (always refine the newest), Greedy (refine the best), UCB1 (an upper-confidence rule over a rank-normalised score with ) and Best-of-N (independent sessions, no sharing). Then comes Malena, which removes the search loop entirely. Then come additions to Malena, then more Malenas.
| base vs this row | Δ medal rate | won A/B/= | sign test p |
|---|---|---|---|
| +D | +10.5 [-1.0, +22.0] | 4/4/6 | 1.000 |
| +P3 | +3.3 [-8.2, +14.9] | 4/3/7 | 1.000 |
| +B+P3 | +22.4 [+10.8, +33.7] * | 7/2/5 | 0.180 |
Δ and its bracket are the paper's paired-by-task bootstrap; * means the bracket excludes zero. The win counts are printed beside each comparison in the paper; the sign test on them (ties dropped) is computed here. It asks whether the result would survive drawing a different set of tasks, which the bootstrap does not.
The widget holds the four tables that matter. Read each tab on its own: the splits differ, so a number in one tab is not comparable with a number in another.
On the search tab, no rule beats any other. The biggest pairwise gap is UCB1 ahead of Best-of-N by 2.71 percentile points, with a 95% interval of −1.44 to +6.60. Malena then beats all four on medal rate in the paired test, by 6.6 to 11.9 points. Its percentile edge is significant against Best-of-N and Chain only.
This is my first disagreement with the summary. "Giving the model a shell was the
only change that clearly mattered" is close to the paper's own framing (the
coding-agent environment is "by far the largest effect we measure"), but it is
not what the tables say. Letting the session run for the whole budget instead of
chaining six-hour sessions also mattered, significantly, on medal rate. And
removing things mattered: Malena without its system and jobs tools loses 8.9
medal points, significant, and adding time nudges at 25/50/75/90% of the run
costs 10.6, also significant. What the paper actually shows is narrower and still
useful: once the agent has a shell and the whole budget, nothing they added
helped. That includes subagent delegation, a Hugging Face search MCP, strict
validation guidelines, eight hand-written ML-engineering skills, and a nudge when
the GPU sits idle. The idle-GPU nudge was the worst of them: 30.4% against 40.3%
in the 14-task, 12-hour runs.
The selection gap column is the other thing to watch on that tab. Best-of-N reaches an oracle percentile of 71.04, as high as Malena's 71.18, but its self-selected number is 64.21. Its best submission was in the pile. It could not tell which one it was. Malena's gap is the smallest of the five at 1.876 points.
Four harnesses, run as they ship
Section 5 is the uncontrolled half, on purpose: it takes released MLE systems and runs them as released, under the same backbone, hardware and budget. The four are not AIDE, ML-Master or R&D-Agent. They are:
- MLEvolve (Du et al., 2026): a UCT tree of complete training scripts with six operators (draft, improve, evolution, debug, fusion, fusion_draft), a hand-designed reward, a schedule that shifts to exploitation after half the budget, a BM25 plus FAISS memory of past attempts, and three candidates running at once on the one GPU. Its unit of work is a chat completion.
- AiScientist (Chen et al., 2026): one conversation stepping through tool calls, with a git repository and two append-only markdown logs as its only memory, and blocking subagents. It is closer to Malena than anything else here.
- Arbor (Jin et al., 2026): a Coordinator that keeps an "Idea Tree" and an Executor that implements each idea in its own git worktree, three executors in parallel, run in its native CLI mode with the timeout and early-stopping defaults from Arbor's own Kaggle plugin.
- ScienceFlow (Zhao et al., 2026): a manifest-driven research agent with parallel exploratory workers sharing one GPU.
So "four multi-agent ML systems" is loose. One is tree search over completions, one is a single conversation, two are genuinely multi-worker. And "re-ran every wrapper on one codebase" mixes up the two halves of the paper: the one-codebase part is the ladder above; these four ran on their own code.
Fidelity is where a comparison like this usually dies, and the authors spend an
appendix on it. They fixed two bugs that broke the baselines on GLM 5.2: MLEvolve
forced a named tool_choice, which threw the server's grammar-constrained
decoder off when the model wanted to reason first, and ScienceFlow sent an empty
tools: [] array, which the deployment rejected. They ran MLEvolve on DeepSeek
V4 Flash at medium reasoning effort instead of max, because at max each call
reasoned about five times longer (37,124 against 7,069 characters), the tree only
grew 18.5 nodes a run instead of 62.8, and the medal rate on a 10-task check went
from 0% to 25.0%. They tried one and two ScienceFlow workers, and longer Arbor
timeouts and no early stopping. None of those moved beyond noise. It is a
reasonable amount of effort spent on someone else's system, and it is more than
most papers report. It is still less than the effort that went into Malena, and
the authors list that asymmetry as a limitation themselves.

The headline numbers are in Table 15, and the summary's pair checks out exactly. With GLM 5.2 on the 30-task split, Malena's self-selected medal rate is 62.5% [58.3, 66.7]. The best external harness by the same measure is AiScientist at 47.1% [42.3, 52.2], with Arbor at 45.8%, MLEvolve at 39.2% and ScienceFlow at 24.4%.
The other backbones are less tidy than "matched or beat" suggests.
- Kimi K3: Malena and Arbor tie on medal rate at 60.0% each. Malena is ahead on percentile, 72.75 against 68.46.
- DeepSeek V4 Flash: Arbor is nominally ahead, 50.0% against 44.4%, and so is AiScientist on percentile. Neither gap is significant, which is what "matches" means here.
- Gemma 4 31B: MLEvolve wins on percentile, 42.88 against 32.35, by a significant 10.53 points. Its medal rate is not significantly different.
That last one is the pattern the authors build their bitter-lesson reading on: hand-built search helps the weak model and stops helping as the model gets better. It is a fair reading of four backbones. It is also four points on a curve, and the ordering of "capability" is the paper's.
NatureBench, a newer set of 40 tasks built from Nature-family papers, is the contamination hedge. There the score is how often the run's best result beats the published state of the art, at 8 hours. With GLM 5.2, Malena gets 21.7% against AiScientist's 17.5% and MLEvolve's 10.8%. With Kimi K3, AiScientist is nominally ahead, 27.1% to 26.7%. Arbor and ScienceFlow were not run on it.
Where 62.5% and 55.7% come from
The summary puts the two numbers in adjacent sentences, which invites reading the second as a drop from the first. They are different experiments.
| Number | Table | Split | Budget | What it is |
|---|---|---|---|---|
| 62.5% | 15 | fixed30, 30 tasks | 24h | Malena, GLM 5.2, self-select |
| 47.1% | 15 | fixed30 | 24h | AiScientist, GLM 5.2, self-select |
| 55.7% | 1 | fixed14, 14 tasks | 24h | Malena base, GLM 5.2, self-select |
| 33.3% | 1 | fixed14 | 24h | +B+P3, GLM 5.2, self-select |
The task sets need a word. MLE-bench has 75 Kaggle competitions. The paper uses 30 of them, following AIRA₂'s 30-task set from the GPT-5 system card with one swap for licensing (5 Lite, 21 Medium, 4 High). The ladder runs on default29, the original MLE-bench tasks minus the swapped pair. The comparison with released harnesses runs on fixed30, where four tasks have their preparation scripts repaired (more on those below). And fixed14 is a "challenging" subset chosen by dropping every task Gemma 4 31B solved and every task Kimi K3 did not.
There is one more 62.5% in the paper, and it is a coincidence worth knowing about if you search the PDF: Malena's oracle medal rate on default29 in Table 9 is also 62.5%. Same agent, same backbone, different split, different selection rule.
Why three agents with a message channel did worse
The multi-agent experiment has three arms. +D turns on
OpenCode's task tool for background subagents and adds a prompt nudge to use it:
the planner and executor pattern. +P3 launches three independent Malena
sessions on the same machine, sharing the one A100. +B+P3 adds a broadcast
tool on top: any agent can post a message, and it lands in each peer's queue at
the end of their current turn.
Switch the widget above to the multi-agent tab, set it to medal rate, and look at the squares rather than the dots.
The three sessions working independently (+P3) are no worse than one: 52.4% against 55.7%, and their oracle, the best submission any of them produced, is 60.7%, level with base. With broadcast, the self-selected rate falls to 33.3%. The oracle only falls to 53.0%. Most of the collapse is in the gap between them. The runs still produced medal-worthy submissions about as often as a lone session. What went wrong is the choice of which one to hand in.
The paper does not dig into the mechanism, and says so ("how to properly coordinate strong agents towards a common goal remains an open question"). But its own appendix on selection has the most plausible explanation. A single session scores all its candidates on one validation split it designed, so its scores are comparable with each other. Independent sessions each build their own split, and the harness then ranks numbers that were never on a common scale. The authors offer that as a hypothesis for Best-of-N's large gap and say they have not tested it. Three sessions writing into one registry inherit the same problem. I would guess broadcast makes it worse by letting a peer's optimistic score pull the others toward the same over-fitted idea, but that is my guess, and nothing in the paper checks it. Shared hardware is the other suspect: three training jobs on one GPU, which the authors note as a contention risk.
Two more things temper the headline drop. The paired test calls it significant, +22.4 medal points with an interval of 10.8 to 33.7. But fixed14 has 14 tasks, and the per-task tally the paper prints is 7 wins for base, 2 for +B+P3 and 5 ties. A sign test on 7 against 2 gives p = 0.18. And the same intervention at 12 hours, with time nudges switched on in both arms, gives +5.2 [−3.7, +14.2], not significant. I believe broadcast hurt. I would not quote 55.7 to 33.3 as the size of the effect.
What the error bars measure
Every interval in the paper comes from the same recipe (Appendix A.5): keep the set of tasks fixed, resample seeds within each task, recompute the macro average, take quantiles. At 24 hours they could afford 3 to 4 seeds per task.
That interval answers "how much would this number move if I re-ran the same 30 tasks with new seeds?" It does not answer "how much would it move on a different 30 tasks of the same kind?", which is what most readers mean when they see "62.5% of Kaggle tasks". Both are legitimate. The paper reports the first and calls the second out only indirectly, in its statistical-power limitation.
Medal outcomes are lumpy. Most tasks are nearly always medalled or nearly never, so seed-to-seed noise is small and task-to-task variation is large. You can estimate how lumpy from the paper's own bracket. With 30 tasks and 4 seeds (the granularity of 62.5% and 65.0% fits 120 runs exactly), [58.3, 66.7] implies that about 76% of a single run's variance is between tasks. Plug that back in with the tasks resampled as well and the interval for 62.5% becomes roughly 47% to 78%.
Try the presets. AiScientist's 47.1% widens to roughly 32% to 63% the same way, so the two marginal intervals overlap heavily once tasks move. The fixed14 numbers widen far more: one task changing its mind moves a 14-task rate by 7.1 points. And on all 75 MLE-bench tasks the same agent would still carry about ±10 points, however many seeds you ran, because only more tasks shrink that part.
None of this sinks the main comparison. Marginal intervals are the wrong tool for comparing two harnesses on the same tasks; pairing by task cancels the difficulty. The paper does pair, and the per-task win counts it prints let you run a test that does let the tasks vary. At GLM 5.2, Malena beats Arbor and AiScientist 12 tasks to 1 each (p = 0.003 on a sign test), MLEvolve 14 to 2 (p = 0.004) and ScienceFlow 19 to 0. Those hold. At Kimi K3, Malena against Arbor is 5 to 5 and against AiScientist and MLEvolve 7 to 2 (p = 0.18 each). The GLM 5.2 result survives a change of tasks. The Kimi K3 medal result is a tie.
My model behind the widget is a simplification. It treats each task as having one medal probability, uses a normal approximation, and assumes 4 seeds when some rows have 3. It gets the order of magnitude right. Treat the exact edges as approximate.
What it does with a day
If Malena is not running a search tree, what is it doing for 24 hours? The trace analysis is the most interesting reading in the paper, and the least tested.

The authors snapshot the code repository at checkpoints (before each long job or submission), have GLM 5.2 label the ML techniques in each snapshot against a 13-way taxonomy, cluster the labels per task with Qwen3-Embedding-8B, and track the mix over the run. Malena builds one general pipeline early and then spends the rest of the day patching it. On the whale task it settles on a two-stage metric-learning embedder plus nearest-neighbour retrieval within its first 17 checkpoints and never changes that architecture over the remaining 63: backbone swaps, resolution bumps, threshold retuning, ensembling. The authors call it hyperparameter search over code patches, which is a good phrase.

The drift toward ensembling late in the run is what a Kaggle competitor does, and nobody told Malena to do it. On the aptos2019 task, its GLM 5.2 and Kimi K3 runs both landed on soft-label pseudo-labelling, a technique no AiScientist, MLEvolve or Oneshot run on that task used. I like this section. I also notice that the labeller and the clusterer's reviewer are GLM 5.2, one of the two models being studied, and that "rarity" is relative to the other runs in the study, not to Kaggle. It is a good description of behaviour, not a measurement of quality.
The contamination check is more convincing than most. Malena's code was compared with 1,350 real Kaggle solutions across the 30 tasks, by token containment and by chunk embeddings. Zero of 924,750 pairs crossed both thresholds, and Malena's upper tail sits below the real-against-real baseline. Combined with the Chat result (a model that memorised the answer would not need the shell), I don't think memorisation explains Malena.
What a single session costs
A single long session is not free. Malena never resets, so every turn resends the whole context, and its cache-read tokens run about an order of magnitude above the others.

At GLM 5.2 prices, Malena's modelled cost is $12.12 a run against AiScientist's $1.95, about 6.2 times more. That figure prices cache reads at a flat tenth of the input price, because none of the harnesses report real cache hits, so treat it as a model. The authors' point stands anyway: 24 hours of an A100 costs between $24 and $85 depending on the provider, so the GPU dominates. If you run this on a pay-per-token API instead of your own vLLM server, the ratio is the number to watch.
Context length is the other worry with one long session. GLM 5.2 runs at its native 1M tokens here. When they cap it at 64k and let OpenCode compact the session whenever it fills (9.5 compactions per run on average, each discarding about 76% of the context), the self-selected medal rate on a 10-task subset falls from 57.5% to 50.0%, with overlapping intervals. The 256k window on Gemma 4 31B is one of their guesses for why the small model prefers Best-of-N.
The four MLE-bench tasks they had to fix
This is a side note, and it is the part I would act on first if I ran MLE-bench. Four of the 30 tasks have preparation-script bugs that the maintainers have documented and chosen not to fix, so as not to invalidate the leaderboard:
- hubmap-kidney-segmentation copies the test images' polygon annotations, which are the target, into the public test folder.
- smartphone-decimeter-2022 deletes
ground_truth.csvbut leaves the reference receiver's NMEA log that the ground truth is interpolated from. - multi-modal-gesture-recognition builds its test archive from a fully labelled training archive, labels included.
- champs-scalar-coupling drops the 3D structures for test molecules that Kaggle actually provides, which makes the task harder than the real one.
The paper prints the diffs. Three of the four are answer keys sitting in the agent's working directory. A strong coding agent that reads every file in its folder, which is the behaviour this paper rewards, will find them. Any MLE-bench number from an agent with a shell should say which version of those tasks it ran. AIRA₂'s own audit found the same class of problem on a different benchmark.
What I'd take from it
I'd start an ML-engineering agent the way Malena is built: one long coding-agent session, a submission registry that records the agent's own validation score, a tool that tells it the time left, a watchdog on memory, and nothing else. On GLM 5.2 that beat four released systems on the same tasks, and the margin survives a test that lets the tasks vary. The harness-effect experiment swapped only the orchestration layer at a fixed model and got cheaper, faster runs at the same quality, and Context Language Models found that letting the model manage its own context beat hand-designed compaction. Three papers is not a law, but it is a direction.
I'd be careful with two readings. The paper does not show "no harness". It shows a small, well-chosen one, and removing two of its three tools cost 8.9 medal points. And it does not show that coordination is useless. It shows that letting three agents broadcast into one registry, without a shared validation protocol, breaks selection. The paper's own data say where the next effort should go: Best-of-N and the broadcast run both made good submissions and then could not pick them. Picking is a validation-design problem, and the one intervention that shrank the gap (a shared baseline workspace with a fixed split) also cut the ceiling. Nobody has solved it.
The result applies to strong models. With Gemma 4 31B the old tree search still wins on percentile, and the scaffolding-gets-eaten argument applies with the same caveat: it is eaten above a capability threshold, and you have to know which side of it you are on.
How I checked
I read the arXiv HTML of 2609.40303v1 end to end, including Appendices A to G, and took every number above from its tables (1, 2, 5, 6, 9, 10, 12, 13, 15, 16, 17, 21, 22) or its text. The Chat medal rate for GLM 5.2 is the one figure I read off a chart, and I say so where it appears. The figures are the paper's own SVGs, rendered to PNG and flattened onto white. There is no released code to read, and neither the X post's thread nor its replies pointed to any. The sign-test p-values are mine, computed exactly from the per-task win counts the paper prints, with ties dropped. The task-resampled intervals are mine too, from a simple model whose between-task share I backed out of the paper's own printed brackets; the widget's comments state the formula. I did not re-run anything.