~/satyajit

Malena: how little harness a strong ML-engineering agent needs

mdjsonmcp

2026-10-07 · 24 min · agents · harness · benchmarks · mle-bench · multi-agent · agentic-coding

Why read this

Notabletop 60%

Traces each viral number to its table and split, shows the multi-agent drop is a selection failure, and widens the paper's seed-only error bars.

  • Original analysis
  • Widely used
  • A lasting reference

Agents & harnessesNeeds datacenter GPUsPractitioner paper

How this was scored
Is it new?
2 of 3: A real new idea, method or capability
Can I trust it?
2 of 3: Measures key facts from files, code or configs
Can I run it?
1 of 3: API-only, gated or restrictive licence
Will I understand it?
2 of 3: Mechanism from first principles with figures
Can I act on it?
2 of 3: A concrete recipe, numbers or comparison
Will it last?
2 of 3: A reference for a year or more
Does it affect many?
2 of 3: A widely used model, tool or lab release
Only here?
2 of 3: A teardown or measurement few others did

Score 65 of 100, ranked 167 of 454 rated articles. Each question is answered 0–3 by hand, and a 3 is rare. How articles are scored

A summary of this paper went round with five crisp claims attached. One coding agent with a shell matched or beat four multi-agent ML systems. Every wrapper was re-run on one codebase, same model, same hardware, same 24 hours. The shell was the only change that clearly mattered. With GLM 5.2 the minimal agent medalled on 62.5% of Kaggle tasks against 47.1% for the best published harness. Adding parallel agents and a message channel dropped it from 55.7% to 33.3%.

I wanted it to be true. This site has written up a lot of harnesses, from AIRA's operator graphs to self-rewriting ones, and a week ago I argued about which half of a scaffold gets eaten by a better base model. A clean, matched experiment saying "most of it" would be useful. So I read the whole paper, appendices included, and checked each of the five claims against the table it came from.

The result mostly holds. The summary does not. The four systems are not the ones people assume, they were not rebuilt in one codebase, the two headline medal rates come from different task splits, and the multi-agent collapse is something more specific and more interesting than "coordination hurts". The error bars also measure less than they appear to.

The paper

How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering? by Kirill Brilliantov, Alejandro Hernández-Cano and Emmanuel Abbé (EPFL and Apple), posted 30 September 2026. No code is released. I looked on the arXiv page and through every link in the HTML; the only repositories cited are OpenCode, which the agent runs on, and MLE-bench's issue tracker. So unlike most pieces here, there is no file and line to quote. Everything below comes from the paper's text and tables.

The question is narrow and well posed. MLE harnesses grew up when the unit of work was one chat completion: the model writes a script, the harness runs it, reads the error, decides what to try next, keeps a tree of candidates, and so on. Modern models are post-trained to do that loop themselves inside a coding agent. So which parts of the harness still do anything?

Their method is to hold everything fixed that usually moves between papers. One backbone at a time, served locally with vLLM. One A100 80GB, 12 CPU cores and 144GB of RAM per agent. A 24-hour budget. The same tasks. Then change one thing.

What Malena actually is

The name is short for "machine learning engineer agent", and the paper calls it a baseline, which undersells how specific it is. It is an OpenCode v1.15.6 session that is allowed to run for the whole 24 hours. When the model ends a turn early, the harness re-prompts it with a short message giving the time left. That loop is the whole orchestration.

The tools are OpenCode's defaults (read, write, edit, task, todowrite) with three switched off (doom_loop, plan_exit, question), and the native bash replaced by their own. Their bash matters more than it sounds. It only runs commands from a short whitelist (pwd, ls*, python*, nvidia-smi and some others), reports duration and average and peak CPU, GPU, RAM and VRAM after every command, and runs a watchdog that kills a command shortly before the machine runs out of memory, so a runaway training job does not get the pod evicted.

On top of that sit three tools, and these are what the paper means by a minimal harness:

ToolWhat it does
submissions_register / submissions_listAdds a submission file to a registry with its validation score and free-form notes. No test score ever comes back.
systemHardware, utilisation, and time remaining.
job_submit, job_wait, job_queue, job_infoRun commands in the background, wait on them, inspect them. With these on, every foreground bash call is capped at 10 minutes.

The prompt is built from a template shared with the other iterations: a persona, some "scientific-approach questions", the resource and time budget, a few general recommendations, and then Malena's own instructions, which say it may design any number of pipelines and register any number of submissions. Unlike Chat, it is not handed a data summary, only the paths. The prompt text itself is not published, so I can't tell you what "well-prompted" amounts to in words.

The registry is the quiet load-bearing part. At the end of the run, the self-selected answer is the registered submission with the best validation score the agent reported. That makes Malena's choice of final answer depend on its own validation split, and that one design fact explains most of what happens when you add more agents.

The ladder

Section 4 is the controlled half of the paper. Everything here is built inside one codebase, and each rung adds one class of intervention. The authors build "the simplest faithful version" of each idea rather than a copy of a particular published system, which is the right call for attribution and the wrong one if you wanted to know how AIDE itself does.

The bottom rung is Chat: one completion, given the task and a harness-computed overview of the data, which must return a whole Python script. If it fails, the error goes back and the model must rewrite the whole script. The next rung is Oneshot: one coding-agent session, six hours, free to explore the data and run cheap tests, but allowed exactly one production run, after which it must write main.py, submission.csv, design.md and score.txt.

That step is the big one, and Figure 2 shows it across nine backbones.

Two dot plots side by side, medal rate on the left and percentile on the right, with nine backbones on the x-axis from GPT OSS 120B to Kimi K3. Each backbone has a blue Chat pair and an orange Oneshot pair, circles for one iteration and squares for best-of-N at six hours. For the weak models the pairs overlap near ten percent medal rate. For DeepSeek V4 Flash, DeepSeek V4 Pro, GLM 5.2 and Kimi K3 the orange Oneshot points sit clearly above the blue Chat points.
Chat against Oneshot on the default29 split. The gap opens only for models with heavy agentic post-training; the DeepSeek V4 Preview releases show almost none, their final releases a clear one (paper, Figure 2).

For GLM 5.2, a single Oneshot iteration reaches a 55.15 mean percentile and a 36.8% medal rate (Table 13). The Chat point for the same model sits at about 21% medal rate and about 40 percentile, which I read off the chart because it is not tabulated. The interesting detail is the DeepSeek V4 pair. The Preview releases show nearly no Chat-to-Oneshot gap; the final releases, which the authors say got much more coding-agent post-training on comparable base models, open a clear one. It is the paper's best evidence for its broader claim that post-training ate the harness, and they flag it as suggestive rather than shown.

Above Oneshot, the harness chains sessions at a matched 24 hours under four search rules: Chain (always refine the newest), Greedy (refine the best), UCB1 (an upper-confidence rule over a rank-normalised score with c=0.75c = 0.75) and Best-of-N (independent sessions, no sharing). Then comes Malena, which removes the search loop entirely. Then come additions to Malena, then more Malenas.

the harness ladder, rung by rungevery number from the paper's tables
fixed14 · 24h · GLM 5.2 · Tables 1 and 10
20%30%40%50%60%70%baseone Malena session55.7%+Dbackground subagents45.2%+P3three sessions, one GPU52.4%+B+P3three sessions + broadcast33.3%
● self-select, 95% CI□ oracle (best hidden-test submission)shaded: selection gap
base vs this rowΔ medal ratewon A/B/=sign test p
+D+10.5 [-1.0, +22.0]4/4/61.000
+P3+3.3 [-8.2, +14.9]4/3/71.000
+B+P3+22.4 [+10.8, +33.7] *7/2/50.180

Δ and its bracket are the paper's paired-by-task bootstrap; * means the bracket excludes zero. The win counts are printed beside each comparison in the paper; the sign test on them (ties dropped) is computed here. It asks whether the result would survive drawing a different set of tasks, which the bootstrap does not.

The widget holds the four tables that matter. Read each tab on its own: the splits differ, so a number in one tab is not comparable with a number in another.

On the search tab, no rule beats any other. The biggest pairwise gap is UCB1 ahead of Best-of-N by 2.71 percentile points, with a 95% interval of −1.44 to +6.60. Malena then beats all four on medal rate in the paired test, by 6.6 to 11.9 points. Its percentile edge is significant against Best-of-N and Chain only.

This is my first disagreement with the summary. "Giving the model a shell was the only change that clearly mattered" is close to the paper's own framing (the coding-agent environment is "by far the largest effect we measure"), but it is not what the tables say. Letting the session run for the whole budget instead of chaining six-hour sessions also mattered, significantly, on medal rate. And removing things mattered: Malena without its system and jobs tools loses 8.9 medal points, significant, and adding time nudges at 25/50/75/90% of the run costs 10.6, also significant. What the paper actually shows is narrower and still useful: once the agent has a shell and the whole budget, nothing they added helped. That includes subagent delegation, a Hugging Face search MCP, strict validation guidelines, eight hand-written ML-engineering skills, and a nudge when the GPU sits idle. The idle-GPU nudge was the worst of them: 30.4% against 40.3% in the 14-task, 12-hour runs.

The selection gap column is the other thing to watch on that tab. Best-of-N reaches an oracle percentile of 71.04, as high as Malena's 71.18, but its self-selected number is 64.21. Its best submission was in the pile. It could not tell which one it was. Malena's gap is the smallest of the five at 1.876 points.

Four harnesses, run as they ship

Section 5 is the uncontrolled half, on purpose: it takes released MLE systems and runs them as released, under the same backbone, hardware and budget. The four are not AIDE, ML-Master or R&D-Agent. They are:

So "four multi-agent ML systems" is loose. One is tree search over completions, one is a single conversation, two are genuinely multi-worker. And "re-ran every wrapper on one codebase" mixes up the two halves of the paper: the one-codebase part is the ladder above; these four ran on their own code.

Fidelity is where a comparison like this usually dies, and the authors spend an appendix on it. They fixed two bugs that broke the baselines on GLM 5.2: MLEvolve forced a named tool_choice, which threw the server's grammar-constrained decoder off when the model wanted to reason first, and ScienceFlow sent an empty tools: [] array, which the deployment rejected. They ran MLEvolve on DeepSeek V4 Flash at medium reasoning effort instead of max, because at max each call reasoned about five times longer (37,124 against 7,069 characters), the tree only grew 18.5 nodes a run instead of 62.8, and the medal rate on a 10-task check went from 0% to 25.0%. They tried one and two ScienceFlow workers, and longer Arbor timeouts and no early stopping. None of those moved beyond noise. It is a reasonable amount of effort spent on someone else's system, and it is more than most papers report. It is still less than the effort that went into Malena, and the authors list that asymmetry as a limitation themselves.

Two dot-and-whisker panels, medal rate above and percentile below, grouped by backbone: Kimi K3, GLM 5.2, DeepSeek V4 Flash and Gemma 4 31B. Colours are Malena blue, Arbor orange, AiScientist green, MLEvolve purple, ScienceFlow red; filled circles are self-select and open squares oracle. At GLM 5.2 the blue Malena points sit near 62 to 65 percent medal rate, well above every other harness. At DeepSeek V4 Flash, Arbor is slightly above Malena. At Gemma 4 31B every harness is low and MLEvolve is highest on percentile.
Malena against the released harnesses on fixed30 at 24 hours, four backbones. Brackets are 95% bootstrap intervals that resample seeds with the tasks held fixed (paper, Figure 1).

The headline numbers are in Table 15, and the summary's pair checks out exactly. With GLM 5.2 on the 30-task split, Malena's self-selected medal rate is 62.5% [58.3, 66.7]. The best external harness by the same measure is AiScientist at 47.1% [42.3, 52.2], with Arbor at 45.8%, MLEvolve at 39.2% and ScienceFlow at 24.4%.

The other backbones are less tidy than "matched or beat" suggests.

That last one is the pattern the authors build their bitter-lesson reading on: hand-built search helps the weak model and stops helping as the model gets better. It is a fair reading of four backbones. It is also four points on a curve, and the ordering of "capability" is the paper's.

NatureBench, a newer set of 40 tasks built from Nature-family papers, is the contamination hedge. There the score is how often the run's best result beats the published state of the art, at 8 hours. With GLM 5.2, Malena gets 21.7% against AiScientist's 17.5% and MLEvolve's 10.8%. With Kimi K3, AiScientist is nominally ahead, 27.1% to 26.7%. Arbor and ScienceFlow were not run on it.

Where 62.5% and 55.7% come from

The summary puts the two numbers in adjacent sentences, which invites reading the second as a drop from the first. They are different experiments.

NumberTableSplitBudgetWhat it is
62.5%15fixed30, 30 tasks24hMalena, GLM 5.2, self-select
47.1%15fixed3024hAiScientist, GLM 5.2, self-select
55.7%1fixed14, 14 tasks24hMalena base, GLM 5.2, self-select
33.3%1fixed1424h+B+P3, GLM 5.2, self-select

The task sets need a word. MLE-bench has 75 Kaggle competitions. The paper uses 30 of them, following AIRA₂'s 30-task set from the GPT-5 system card with one swap for licensing (5 Lite, 21 Medium, 4 High). The ladder runs on default29, the original MLE-bench tasks minus the swapped pair. The comparison with released harnesses runs on fixed30, where four tasks have their preparation scripts repaired (more on those below). And fixed14 is a "challenging" subset chosen by dropping every task Gemma 4 31B solved and every task Kimi K3 did not.

There is one more 62.5% in the paper, and it is a coincidence worth knowing about if you search the PDF: Malena's oracle medal rate on default29 in Table 9 is also 62.5%. Same agent, same backbone, different split, different selection rule.

Why three agents with a message channel did worse

The multi-agent experiment has three arms. +D turns on OpenCode's task tool for background subagents and adds a prompt nudge to use it: the planner and executor pattern. +P3 launches three independent Malena sessions on the same machine, sharing the one A100. +B+P3 adds a broadcast tool on top: any agent can post a message, and it lands in each peer's queue at the end of their current turn.

Switch the widget above to the multi-agent tab, set it to medal rate, and look at the squares rather than the dots.

The three sessions working independently (+P3) are no worse than one: 52.4% against 55.7%, and their oracle, the best submission any of them produced, is 60.7%, level with base. With broadcast, the self-selected rate falls to 33.3%. The oracle only falls to 53.0%. Most of the collapse is in the gap between them. The runs still produced medal-worthy submissions about as often as a lone session. What went wrong is the choice of which one to hand in.

The paper does not dig into the mechanism, and says so ("how to properly coordinate strong agents towards a common goal remains an open question"). But its own appendix on selection has the most plausible explanation. A single session scores all its candidates on one validation split it designed, so its scores are comparable with each other. Independent sessions each build their own split, and the harness then ranks numbers that were never on a common scale. The authors offer that as a hypothesis for Best-of-N's large gap and say they have not tested it. Three sessions writing into one registry inherit the same problem. I would guess broadcast makes it worse by letting a peer's optimistic score pull the others toward the same over-fitted idea, but that is my guess, and nothing in the paper checks it. Shared hardware is the other suspect: three training jobs on one GPU, which the authors note as a contention risk.

Two more things temper the headline drop. The paired test calls it significant, +22.4 medal points with an interval of 10.8 to 33.7. But fixed14 has 14 tasks, and the per-task tally the paper prints is 7 wins for base, 2 for +B+P3 and 5 ties. A sign test on 7 against 2 gives p = 0.18. And the same intervention at 12 hours, with time nudges switched on in both arms, gives +5.2 [−3.7, +14.2], not significant. I believe broadcast hurt. I would not quote 55.7 to 33.3 as the size of the effect.

What the error bars measure

Every interval in the paper comes from the same recipe (Appendix A.5): keep the set of tasks fixed, resample seeds within each task, recompute the macro average, take quantiles. At 24 hours they could afford 3 to 4 seeds per task.

That interval answers "how much would this number move if I re-ran the same 30 tasks with new seeds?" It does not answer "how much would it move on a different 30 tasks of the same kind?", which is what most readers mean when they see "62.5% of Kaggle tasks". Both are legitimate. The paper reports the first and calls the second out only indirectly, in its statistical-power limitation.

Medal outcomes are lumpy. Most tasks are nearly always medalled or nearly never, so seed-to-seed noise is small and task-to-task variation is large. You can estimate how lumpy from the paper's own bracket. With 30 tasks and 4 seeds (the granularity of 62.5% and 65.0% fits 120 runs exactly), [58.3, 66.7] implies that about 76% of a single run's variance is between tasks. Plug that back in with the tasks resampled as well and the interval for 62.5% becomes roughly 47% to 78%.

how wide is a medal rate?normal approximation · numbers computed here
0%20%40%60%80%100%as printedthe paper's own bracket58.3–66.7tasks fixedwhat the paper's bootstrap varies58.3–66.7tasks resampleda fresh draw of tasks46.8–78.2
one task flipping moves the rate by 3.3 pts
half-width, tasks fixed ±4.2; resampled ±15.7
with infinite seeds, resampled still ±15.1
ρ is backed out of each preset's printed bracket, so the “as printed” and “tasks fixed” bars match by construction. The orange bar is the one the paper does not report: how far the number could move on another set of tasks of the same kind. More seeds shrink the blue bar toward nothing; only more tasks shrink the orange one.

Try the presets. AiScientist's 47.1% widens to roughly 32% to 63% the same way, so the two marginal intervals overlap heavily once tasks move. The fixed14 numbers widen far more: one task changing its mind moves a 14-task rate by 7.1 points. And on all 75 MLE-bench tasks the same agent would still carry about ±10 points, however many seeds you ran, because only more tasks shrink that part.

None of this sinks the main comparison. Marginal intervals are the wrong tool for comparing two harnesses on the same tasks; pairing by task cancels the difficulty. The paper does pair, and the per-task win counts it prints let you run a test that does let the tasks vary. At GLM 5.2, Malena beats Arbor and AiScientist 12 tasks to 1 each (p = 0.003 on a sign test), MLEvolve 14 to 2 (p = 0.004) and ScienceFlow 19 to 0. Those hold. At Kimi K3, Malena against Arbor is 5 to 5 and against AiScientist and MLEvolve 7 to 2 (p = 0.18 each). The GLM 5.2 result survives a change of tasks. The Kimi K3 medal result is a tie.

My model behind the widget is a simplification. It treats each task as having one medal probability, uses a normal approximation, and assumes 4 seeds when some rows have 3. It gets the order of magnitude right. Treat the exact edges as approximate.

What it does with a day

If Malena is not running a search tree, what is it doing for 24 hours? The trace analysis is the most interesting reading in the paper, and the least tested.

Line chart of oracle best-so-far percentile against wall-clock time from 0 to 24 hours for four harnesses with shaded bands. All rise steeply in the first four hours. Malena, in blue, plateaus near 60 by hour 5 and creeps up to about 72 by hour 24. MLEvolve in purple ends near 65, Arbor in orange and AiScientist in green near 62 to 63, with AiScientist the slowest early.
Best submission so far against wall-clock time, GLM 5.2, fixed30. Every harness does most of its work in the first five to ten hours; later gains arrive as occasional steps (paper, Figure 3).

The authors snapshot the code repository at checkpoints (before each long job or submission), have GLM 5.2 label the ML techniques in each snapshot against a 13-way taxonomy, cluster the labels per task with Qwen3-Embedding-8B, and track the mix over the run. Malena builds one general pipeline early and then spends the rest of the day patching it. On the whale task it settles on a two-stage metric-learning embedder plus nearest-neighbour retrieval within its first 17 checkpoints and never changes that architecture over the remaining 63: backbone swaps, resolution bumps, threshold retuning, ensembling. The authors call it hyperparameter search over code patches, which is a good phrase.

Three stacked area charts, AiScientist, Malena and MLEvolve, showing the share of technique categories from the start of a run to the end. In Malena the pink ensembling and brown model-selection bands grow steadily and take about 60 percent of the share by the end. AiScientist shows a milder version of the same drift. MLEvolve's mix stays roughly flat after the first tenth of the run, dominated by optimisation and architecture.
Technique mix over a run, pooled over 10 tasks. Malena drifts toward ensembling, model selection and post-processing; MLEvolve's mix settles early and stays put (paper, Figure 5).

The drift toward ensembling late in the run is what a Kaggle competitor does, and nobody told Malena to do it. On the aptos2019 task, its GLM 5.2 and Kimi K3 runs both landed on soft-label pseudo-labelling, a technique no AiScientist, MLEvolve or Oneshot run on that task used. I like this section. I also notice that the labeller and the clusterer's reviewer are GLM 5.2, one of the two models being studied, and that "rarity" is relative to the other runs in the study, not to Kaggle. It is a good description of behaviour, not a measurement of quality.

The contamination check is more convincing than most. Malena's code was compared with 1,350 real Kaggle solutions across the 30 tasks, by token containment and by chunk embeddings. Zero of 924,750 pairs crossed both thresholds, and Malena's upper tail sits below the real-against-real baseline. Combined with the Chat result (a model that memorised the answer would not need the shell), I don't think memorisation explains Malena.

What a single session costs

A single long session is not free. Malena never resets, so every turn resends the whole context, and its cache-read tokens run about an order of magnitude above the others.

Box plots on a log scale of tokens per run for input, output and cache-read tokens, plus inference cost in dollars on a right-hand axis, for Malena, Arbor, AiScientist and MLEvolve. Input and output tokens are similar across the agentic harnesses, with MLEvolve highest. Malena's cache-read tokens are around ten to the eighth, roughly ten times the others. Malena's median cost is a few dollars with a long tail to about 100 dollars.
Tokens and modelled cost per 24-hour run at GLM 5.2 prices. Malena's cost is driven by cache reads from one ever-growing session; MLEvolve has none because each tree node is an independent call (paper, Figure 3, right panel).

At GLM 5.2 prices, Malena's modelled cost is $12.12 a run against AiScientist's $1.95, about 6.2 times more. That figure prices cache reads at a flat tenth of the input price, because none of the harnesses report real cache hits, so treat it as a model. The authors' point stands anyway: 24 hours of an A100 costs between $24 and $85 depending on the provider, so the GPU dominates. If you run this on a pay-per-token API instead of your own vLLM server, the ratio is the number to watch.

Context length is the other worry with one long session. GLM 5.2 runs at its native 1M tokens here. When they cap it at 64k and let OpenCode compact the session whenever it fills (9.5 compactions per run on average, each discarding about 76% of the context), the self-selected medal rate on a 10-task subset falls from 57.5% to 50.0%, with overlapping intervals. The 256k window on Gemma 4 31B is one of their guesses for why the small model prefers Best-of-N.

The four MLE-bench tasks they had to fix

This is a side note, and it is the part I would act on first if I ran MLE-bench. Four of the 30 tasks have preparation-script bugs that the maintainers have documented and chosen not to fix, so as not to invalidate the leaderboard:

The paper prints the diffs. Three of the four are answer keys sitting in the agent's working directory. A strong coding agent that reads every file in its folder, which is the behaviour this paper rewards, will find them. Any MLE-bench number from an agent with a shell should say which version of those tasks it ran. AIRA₂'s own audit found the same class of problem on a different benchmark.

What I'd take from it

I'd start an ML-engineering agent the way Malena is built: one long coding-agent session, a submission registry that records the agent's own validation score, a tool that tells it the time left, a watchdog on memory, and nothing else. On GLM 5.2 that beat four released systems on the same tasks, and the margin survives a test that lets the tasks vary. The harness-effect experiment swapped only the orchestration layer at a fixed model and got cheaper, faster runs at the same quality, and Context Language Models found that letting the model manage its own context beat hand-designed compaction. Three papers is not a law, but it is a direction.

I'd be careful with two readings. The paper does not show "no harness". It shows a small, well-chosen one, and removing two of its three tools cost 8.9 medal points. And it does not show that coordination is useless. It shows that letting three agents broadcast into one registry, without a shared validation protocol, breaks selection. The paper's own data say where the next effort should go: Best-of-N and the broadcast run both made good submissions and then could not pick them. Picking is a validation-design problem, and the one intervention that shrank the gap (a shared baseline workspace with a fixed split) also cut the ceiling. Nobody has solved it.

The result applies to strong models. With Gemma 4 31B the old tree search still wins on percentile, and the scaffolding-gets-eaten argument applies with the same caveat: it is eaten above a capability threshold, and you have to know which side of it you are on.

How I checked

I read the arXiv HTML of 2609.40303v1 end to end, including Appendices A to G, and took every number above from its tables (1, 2, 5, 6, 9, 10, 12, 13, 15, 16, 17, 21, 22) or its text. The Chat medal rate for GLM 5.2 is the one figure I read off a chart, and I say so where it appears. The figures are the paper's own SVGs, rendered to PNG and flattened onto white. There is no released code to read, and neither the X post's thread nor its replies pointed to any. The sign-test p-values are mine, computed exactly from the per-task win counts the paper prints, with ties dropped. The task-resampled intervals are mine too, from a simple model whose between-task share I backed out of the paper's own printed brackets; the widget's comments state the formula. I did not re-run anything.

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "Malena: how little harness a strong ML-engineering agent needs", ai.thesatyajit.com, October 2026.

bibtex
@misc{ghana2026minimalmlengineeringagent,
  author = {Satyajit Ghana},
  title  = {Malena: how little harness a strong ML-engineering agent needs},
  url    = {https://ai.thesatyajit.com/articles/minimal-ml-engineering-agent},
  year   = {2026}
}
share