Apple study finds minimal coding agent matches or beats multi-agent ML systems

Editorial illustration: A single robotic arm on the left and four robotic arms on the right assemble similarly sized towers of pale blue glass blocks on separate metal platforms.

In brief

  • Apple researchers say one well-prompted coding agent matched or beat four multi-agent systems.
  • Malena, the minimal agent, could only read files, write files and run bash commands.
  • MLE-bench: Malena posted a 62.5% any-medal rate versus AiScientist's 47.1%, per the paper.
  • Model strength, not harness complexity, drove performance on current benchmarks, the authors say.
  • Constrained compute kept seed counts modest, a limitation the researchers flagged.

A stripped-down agent called Malena

The paper is titled "How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?" and its research team includes Alejandro Hernández-Cano, Kirill Brilliantov and Emmanuel Abbé. A harness, in this context, is the software wrapped around a large language model that lets it act (tools, memory, planning loops and sometimes multiple cooperating agents).

Apple's minimal agent doesn't have much of one. Malena ran in a single session and could only read files, write files and run bash commands. The researchers compared it against MLEvolve, AiScientist, Arbor and ScienceFlow, testing each under matched conditions on frontier large language models so the comparison isolated the harness rather than the model underneath.

The numbers

On MLE-bench, Malena posted a 62.5% any-medal rate, according to the paper's own results. That metric tracks how often an agent's results were strong enough to earn a medal on a given task. AiScientist, the best-performing external harness, scored 47.1%.

That's a 15.4-point gap.

The evaluation covered MLE-bench and NatureBench in two settings: 30 tasks with a 24-hour budget and 40 tasks with an 8-hour budget. The team ran systematic ablations, removing or swapping components one at a time, and tested different combinations of harnesses and backbone models. Their finding: once an agent had direct access to its environment, adding multi-agent coordination or extra autonomy mechanisms produced no significant gains across the architectures tested.

Caveats the authors flag

The study concluded that the strength of the underlying model was the main driver of performance, and that additional harness complexity was often futile on current benchmarks. It's a narrow claim, though. Compute budgets were constrained, which kept the number of seeds (repeated runs) modest, and the researchers framed their conclusion around current benchmarks; Crypto Briefing noted some margins could narrow with more runs.

The outlet also argued that when harness improvements are reported without matched backbone models, gains attributed to architecture may actually come from a stronger underlying LLM. That's Crypto Briefing's analysis, not a result from the paper. It does leave a simple question for the next multi-agent paper: which model was it running on?

Frequently asked questions

What is an AI agent harness?

A harness is the software wrapped around a large language model that lets it act. It can include tools, memory, planning loops and sometimes multiple cooperating agents. Apple's study tested how much of that scaffolding a strong agent actually needs for ML engineering work.

How did Apple's minimal agent Malena perform against multi-agent systems?

According to the paper's own results, Malena posted a 62.5% any-medal rate on MLE-bench. The best-performing external harness, AiScientist, scored 47.1%, a gap of 15.4 percentage points. The other systems tested were MLEvolve, Arbor and ScienceFlow.

What are the limitations of Apple's harness study?

The researchers said constrained compute budgets kept the number of seeds, or repeated runs, modest. They framed their conclusion around current benchmarks, and Crypto Briefing noted that some margins could narrow with more runs.