Meta, Duke and UC Davis researchers detail branching method for AI agent harnesses
In brief
- arXiv preprint 2609.37834v1, posted September 29, 2026, hasn't been peer reviewed.
- Agent harness, not the model itself, is what the method improves.
- Gemini 3 Flash Olympiad math accuracy rose from 46.0% to 62.0%, authors report.
- Terminal-Bench 2.0 gain reported at 11.6%; SWE-bench Lite gain at 3.8%.
- Router picks the best branch for each input at deployment.
What's being optimized
An agent harness covers the prompts a model receives and the tools it can call, along with the context it gets to see and how its actions are executed. Harness optimization is the practice of searching for a better version of that wrapper automatically.
The paper, arXiv:2609.37834v1, is titled "Mixture of Self-Improving Branches for Agent Harness Optimization." It builds directly on Meta-Harness, an earlier system released March 30, 2026 that had outperformed traditional methods across various benchmarks.
The new work's argument is simple: a single search path leaves performance on the table.
How the branches work
The researchers split the search into multiple specialized branches. Each one evolves on its own, with a separate subset of development data and its own policy for proposing changes. A branch keeps the cases where it beats its siblings and refines its strategy based on how earlier attempts performed.
So which branch handles a given task? That's the router's job. At deployment, it selects the best branch head for each incoming input.
It's all done on development-set data, and the authors said the gains came without access to the test set.
The numbers (and the caveats)
According to the preprint, as reported by Crypto Briefing, the system beat Meta-Harness on math, terminal and coding benchmarks. The authors report a 34.8% relative improvement on Olympiad-level math reasoning. On Terminal-Bench 2.0, which tests agents working in a command-line environment, they report an 11.6% gain; SWE-bench Lite, a software engineering benchmark, showed a 3.8% increase.
None of this has cleared formal peer review.
It's a preprint, and LeoDex News hasn't independently checked the results. Crypto Briefing noted that running multiple evolving branches plus a router likely adds complexity, and that the results come from specific benchmarks and a specific model.
Haoyu Dong (affiliated with Meta and Duke) and Zihao Lin (affiliated with Meta and UC Davis) are on the author list, with Lizhu Zhang and Zhuokai Zhao listed as co-last authors. The work comes from individual researchers, and it isn't a Meta product launch.
Frequently asked questions
What is an agent harness in AI?
An agent harness is the code framework wrapped around a large language model. It covers the prompts the model receives, the tools it can call, the context it sees and how its actions are executed. Harness optimization automatically searches for a better version of that wrapper.
How does the Mixture of Self-Improving Branches method work?
The researchers split the harness search into multiple specialized branches, each evolving with its own development data and its own policy for proposing changes. Branches keep the cases where they beat their siblings, and at deployment a router selects the best branch head for each input.
Have the results been peer reviewed?
No. The paper is a preprint posted to arXiv as arXiv:2609.37834v1 on September 29, 2026, and it hasn't necessarily undergone formal peer review. The authors reported the gains were achieved using development-set data only, without access to the test set.


