Meta, Duke and UC Davis Researchers Unveil Self-Improving Branch Method for Agent Harness Optimization
Key Takeaways
- •The method delivered a 34.8% relative improvement on Olympiad-level math reasoning with Gemini 3 Flash, raising accuracy from 46.0% to 62.0% without retraining the underlying model.
- •The approach divides harness optimization into multiple self-improving branches, each using a distinct subset of development data and its own policy for proposing changes, with a router selecting the strongest branch for each input at deployment.
- •Performance gains extended beyond math to agentic tasks, including an 11.6% improvement on Terminal-Bench 2.0 and a 3.8% increase on SWE-bench Lite.
- •The paper, posted as arXiv:2609.37834v1 on September 29, 2026, builds directly on Meta-Harness, an earlier system released on March 30, 2026 that had already outperformed traditional optimization methods.
- •According to the authors, the gains were achieved using only development-set data with no access to the test set, but the preprint has not yet cleared formal peer review.

Researchers from Meta, Duke University and the University of California, Davis have proposed a new approach to making AI agents better: leave the underlying model alone and improve the scaffolding around it, using several self-improving teams instead of a single one.
In a preprint titled “Mixture of Self-Improving Branches for Agent Harness Optimization,” the researchers report a 34.8% relative improvement on Olympiad-level math reasoning, lifting accuracy from 46.0% to 62.0% with the Gemini 3 Flash model — all without retraining the model itself.
What a Harness Is, and Why It Matters
More formally, an agent harness is the code framework wrapped around a large language model. It governs the prompts the model receives, the tools it can call, the context it gets to see, and how its actions are executed.
Because that scaffolding is code rather than model weights, it can be revised without a new training run — the lever this line of work pulls. Harness optimization is the practice of automatically searching for a better version of that wrapper. The new paper, published on September 29, 2026 as arXiv:2609.37834v1, builds directly on an earlier system called Meta-Harness.
How the Branching Approach Works
Meta-Harness, released on March 30, 2026, had previously outperformed traditional methods across various benchmarks. The new work argues that a single search path leaves performance on the table.
The researchers therefore split the search into multiple specialized branches. Each branch evolves on its own, working with a distinct subset of development data and applying its own policy for proposing changes.
The branches also learn from their history. Each one retains the cases where it beats its siblings and refines its strategy based on how earlier attempts performed.
That an obvious question: who picks the specialist? The answer is a router. At deployment, the router selects the best branch head for each incoming input.
The entire process runs only on development-set data. According to the authors, the complementary harnesses delivered their gains without any access to the test set.
The Benchmark Numbers
The headline result came on Olympiad-level mathematical reasoning. Accuracy climbed from 46.0% to 62.0% using Gemini 3 Flash, which the authors frame as a 34.8% relative improvement.
The gains carried over to agentic tasks. On Terminal-Bench 2.0, which tests agents working in a command-line environment, the system posted an 11.6% gain. SWE-bench Lite, a software engineering benchmark, showed a 3.8% increase.
Together, the three evaluations span math reasoning, command-line agent work, and software engineering, so the reported gains are not confined to a single task type.
Who Is Behind the Work
The author list includes Haoyu Dong, affiliated with Meta and Duke University, and Zihao Lin, affiliated with Meta and UC Davis. Lizhu Zhang and Zhuokai Zhao are listed as co-last authors.
The paper is a preprint posted to arXiv, meaning it has been shared publicly but has not necessarily cleared formal peer review.
What This Means
A jump from 46.0% to 62.0% on hard math, achieved without touching model weights, suggests that meaningful improvements can come from engineering the environment a model operates in. For teams building on top of large language models, it frames the harness itself as an optimizable artifact — one that can move alongside, rather than wait on, model releases.
There are open questions worth watching. Running and maintaining multiple evolving branches plus a router likely adds complexity, and the paper’s results come from specific benchmarks and a specific model. Whether the approach clears peer review, holds up under independent testing, and extends beyond Gemini 3 Flash are the natural checkpoints to track.
Meta-Harness arrived in March 2026 and was already beating traditional methods. Roughly six months later, its successor claims to beat Meta-Harness itself.