NewsMacroMIT and Sakana AI's SIFT Framework Cuts the Cost of Judging Self-Improving Coding Agents

MIT and Sakana AI's SIFT Framework Cuts the Cost of Judging Self-Improving Coding Agents

Author: CryptoBriefing·

Key Takeaways

  • •SIFT achieved 35.1% on the Polyglot coding benchmark in just 30 expansion steps, surpassing the Darwin Gödel Machine's 30.7%, which needed an 80-node search.
  • •The framework replaces full benchmarking of every proposed change with a large language model judge whose pairwise preferences are converted into rankings via a regularized Bradley-Terry model.
  • •A configuration built on Alibaba's open-weight Qwen3-Coder-30B completed its entire search in 224 CPU hours at approximately $34 in API costs, roughly one-tenth of the resources DGM consumed.
  • •SIFT also improved performance on TerminalBench 2.1 from 29.2% to 36.7% and on S-60 from 40.0% to 52.1%.
  • •The method's effectiveness hinges on the accuracy of the language model judge, which researchers highlight as the central open question as search scale grows.
MIT and Sakana AI's SIFT Framework Cuts the Cost of Judging Self-Improving Coding Agents

Coding agents capable of rewriting their own source code carry a hidden expense that has little to do with generating the changes in the first place: every self-modification must be verified before anyone can know whether it actually helped. Researchers from MIT and Sakana AI believe they have found a substantially cheaper way to perform that verification.

Their framework, SIFT—short for Self-Improvement via Fast Tree-search—scored 35.1% on the Polyglot coding benchmark after only 30 expansion steps. The name is literal: tree-search grows a branching structure of candidate modifications, with each expansion step adding a node to explore. The baseline for comparison matters: the earlier Darwin Gödel Machine (DGM) needed 80 nodes of search to reach 30.7% on the same benchmark.

Judging candidates without running the full benchmark

Recursive self-improving coding agents operate much like a writer revising their own drafts. The agent proposes a modification to its own code, hoping the new version will perform better on real tasks. The difficulty lies in verification: testing each proposed patch against a full benchmark consumes serious compute, and the bill compounds quickly when an agent generates many candidates.

SIFT sidesteps much of that cost by introducing a referee. Rather than benchmarking every change, the framework asks a large language model to compare two candidate modifications and determine which one looks better. These head-to-head verdicts are then aggregated through a regularized Bradley-Terry model, a statistical method for converting pairwise preferences into an overall ranking.

The framework also runs its evaluations asynchronously, so candidates do not have to wait in a single-file line for their turn on the test bench. The result is a hybrid pipeline: the language model filters the field cheaply, and only the shortlisted modifications proceed to the costly downstream evaluation.

The numbers behind the efficiency claim

An o3-mini coding agent produced the headline result. On Polyglot, SIFT reached its 35.1% score in 30 expansion steps, edging past DGM's 30.7%, which was achieved over a far longer 80-node search.

The cost story sharpens further with open-weight models. A configuration built on Qwen3-Coder-30B, an open-weight release from Alibaba's Qwen family, completed its entire search in 224 CPU hours with approximately $34 in API costs—about one-tenth of the resources DGM consumed.

SIFT also delivered gains on other benchmarks. On TerminalBench 2.1, performance climbed from 29.2% to 36.7%, an improvement of 7.5 percentage points. The jump was larger on S-60, where scores rose from 40.0% to 52.1%, a gain of 12.1 percentage points over the starting baseline.

Who built it and where it fits

The paper was authored by Xinghong Fu of MIT, alongside Aravinth Kulanthaivelu and Yutaro Yamada, with Sakana AI as the collaborating lab. It was published on arXiv, the open-access preprint server where research typically appears ahead of formal peer review, around September 18, 2026, and discussion of the work began spreading across various platforms in late September 2026. Much of that conversation has centered on two themes: the cost savings and the growing importance of how good the language model judge actually is.

The approach fits neatly with Sakana AI's broader philosophy. The Tokyo-based lab, founded in 2023 by former Google researchers including Transformer paper co-author Llion Jones, has emphasized evolutionary discovery methods over brute-force strategies in AI development, and SIFT is very much a case of working smarter rather than throwing more hardware at the problem. It also builds directly on the lineage of the Darwin Gödel Machine: DGM established that agents could improve themselves through iterative search, and SIFT targets the evaluation bottleneck that made that search so expensive.

What this means for developers and the self-improvement race

The most immediate implication is accessibility. If a full self-improvement search can run for around $34 in API costs, the field may no longer belong only to labs with compute budgets.

One caveat deserves attention. SIFT's entire premise rests on the language model judge making good calls, which is why judge quality has emerged as a central talking point in discussions of the paper. A judge that consistently prefers the wrong patch would steer the search in the wrong direction—just more cheaply. Whether that judgment stays reliable as searches scale up is the open question to watch in follow-up work.

Source: CryptoBriefing