MIT and Sakana AI’s SIFT framework cuts the cost of judging self-improving coding agents

Photo: Steve A Johnson / Pexels

MIT and Sakana AI’s SIFT framework cuts the cost of judging self-improving coding agents

A new method uses a language model to referee code changes head-to-head, hitting 35.1% on Polyglot at a fraction of the compute bill

Coding agents that rewrite their own code have a quiet problem. Every time they tweak themselves, someone has to check whether the tweak helped.

Researchers from MIT and Sakana AI think they have found a cheaper way to do that checking. Their framework, called SIFT, short for Self-Improvement via Fast Tree-search, posted a full score of 35.1% on the Polyglot coding benchmark after just 30 expansion steps.

The comparison point matters. The earlier Darwin Gƶdel Machine (DGM) approach reached 30.7% on the same benchmark, but only after 80 nodes.

How SIFT picks winners without running the full gauntlet

Recursive self-improving coding agents work like a writer revising their own drafts. The agent proposes a change to its own source code, hoping the new version performs better on real tasks.

The catch is verification. Testing every proposed patch against a full benchmark burns serious compute, and the bill compounds quickly when an agent generates many candidates.

SIFT sidesteps much of that cost with a referee. Instead of benchmarking every change, it asks a large language model to compare two candidate modifications and say which one looks better.

Those head-to-head verdicts are then combined using a regularized Bradley-Terry model, a statistical method for turning pairwise preferences into an overall ranking.

Advertisement

The framework also runs its evaluations asynchronously. Candidates do not have to wait in a single-file line for their turn on the test bench.

The result is a hybrid pipeline. The language model filters the field cheaply, and only the shortlisted modifications get the costly downstream evaluation.

The numbers behind the efficiency claim

The headline result came from an o3-mini coding agent. On Polyglot, SIFT reached its 35.1% score in 30 expansion steps, edging past DGM’s 30.7% achieved over a much longer 80-node search.

The cost story gets sharper with open-weight models. A configuration built on Qwen3-Coder-30B finished its entire search in 224 CPU hours with approximately $34 in API costs, about one-tenth of the resources DGM consumed.

SIFT also showed gains on other benchmarks. On TerminalBench 2.1, performance climbed from 29.2% to 36.7%, an improvement of 7.5 percentage points.

On SWE-60, the jump was larger. Scores rose from 40.0% to 52.1%, a gain of 12.1 percentage points over the starting baseline.

Who built it and where it fits

The paper was authored by Xinghong Fu of MIT, alongside Aravinth Kulanthaivelu and Yutaro Yamada, with Sakana AI as the collaborating lab. It was published on arXiv around September 18, 2026.

Discussion of the work began spreading across various platforms in late September 2026. Much of the conversation has centered on two themes: the cost savings and the growing importance of how good the language model judge actually is.

The approach fits neatly with Sakana AI’s broader philosophy. The Tokyo-based lab has emphasized evolutionary discovery methods over brute-force strategies in AI development, and SIFT is very much a case of working smarter rather than throwing more hardware at the problem.

It also builds directly on the lineage of the Darwin Gƶdel Machine. DGM established that agents could improve themselves through iterative search, and SIFT targets the evaluation bottleneck that made that search so expensive.

What this means for AI developers and the self-improvement race

The most immediate implication is accessibility. If a full self-improvement search can run for around $34 in API costs, the field may no longer belong only to labs with deep compute budgets.

That caveat deserves attention. SIFT’s entire premise rests on the language model judge making good calls, which is why judge quality has emerged as a central talking point in discussions of the paper.

A judge that consistently prefers the wrong patch would steer the search in the wrong direction, just more cheaply.

Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our Editorial Policy.
MIT and Sakana AI’s SIFT framework cuts the cost of judging self-improving coding agents
MIT and Sakana AI’s SIFT framework cuts the cost of judging self-improving coding agents

A new method uses a language model to referee code changes head-to-head, hitting 35.1% on Polyglot at a fraction of the compute bill

Photo: Steve A Johnson / Pexels

Coding agents that rewrite their own code have a quiet problem. Every time they tweak themselves, someone has to check whether the tweak helped.

Researchers from MIT and Sakana AI think they have found a cheaper way to do that checking. Their framework, called SIFT, short for Self-Improvement via Fast Tree-search, posted a full score of 35.1% on the Polyglot coding benchmark after just 30 expansion steps.

The comparison point matters. The earlier Darwin Gƶdel Machine (DGM) approach reached 30.7% on the same benchmark, but only after 80 nodes.

How SIFT picks winners without running the full gauntlet

Recursive self-improving coding agents work like a writer revising their own drafts. The agent proposes a change to its own source code, hoping the new version performs better on real tasks.

The catch is verification. Testing every proposed patch against a full benchmark burns serious compute, and the bill compounds quickly when an agent generates many candidates.

SIFT sidesteps much of that cost with a referee. Instead of benchmarking every change, it asks a large language model to compare two candidate modifications and say which one looks better.

Those head-to-head verdicts are then combined using a regularized Bradley-Terry model, a statistical method for turning pairwise preferences into an overall ranking.

Advertisement

The framework also runs its evaluations asynchronously. Candidates do not have to wait in a single-file line for their turn on the test bench.

The result is a hybrid pipeline. The language model filters the field cheaply, and only the shortlisted modifications get the costly downstream evaluation.

The numbers behind the efficiency claim

The headline result came from an o3-mini coding agent. On Polyglot, SIFT reached its 35.1% score in 30 expansion steps, edging past DGM’s 30.7% achieved over a much longer 80-node search.

The cost story gets sharper with open-weight models. A configuration built on Qwen3-Coder-30B finished its entire search in 224 CPU hours with approximately $34 in API costs, about one-tenth of the resources DGM consumed.

SIFT also showed gains on other benchmarks. On TerminalBench 2.1, performance climbed from 29.2% to 36.7%, an improvement of 7.5 percentage points.

On SWE-60, the jump was larger. Scores rose from 40.0% to 52.1%, a gain of 12.1 percentage points over the starting baseline.

Who built it and where it fits

The paper was authored by Xinghong Fu of MIT, alongside Aravinth Kulanthaivelu and Yutaro Yamada, with Sakana AI as the collaborating lab. It was published on arXiv around September 18, 2026.

Discussion of the work began spreading across various platforms in late September 2026. Much of the conversation has centered on two themes: the cost savings and the growing importance of how good the language model judge actually is.

The approach fits neatly with Sakana AI’s broader philosophy. The Tokyo-based lab has emphasized evolutionary discovery methods over brute-force strategies in AI development, and SIFT is very much a case of working smarter rather than throwing more hardware at the problem.

It also builds directly on the lineage of the Darwin Gƶdel Machine. DGM established that agents could improve themselves through iterative search, and SIFT targets the evaluation bottleneck that made that search so expensive.

What this means for AI developers and the self-improvement race

The most immediate implication is accessibility. If a full self-improvement search can run for around $34 in API costs, the field may no longer belong only to labs with deep compute budgets.

That caveat deserves attention. SIFT’s entire premise rests on the language model judge making good calls, which is why judge quality has emerged as a central talking point in discussions of the paper.

A judge that consistently prefers the wrong patch would steer the search in the wrong direction, just more cheaply.

Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our Editorial Policy.