Googleās RRSI framework lets AI agents rebuild their own harnesses
Google Cloud AI Research and university partners released an open-source tool that evolves the scaffolding around a fixed language model while trying not to game the benchmarks
Google Cloud AI Research has released RRSI, an open-source framework that lets AI agents keep rewriting the machinery around their own language model. The model itself stays fixed. Everything bolted onto it is fair game.
What RRSI actually does
RRSI stands for Regularized Recursive Self-Improvement of Agent Harnesses. It was introduced on September 21, 2026, in collaboration with UNC-Chapel Hill, Stanford, and Washington University in St. Louis. The code is available under the Apache 2.0 license, which is permissive enough for commercial use.
A quick primer on the word “harness.” If the language model is an engine, the harness is the rest of the car: steering, transmission, dashboard, glovebox. Technically, it covers the prompts, control flow, tools, and memory management that turn a raw model into an agent capable of completing tasks.
RRSI lets the agent propose edits to that harness, test them, and keep the ones that work. Then it repeats the loop. The “recursive” part means the improved harness becomes the starting point for the next round of improvements.
The “regularized” part of the name addresses the overfitting problem. The framework opens the full edit space, then constrains how edits are proposed and which ones survive.
On the proposal side, RRSI uses temporally annealed edit budgets. Put simply, the agent gets more room to make sweeping changes early on and is gradually pushed toward smaller tweaks as the process matures. It also uses history-conditioned exploration, so new proposals take into account what has already been tried.
On the selection side, there are four filters:
AI, tech, and the markets they move—in one daily briefing.
Daily. Free. Join 34,000+ readers across crypto, finance, and policy.
- Leakage critic: screens for edits that sneak benchmark-specific knowledge into the harness.
- Noise floor: ignores gains too small to distinguish from random variation.
- Cost rule: weighs whether an improvement is worth what it costs to run.
- Pruning: trims away components that are not pulling their weight.
The numbers
The headline efficiency figure: RRSI uses approximately 30% fewer policy tokens than unregularized evolution methods.
On Terminal-Bench 2.1, a benchmark for command-line coding tasks, scores climbed from 74.2% to 80.2%, a gain of 6.0 points.
On SWE-bench Verified, results moved from 82.0% to 83.8%, a 1.8-point improvement.
Across eight benchmarks overall, the researchers reported gains of up to +14.1 points on evolve splits. On five out-of-distribution benchmarks, tasks the agent did not train against, the reported improvement was +4.7 points.
The test domains were deliberately varied. They included coding through Terminal-Bench, workspace tasks through Harvey LAB, and engineering design through EngDesign.
Transfer across models
Harnesses evolved using Gemini 3.5 Flash were able to improve performance when paired with other model variants, including smaller Gemini models.
The full research paper is posted on arXiv under number 2609.24972. Google has also published a GitHub repository and a dedicated project website.