Google Research Opens Sources RRSI for AI Agent Harnesses
Google Research releases RRSI under Apache 2.0, an open-source framework allowing AI agents to improve their own harnesses without overfitting, boosting Terminal-Bench 2.1 to 78.7.

Stock photo for illustration only, not from the actual event
- Google Research open-sources RRSI under the Apache 2.0 license for AI agents
- The system uses harness evolution loops to propose edits and score them
- Prevents three core failure modes: benchmark fitting, noise chasing, and complexity
- Boosts Terminal-Bench 2.1 from 64.6 to 78.7 while saving 30-36% in policy tokens
Google Research has officially unveiled RRSI, a new open-source research framework released under the Apache 2.0 license designed to let AI agents improve their own evaluation harnesses while avoiding severe overfitting. The codebase requires Python 3.10 or higher and accepts any LiteLLM model string, with defaults pointing directly to Claude Opus 4.8 hosted on Vertex AI.
Harness evolution loops operate by proposing specific code modifications, scoring them against a fixed evaluation set, and retaining the winning version. Because identical tasks are reused across rounds, the loop risks memorizing them. The RRSI research highlights three distinct failure modes: benchmark-specific fitting, noise chasing, and complexity accumulation. Each of these factors widens the performance gap between evolve-set scores and genuine out-of-distribution transfer.
Instead of locking down components, RRSI keeps every harness element fully editable while regularizing how the search space transitions. The researchers map these mechanics directly to classic regularizers: edit budgets correspond to L0, code pruning maps to Lasso (L1), and cost rules align with Ridge (L2).

Stock photo for illustration only, not from the actual event
Performance metrics across all six held-out splits showed notable improvements. Utilizing Gemini 3.5 Flash as the underlying policy, Terminal-Bench 2.1 scores rose from 64.6 to 78.7, while SWE-bench Verified advanced from 76.8 to 79.0. Furthermore, the overall harness footprint became significantly lighter. Running on agentic workspace instances, RRSI consumed 2.42 million policy tokens per trial compared to 3.80 million used by unregularized evolution loops—representing a 30 percent reduction in the abstract and up to 36 percent on the project page.
Applying traditional mathematical regularizers like L1 (Lasso) and L2 (Ridge) to AI agent harness evolution represents a sophisticated method for tackling generalization roadblocks. By mathematically regularizing structural search directions rather than restricting code flexibility, researchers enable agents to adapt dynamically without succumbing to memorization traps or excessive resource consumption.
Developers looking to deploy the framework can set it up via GitHub repository cloning, install developer dependencies, and execute coding domain benchmarks. Each iteration drafts two candidate modifications within separate git worktrees, screens them, evaluates outcomes, and fast-forwards the branch to the top-performing winner.
Source: MarkTechPost
Found something wrong in this article? Report an issue with this article
Comments
Leave a Comment