Skip to main content

Google Research Opens Sources RRSI for AI Agent Harnesses

Google Research releases RRSI under Apache 2.0, an open-source framework allowing AI agents to improve their own harnesses without overfitting, boosting Terminal-Bench 2.1 to 78.7.

AI-written
Inewgen
29 Sep 2026Source: MarkTechPost3 min read (0 views)
Share
Google Research Opens Sources RRSI for AI Agent Harnesses

Stock photo for illustration only, not from the actual event

Font size
  • Google Research open-sources RRSI under the Apache 2.0 license for AI agents
  • The system uses harness evolution loops to propose edits and score them
  • Prevents three core failure modes: benchmark fitting, noise chasing, and complexity
  • Boosts Terminal-Bench 2.1 from 64.6 to 78.7 while saving 30-36% in policy tokens

Google Research has officially unveiled RRSI, a new open-source research framework released under the Apache 2.0 license designed to let AI agents improve their own evaluation harnesses while avoiding severe overfitting. The codebase requires Python 3.10 or higher and accepts any LiteLLM model string, with defaults pointing directly to Claude Opus 4.8 hosted on Vertex AI.

Harness evolution loops operate by proposing specific code modifications, scoring them against a fixed evaluation set, and retaining the winning version. Because identical tasks are reused across rounds, the loop risks memorizing them. The RRSI research highlights three distinct failure modes: benchmark-specific fitting, noise chasing, and complexity accumulation. Each of these factors widens the performance gap between evolve-set scores and genuine out-of-distribution transfer.

78.7Terminal-Bench 2.1 (up from 64.6)
79.0SWE-bench Verified (up from 76.8)
30-36%Reduction in policy tokens per trial

Instead of locking down components, RRSI keeps every harness element fully editable while regularizing how the search space transitions. The researchers map these mechanics directly to classic regularizers: edit budgets correspond to L0, code pruning maps to Lasso (L1), and cost rules align with Ridge (L2).

software code programming computer screen no logo

Stock photo for illustration only, not from the actual event

Performance metrics across all six held-out splits showed notable improvements. Utilizing Gemini 3.5 Flash as the underlying policy, Terminal-Bench 2.1 scores rose from 64.6 to 78.7, while SWE-bench Verified advanced from 76.8 to 79.0. Furthermore, the overall harness footprint became significantly lighter. Running on agentic workspace instances, RRSI consumed 2.42 million policy tokens per trial compared to 3.80 million used by unregularized evolution loops—representing a 30 percent reduction in the abstract and up to 36 percent on the project page.

Never miss the latest news?

Subscribe to get news summaries by email - not often enough to be annoying.

โฆษณา

Applying traditional mathematical regularizers like L1 (Lasso) and L2 (Ridge) to AI agent harness evolution represents a sophisticated method for tackling generalization roadblocks. By mathematically regularizing structural search directions rather than restricting code flexibility, researchers enable agents to adapt dynamically without succumbing to memorization traps or excessive resource consumption.

Developers looking to deploy the framework can set it up via GitHub repository cloning, install developer dependencies, and execute coding domain benchmarks. Each iteration drafts two candidate modifications within separate git worktrees, screens them, evaluates outcomes, and fast-forwards the branch to the top-performing winner.

Source: MarkTechPost

Comments

Leave a Comment
0/2000

Found something wrong in this article? Report an issue with this article