Read as article
Google Open-Sources RRSI: Agents That Improve Their Own Harness
By @sharedot · · 7 pages
Google Research released RRSI under Apache 2.0, letting a frozen LLM agent rewrite its own prompts, tools and memory while regularizers keep gains transferring.
What happened: a self-editing harness, open sourced
Google Cloud AI Research, working with UNC-Chapel Hill, Stanford and Washington University in St. Louis, released RRSI (Regularized Recursive Self-Improvement). MarkTechPost reports the framework lets an LLM agent rewrite its own harness — prompts, tools, memory, control flow and sub-agents — while the model weights stay frozen. Instead of regularizing the model, RRSI regularizes the improvement loop itself, so gains hold on benchmarks the agent never optimized against. The code ships under Apache 2.0, needs Python 3.10+, accepts any LiteLLM model string, and defaults to Claude Opus 4.8 on Vertex AI.
Why it is surprising: overfitting, attacked directly
Harness-evolution loops normally propose edits, score them on a fixed evolve set and keep the winner — which invites memorization. According to MarkTechPost, the RRSI research names three failure modes: benchmark-specific fitting, noise chasing and complexity accumulation. RRSI counters each with rules mapped to classic regularizers: an annealed edit budget (cosine schedule) that narrows late rounds to one attributable change, an evidence ledger so falsified ideas are not retried, a leakage critic that screens diffs for benchmark-specific logic, a noise-adjusted floor, a cost rule forcing extra inference tokens to be paid for by measured gain, and pruning of stale components.
The evidence: all six held-out splits improved
MarkTechPost reports Terminal-Bench 2.1 rose from 74.2% to 80.2% with Claude Opus 4.8, while SWE-bench Verified — never used for selection — improved from 82.0% to 83.8%. Out-of-distribution, JobBench gained 4.7, GDPval 3.5 and APEX-Agents 3.7 points, and all six held-out splits improved. With Gemini 3.5 Flash as the policy, Terminal-Bench 2.1 rose from 64.6 to 78.7. Against competitors sharing the same starting harness, MarkTechPost reports Meta-Harness leads on the Harvey LAB evolve split (93.0 vs RRSI's 90.5), but RRSI posts the only OOD average more than a point above the 39.7 baseline, at 43.6. The harness is also lighter: 2.42M policy tokens per trial versus 3.80M unregularized, reported by the abstract as 30% fewer and by the project page as 36%.
The stakes: self-improvement you can actually trust
The core tension in agentic self-improvement has been that automated gains rarely transfer — RRSI's regularizers are a direct attempt to make them stick. MarkTechPost notes the release is a research framework, not an official Google product, which tempers any production claims. It also lands amid a broader wave of agent infrastructure: AutoTrust AI separately released JEV-27B, an Apache-2.0 open decision model that adds a 108.9-million-parameter decision block to a frozen Qwen3.8-27B backbone, trained in about 9.2 hours on one NVIDIA B200 per PR Newswire. And PPC Land reports Google researchers authored a paper on SAFE, a multi-agent system for forensic investigation of coordinated synthetic-video channel networks on YouTube.
What comes next: clone it and wire in a domain
Getting started is a git clone of google-research/rrsi plus a pip install, with MarkTechPost noting each round drafts two candidates in separate git worktrees, screens them, evaluates both and fast-forwards the branch to the winner. The coding instance also needs Docker and harbor, and new domains plug in through a single adapter module. For builders, the interesting experiment is swapping the default Claude Opus 4.8 policy for any LiteLLM string and watching whether the regularized loop still transfers — the Gemini 3.5 Flash results suggest it does. In parallel, AutoTrust's sovereign self-hosted decision models and Google's own agentic forensics work signal that the harness, not the weights, is where capability is now being pushed.