Running RRSI as a service, AI Agents
RRSI Agency Implementation, Productizing AI Agent Evolution
Learn how to deliver AI agent optimization as a recurring service by setting up RRSI's regularized evolution loop, configuring leakage detection and noise-adjusted acceptance criteria, and building client-facing dashboards that show which prompt and tool changes actually transfer to production. Agencies will master harness auditing, benchmark design, and cost-justified iteration to retain clients through measurable agent performance gains.
Open the decision record for RRSIWhat does running RRSI for clients commit you to?
Published figures for this service. Blank fields are not published.
- Monthly tool cost
- Not published, RRSI is open-source with no published pricing tiers; the agency must supply its own compute, inference, and labor cost basis.
- Time to first value
- Not published
- Payback
- Not modeled
- Guided implementation
- 8 hours
Is RRSI worth running as a client service?
The evidence supports that RRSI can produce agent harnesses whose gains transfer to held-out benchmarks while using fewer policy tokens, which is a credible basis for a managed harness-evolution service. What remains unknown is any vendor cost basis (RRSI is open-source), any client price, labor rate, or delivery volume, all of which the agency must supply before an ROI timeline can be modeled.
An agency-fit judgement for reselling this service. It is separate from the tool description on the decision record.
Before you start
What has to be in place before the first client engagement.
Tools and subscriptions
- RRSI open-source code from github.com/google-research/rrsi
- A base agent harness (H0) with editable prompts, tools, control flow, memory, and context
- An evolve-set benchmark suite (e.g. Terminal-Bench 2.1) plus held-out benchmarks (SWE-bench Verified, Harvey LAB, JobBench, GDPval, APEX-Agents, Frontier-Eng)
- A frozen LLM policy model (e.g. Claude Opus 4.8 or Gemini 3.5 Flash) plus proposer, analyst, and critic models
- Compute and inference budget sufficient to run recursive evolution rounds with k=2 trials across benchmark tasks
People and inputs
- ML engineering staff who can configure proposer/analyst/critic models and benchmark suites
- A defined noise band δ and annealed edit-budget schedule for each client engagement
- Per-client isolation tooling, since RRSI provides no multi-account or agency dashboard
- Time to run multiple evolution rounds before a transferable harness exists
Included with the course
7 working documents for delivering this service.
- RRSI Harness Audit Checklist for Client Onboardingchecklist
- Benchmark Suite Design Template for Coding and Workspace Agentstemplate
- Leakage Detection SOP and False-Gain Prevention Workflowsop
- Cost-Justified Edit Budget Worksheet for Monthly Retainersworksheet
- Evidence Ledger Report Template for Client Performance Reviewstemplate
- Noise-Adjusted Floor Configuration Guide by Task Domainguide
- Harness Component Pruning Decision Matrixworksheet
Listed by name. These documents are not yet published as individual downloads.