RRSI

An agent that rewrites its own harness gets better at the tasks it practices on and barely better anywhere else. RRSI regularizes the rewriting: the harness it keeps scores 43.6 on benchmarks it never saw, where unregularized evolution keeps 40.3, and runs on a third fewer tokens.

Pick a set of rules and watch twenty rounds of self-improvement: which edits get into the harness, and where the two scores go. Then we build, piece by piece, the three proposal habits and four selection gates that decide what a self-improving harness is allowed to keep.

You need what an LLM agent is and what overfitting means. We build the rest from zero.

Evolve a harness

Ready

Each round the proposer drafts two candidate harnesses. Watch which ones get in, and what they really were.

Final scores and token costs are the paper's Table 2 (agentic workspace: evolve split = Harvey LAB, out of distribution = mean of JobBench, GDPval and APEX-Agents). The paper reports only each arm's final harness, so the round-by-round path and the edit lines are a toy that shows the kinds of edits each gate stops; the gate arithmetic uses the workspace values of Table 5. The grey dots at round 20 mark where the other three rule sets finish.

Chapter 0

Better on the test it studied

Why a harness that rewrites itself can improve at everything it is graded on, and at almost nothing else

You run an agent that does legal work. It opens a folder of contracts, spreadsheets and PDFs, and it has to hand back a memo, a redline or a table, saved under the exact file name the client asked for. Underneath it sits a frozen language model, Claude Opus 4.8. Nothing about that model will change today.

What you can change is everything around it. The harness is the program wrapped around the model: the system and task prompts, the control flow that decides when the agent plans, acts, checks its work and stops, the tools it can call together with the descriptions that tell it how, the memory and skill files it may consult, and the context management that decides what the model sees at each step. The paper's own definition is shorter: the harness is everything around the weights.

The harness decides whether the same model reads the right file before editing it, recovers when a command fails, keeps its working context lean, and actually writes its findings into the deliverable. A lot of the recent progress in agent products came from this layer rather than from new weights. If you have read Harness Engineering, this is that layer.

Until recently a person did this work. They read a failed trajectory, spotted the mistake, changed the scaffold, and ran the tasks again. The paper names the bottleneck bluntly: progress is limited by how many trajectories an engineer can read.

So let a language model do the reading. A proposer model reads the failures and writes a change to the harness; the changed harness is scored on a set of practice tasks; the change is kept if the score went up. Then the loop runs again from the improved harness. Because the improved agent's own behavior produces the feedback that improves it further, the paper calls this recursive self-improvement (RSI) at the level of the agent system. No weight changes. The system keeps rewriting the part of itself it is allowed to edit.

The practice tasks are the evolve set. The loop scores every candidate on the same evolve set, round after round, and keeps whatever scores best on it. That is the recipe for fitting a model to its training data, except that the thing being fitted is now a program. The question that matters is what happens when the evolved harness meets tasks it never saw. The paper's Figure 1(a) answers it, and here it is to play with.

Evolve gain against transfer

Each point is one way of evolving the same starting harness on Harvey LAB's 120 evolve tasks. Across: how much it improved on those tasks. Up: how much it improved on three benchmarks it never saw. Tap a method to see its scores, benchmark by benchmark.

Scores from Table 1 (the four prior methods and RRSI) and Table 2 (unregularized evolution, which the paper reports only as an average out of distribution). "Out of distribution" is the mean of JobBench, GDPval and APEX-Agents, computed from Table 1's columns; the relative gains are computed from the same numbers and reproduce the positions in the paper's Figure 1(a).

Read the picture from right to left. Meta-Harness, the strongest method on the evolve split, climbs from 89.4 to 93.0 there, and adds just 0.9 points on the benchmarks it never saw. HarnessX lands exactly on the starting harness. AHE and TTHE end below it: TTHE by 1.7 points. Unregularized evolution, the plain loop with no rules, posts one of the largest evolve gains and one of the smallest transfers.

RRSI sits alone at the top left. It has the smallest evolve-set gain of any evolved harness (89.4 to 90.5) and the only out-of-distribution average that clears the base harness by more than a point: 43.6 against 39.7. The paper calls this the trade the regularizers are designed to make.

So a bigger number on the evolve split is not evidence of a better harness. It can be evidence of a harness that learned the evolve split.

Why would a self-improving loop do this? The paper names three behaviors, and each has an everyday cause.

  1. Benchmark-specific fitting. The proposer reads failed trajectories that are full of task names, company names, file names and expected values. Writing those into the harness raises the evolve score at once and helps on no other benchmark.
  2. Noise chasing. Every evaluation of an agent is noisy: the same harness scores a little differently each time. Keep whichever candidate got the luckiest score, and a change that did nothing looks like progress.
  3. Complexity accumulation. A second review pass, another subagent, a longer context: each buys a sliver of evolve score with a lot of extra tokens, and a loop that only reads the score has no reason ever to take them out again.

In the paper's words, benchmark-specific fitting, noise chasing and complexity accumulation all widen the evolve-to-transfer gap.

RRSI, Regularized Recursive Self-Improvement of Agent Harnesses (Xia, Han, Wang, Chen and colleagues at Google Cloud AI Research, UNC-Chapel Hill, Stanford and Washington University in St. Louis, 2026), keeps the loop and keeps the harness fully editable: prompts, control flow, tools, skills, memory and subagents may all change. What it regularizes is how finite, noisy feedback is allowed to become permanent. On the proposal side it limits how many edits one candidate may bundle, remembers the evidence for every edit it ever tried, and, when progress stalls, sends effort to parts of the harness nobody has touched. On the selection side a critic rejects edits that name the benchmark before they are ever scored, a floor rejects gains the noise could explain, a cost rule makes extra tokens pay for themselves, and a pruner deletes machinery that stopped earning its place.

The idea in one line: regularize the search, not the harness. Nothing is forbidden from changing. What RRSI controls is how much can change at once, which measured gains it believes, and what those gains are allowed to cost. The hero at the top of the page is this whole story in twenty rounds.

Here is the road, in order. Each chapter explains one part of the instrument in the hero.

  1. One round, measured: what exactly is scored, on what, and why reusing the same tasks every round changes everything.
  2. The winner's curse: how a loop that changes nothing still reports progress.
  3. Regularize the search: three ideas borrowed from curve fitting, aimed at the loop instead of the harness.
  4. Fewer edits, clearer credit: the annealed edit budget, the paper's L0 idea.
  5. Remember what failed: the edit ledger, and exploring what nobody has tried.
  6. Screen before you score: the critic that stops leaks.
  7. Don't walk downhill: the noise floor.
  8. Growth must pay: the cost rule and the pruner.
  9. Run the whole loop: every regularizer as a switch you can flip.
  10. What it buys, then connections.
In Figure 1(a), Meta-Harness has the largest gain on the evolve split. What does the figure say about it?

Chapter 1

One round, measured

Follow a harness through one round of evolution: run it, score it, summarize its failures, propose, and pick

Chapter 0 said the evolved harnesses "got better" on the evolve split. Before we can argue about whether that improvement is real, we need to be exact about two things: what number the loop is climbing, and which data it looks at to compute that number. This chapter builds both, one piece at a time, and then runs one round of the loop in front of you.

Start with the agent itself. The paper writes it as a pair, A = (π, H). The first half, π, is the policy: the language model's weights, the thing that reads a context and writes the next action. In every experiment of this paper the policy is frozen: Claude Opus 4.8 in all three domains, never fine-tuned, never touched. The second half, H, is the harness: the paper's definition is simply "everything around the weights".

Concretely, "everything around the weights" is five kinds of thing:

  1. Prompts. The system prompt and the task prompt: who the agent is told it is, what it is told to do, what rules it is given.
  2. Control flow. The program that decides when the agent plans, when it acts, when it reflects on what went wrong, and when it stops.
  3. Tool interfaces. Which tools exist (a shell, a file reader, a spreadsheet reader, a web search) and the descriptions the model reads to decide how to call them.
  4. Memory and skill files. Notes the agent may consult: saved procedures, lessons from earlier tasks, how-to documents.
  5. Context management. What the policy actually sees at each step: what gets summarized, truncated, or dropped when the conversation grows long.

The paper's three base harnesses are real, public ones. For coding it is Terminus-2, the shell-driving agent from the Terminal-Bench authors. For the legal-work and engineering tasks it is a ReAct loop (think, act, observe, repeat) over an MCP tool gateway, with a dynamic toolbelt and ReSum-style context management, which summarizes a long history so it keeps fitting in the context. Every run of RRSI starts from one of these, called H0, the unevolved harness.

That last phrase is literal in the released code. A harness is a directory of source files, and every candidate harness is a git commit: it is drafted, screened and evaluated in its own worktree on a branch off evolve/<domain>, and accepting it fast-forwards that branch, "so the incumbent is always a commit". The word incumbent just means the harness currently in charge, the one every new candidate must beat.

What gets measured

Give the agent a task x: a folder of legal documents and a request for a memo, or a container with a broken program and a description of the bug. The agent works: it reads, calls tools, writes files. The whole record of that work is a trajectory, written τ. At the end it leaves a deliverable, and a verifier grades it with a number r(x, τ) between 0 and 1.

What the verifier is depends on the domain, and it matters. On Terminal-Bench 2.1 the task comes with unit tests hidden from the agent; the task counts as solved only if those tests pass after the agent stops, so r is exactly 1 or 0, and "cannot be produced by a plausible-looking answer". On Harvey LAB, the legal benchmark, each task has a rubric of 20 to 100 independent criteria; each criterion is judged in isolation by an LLM judge (Gemini-3.5-Flash) reading the deliverable and that one criterion; a full evaluation is roughly 14,000 criterion verdicts. A missing deliverable fails every criterion it was supposed to satisfy.

The quantity the paper cares about is the expected score of a harness over a set of tasks D, and its expected cost in policy tokens, the tokens the frozen model reads and writes while solving. That is the paper's Equation 1:

S(H; D)
The harness's true score on the task set: what you would get averaging over infinitely many runs.
D
A set of tasks, for example the 120 Harvey LAB tasks of the evolve split.
𝔼
An average twice over: over tasks drawn from D, and over the trajectories the agent might produce on each (the policy samples, so two runs of the same task differ).
r(x, τ)
The verifier's grade of one trajectory on one task, between 0 and 1.
C(H; D)
The harness's expected cost on the same tasks.
c(τ)
The number of policy tokens one trajectory consumed.

Nobody can compute an expectation over infinitely many runs. The loop runs each task a small number of times, k trials, and averages what it sees. That average, with a hat to mark it as an estimate, is Equation 3:

Ŝ(H)
The empirical score: what the loop actually sees and climbs.
k · |D|
The total number of trials: k trials for each of the |Devolve| tasks. Ŝ divides the summed rewards by it.
j = 1..k
The trial index. k = 2 on the coding and agentic-workspace instances, k = 4 on engineering design (Table 5).
τx(j)
The j-th trajectory the agent produced on task x.
Ĉ(H)
The empirical cost: average policy tokens per trial.

Let's compute one by hand. A toy evolve set of three pass-or-fail tasks, k = 2 trials each:

TaskTrial 1: r, tokensTrial 2: r, tokens
A1, 1.2 M0, 1.8 M
B1, 0.9 M1, 1.1 M
C0, 2.4 M1, 2.0 M

There are k · |Devolve| = 2 × 3 = 6 trials. The rewards add to 1 + 0 + 1 + 1 + 0 + 1 = 4, so Ŝ = 4 / 6 = 0.667. The tokens add to 1.2 + 1.8 + 0.9 + 1.1 + 2.4 + 2.0 = 9.4 million, so Ĉ = 9.4 / 6 = 1.57 million tokens per trial. The paper reports scores in points, so this harness "scores 66.7".

Ŝ = (1 + 0 + 1 + 1 + 0 + 1) / (2 × 3) = 4 / 6 = 0.667 (reported as 66.7 points)

Two engineering details in the released code change what this average means, and both are there to stop a harness from gaming it.

Long rubrics count more. On Harvey LAB a trial's reward is criteria passed over criteria total, and its weight is the criteria total, so Ŝ is the fraction of all criteria passed, the benchmark's own metric. Take two tasks: one with a 20-criterion rubric where the agent passes 18 (reward 0.9), one with a 100-criterion rubric where it passes 60 (reward 0.6). The plain average of the two rewards is (0.9 + 0.6) / 2 = 0.75. The criteria-weighted score is (18 + 60) / (20 + 100) = 78 / 120 = 0.65. The long task carries five times the evidence, so it gets five times the say.

A missing trial is a zero, not a gap. If a trial crashes or times out, it contributes r = 0 and still counts in the denominator. Suppose 6 trials, 4 successes and 1 crash: Ŝ = 4 / 6 = 0.667, not 4 / 5 = 0.8. The code's comment gives the reason: this way "a candidate cannot look better by destroying the trials it finds hard". The paper applies the same rule to APEX-Agents, where a rollout lost to infrastructure counts as a failure, "which prevents a harness that crashes on hard worlds from looking better than one that attempts them".

python# Equation 3 as the released rrsi/evaluate.py computes it (simplified).
def aggregate(per_task):
    num = den = 0.0; toks = []
    for task in per_task.values():               # one entry per task in D_evolve
        for r, w in zip(task.rewards, task.weights): # k trials; a missing trial has r = 0
            num += r * w                              # Harvey LAB: r = passed/total, w = total
            den += w                                  # coding, engineering: w = 1
        toks += [t for t in task.tokens if t]
    S_hat = num / den                                 # the score the loop climbs
    C_hat = sum(toks) / len(toks)                     # policy tokens per trial
    return S_hat, C_hat

# the toy table above, by hand and by code: 4/6 = 0.667 and 9.4/6 = 1.57 million

One round of the loop

Now the loop. Almost every harness-evolution method, RRSI included, instantiates the same generic round. At round t the current harness Ht is run on the evolve set Devolve, the finite set of tasks the search is allowed to learn from. Its trajectories are summarized into feedback Ft. A proposer LLM reads the feedback and the harness source and writes candidate harnesses. The candidates are evaluated on the same evolve set, and the best one becomes the next incumbent. In symbols, Equation 2:

𝓗t
This round's candidates, mt of them. The released code drafts 2 per round, labeled A and B.
P0
The unconstrained proposal process: an LLM asked to improve the harness, with nothing limiting how. RRSI will replace it with a regularized Preg.
Ft
The feedback: an analyst model's summary of where Ht failed on the evolve set.
Ht+1
Next round's incumbent.
arg max
Over 𝓗t ∪ {Ht}: pick the highest measured score among the candidates and the incumbent. If no candidate beats it, the incumbent stays.

The feedback is not the raw trajectories; there are far too many. In the released code the analyst is handed the worst trial of each of the lowest-scoring tasks plus the best trial of a few top-scoring ones, which the code calls "success habits the proposer must not break", and writes a failure summary from them. Step through a round below and watch what flows where.

Step one round

Round 0

Six toy tasks, two trials each. Press Step to move through the loop one stage at a time: run the incumbent, score it, summarize its failures, propose two candidates, evaluate each on the same tasks, pick. Keep stepping into later rounds and watch the counter on the right.

The tasks, trial outcomes, token counts and edit descriptions are a toy. The stages are the paper's generic loop (Equation 2) as RRSI's Algorithms 1 and 2 run it, without the regularizers yet. The real Harvey LAB evolve set is 120 tasks × k = 2 trials, about 14,100 criterion verdicts per evaluation.

Look at the counter. After one round the loop has evaluated three harnesses, the incumbent and two candidates, and all three were measured on the same six tasks. By round 3 those six tasks have been looked at seven times, and every candidate after the first round was written by a proposer that had read the results of the earlier looks.

That is the property the paper names. "Unlike ordinary evaluation, this reuse of Devolve is adaptive: the candidates proposed at round t depend on measurements obtained from the same tasks in earlier rounds." An ordinary evaluation is a fixed harness measured once on tasks it was not built from, so its score is an honest estimate. Here the harness under test was built by reading those tasks' failures, and chosen for scoring well on them. Its score on them is no longer an honest estimate of how it does anywhere else.

What the loop really is. The paper's summary: "harness evolution can therefore be viewed as adaptive empirical optimization over an unusually expressive search space." Empirical, because it climbs Ŝ, a noisy average over a few trials. Adaptive, because each step is chosen by looking at the same data again. Unusually expressive, because a proposer may write any program. Machine learning has a name for an expressive model fitted to a finite sample: overfitting risk. The rest of this lesson is about controlling it.

Notice what candidate B did in round 0: it only reworded the system prompt, and it still measured one trial higher than the incumbent. Was the rewording helpful, or did one of its twelve trials simply get lucky? Nothing in Equation 2 can tell the difference; it keeps whatever measured highest. Chapter 2 runs that question to its end.

Why is reusing the evolve set every round different from ordinary evaluation?

Chapter 2

The winner's curse

Pick the best of several noisy scores, and a harness that changed nothing looks better

Run the same harness on the same tasks twice and you will not get the same score. The policy samples its tokens, so two trajectories on one task differ. On Harvey LAB the grader is itself an LLM judge. And the loop only runs each task k = 2 times, so one lucky or unlucky trial moves the average. The score Ŝ from Chapter 1 is a noisy measurement of the true score S.

The paper measures that noise directly. Before evolution starts, it evaluates the unchanged base harness repeatedly and records how much its score wanders. The result is the noise band δ: a score difference smaller than δ is the kind of difference an unchanged harness produces by itself. Its values, in the paper's own units (Table 5):

Instanceδ as a fractionWhat that isIn points
Coding (Terminal-Bench 2.1)0.0173 passes out of 89 tasks × k = 2 = 178 trials1.7
Agentic workspace (Harvey LAB)0.00460 criteria out of roughly 14,100 verdicts0.4
Engineering design (EngDesign)0.0205 passes out of 61 tasks × k = 4 = 244 trials2.0

Check one row by hand: 3 / 178 = 0.0169, which the paper rounds to 0.017. In words: on the coding instance, three extra passed trials out of 178 is within what luck alone can produce. That is sobering next to the paper's own table of edits, where a whole round's candidate gained +1.69 points on coding, just under this band (Chapter 8 comes back to that one).

A thought experiment: evolve nothing

Suppose the proposer were useless. Every candidate it writes is secretly an exact copy of the incumbent: same prompts, same tools, same true score. The only thing that differs between candidates is the luck of their measurement. Now run the loop of Equation 2 on these copies for twenty rounds. What happens to the score of the harness it keeps?

Commit to a guess before you press Run. The obvious answer is "nothing, they are all the same harness".

Evolve nothing

Ready

Every candidate is a copy of H0, true score 89.4. Each dot is one noisy measurement; the loop keeps the highest, exactly as Equation 2 says. The yellow line is the score the loop believes its harness has. The green dots measure the kept harness again on fresh trials. Press Run 20 rounds.

2
0.14 pts

A simulation with Gaussian noise, not the paper's data. The true score 89.4 is H0 on the Harvey LAB evolve split (Table 2). The presets turn each instance's δ from Table 5 into the noise of one evaluation using the released calibration rule, δ = 2 × √2 × that noise (Chapter 7 derives it): 0.4 / 2.83 = 0.14, 1.7 / 2.83 = 0.60, 2.0 / 2.83 = 0.71 points. The released code drafts 2 candidates per round, the default here.

The kept score climbs, round after round, although nothing was ever changed. At the paper's workspace noise level the climb is small, about a third of a point after twenty rounds. Switch to the coding preset and it is over a point. Push the candidates per round to 8 and it grows again. Meanwhile the green dots, which measure the kept harness on fresh trials, stay scattered around 89.4. The gain exists only in the loop's own bookkeeping.

This is the winner's curse: whenever you select the best of several noisy measurements, the winner's measurement is, on average, too high. Not because the winner is special, but because being selected is itself evidence of good luck. Auctions have the same curse: the bidder who wins is the one who overestimated the item's value the most.

How big is the curse? Derive it

Take the simplest case: two candidates, both with true score μ, each measured once with independent Gaussian noise of standard deviation σ. The loop keeps the larger measurement. What is it on average?

There is a neat identity for the larger of two numbers: the larger one is their average plus half their gap. Try it: for 3 and 7, the average is 5 and half the gap is 2, and 5 + 2 = 7. So:

max(X1, X2)
The score the loop keeps: the higher of two noisy measurements of the same true score.
(X1 + X2) / 2
Their average. Each measurement is unbiased, so their average is too: on average this term equals the true score μ.
|X1 − X2| / 2
Half the gap between them. A gap is never negative, so this term can only push the kept score up. The whole curse lives here.

Now take averages, one term at a time. The first term averages to μ. For the second, the difference of two independent measurements has noise from both, so its variance adds: X1 − X2 is Gaussian with mean 0 and standard deviation √2 · σ. The average size of a zero-mean Gaussian with standard deviation s is a known constant times s: s · √(2/π), about 0.80 s. So the average gap is √2 · σ · √(2/π) = 2σ / √π, and half of it is:

𝔼[max]
What the kept score is, averaged over many repetitions of the same round.
μ
The true score both candidates share.
σ
The noise of one evaluation: how much one measured Ŝ wanders around the true S.

With the workspace noise, σ = 0.14 points, one round of two identical candidates produces a phantom gain of 0.564 × 0.14 = 0.079 points. More candidates make it worse, because the maximum of more draws is larger. For m draws the average excess, in units of σ, is:

Noisy draws comparedAverage excess of the winnerAt σ = 0.14 points
100
20.56 σ+0.08
41.03 σ+0.14
81.42 σ+0.20
412.17 σ+0.30

The last column is the device's default run. Under Equation 2 the incumbent competes with the score it already has, so the kept score after twenty rounds is simply the largest of all the measurements ever taken: the starting one plus 2 per round, 1 + 2 × 20 = 41 draws. Its average excess is 2.17 σ: 2.17 × 0.14 = 0.30 points on the workspace, 2.17 × 0.60 = 1.30 points on coding. The luck compounds because a lucky score, once kept, is never checked again.

0 true gain + 2.17 × 0.60 ≈ +1.3 points (coding noise, 20 rounds of 2 copies)

Flip the device to Re-measure the incumbent. Now the incumbent gets fresh trials every round instead of keeping its lucky score. The kept line still sits a little above 89.4, because each round still keeps the best of three fresh draws, but it no longer ratchets: the luck of round 5 is forgotten by round 6. Re-measuring costs a full evaluation per round, which is why real loops rarely do it; RRSI's answer, as we will see, is not to re-measure but to refuse any "gain" smaller than δ.

The deeper version of this point is from statistics, and the paper cites it: "every evaluation is another adaptive look at the same finite evolve set" (Dwork et al., 2015, on adaptive data analysis). Each look leaks a little information about the evolve set's particular luck into the choices that follow. A held-out set stays honest only as long as nothing is chosen by looking at it, and the evolve set is looked at every single round.

The loop is an optimizer, and noise is something it can optimize. Nothing in Equation 2 distinguishes "this edit works" from "this measurement was lucky". Give an arg max enough draws and it will find the luck. The rest of the paper is about giving the loop a way to tell the difference.

Two more ways to fool yourself

Noise chasing is one of three behaviors the paper names as the sources of the gap between evolve-set gains and transfer: "These benchmark-specific fitting, noise chasing, and complexity accumulation all widen the evolve-to-transfer gap." The other two are not luck at all. They are edits that really raise the evolve score, for reasons that do not travel.

Noise chasing

this chapter

Keeping candidates whose "gain" is the luck of the measurement. It inflates the evolve score and transfers nothing. RRSI's answer is the noise floor and the within-band rule (Chapters 7 and 8).

Benchmark-specific fitting

Chapter 6

The proposer reads failed trajectories full of task names, file names and expected values, and the cheapest fix is to write them into the harness: "if the matter mentions Delaware, cite the merger statute" (an illustrative example). A real gain on those tasks; zero anywhere else. RRSI's answer is a critic that reads every diff before it is scored.

Complexity accumulation

Chapter 8

An edit that buys a little score with a lot of compute, for instance a second full review pass on every deliverable (illustrative), is kept, and nothing ever removes it. Unregularized evolution on the workspace ended at 3.80 million tokens a trial against the base harness's 1.56 (Table 2). RRSI's answer is a cost rule and a pruner.

All three share one shape: the evolve score goes up and the reason does not travel. Chapter 3 names the classical cure for that shape, regularization, and shows how the paper turns it on the loop itself.

In the "evolve nothing" run, every candidate is an exact copy of H0. Why does the kept score still climb?

Chapter 3

Regularize the search, not the harness

Borrow three cures for overfitting from curve fitting, and aim them at the loop instead of the file

Chapter 2 ended with a harness that looked better every round while nothing about it had changed. If that feels familiar, it should. It is the oldest failure in machine learning in new clothes: anything flexible enough to fit its training data will also fit the noise in that data, and the noise does not come back on new data.

Machine learning has long had a family of cures for this, all under one name: regularization, anything that limits how far a model can bend toward the particular examples it was trained on. RRSI's central move is to borrow three of those cures and point them somewhere new. To see what it borrows, and what it changes, we first need the cures in their home setting.

The home setting: a curve through twelve points

Here is the smallest overfitting problem there is. We have twelve points, each measured with noise, from some smooth curve, and we want a formula that predicts points we have not seen. We give the fit eight features to build its curve from: eight wiggly basis curves, sines and cosines of rising frequency. The fit chooses one weight per feature, which says how much of that basis curve to add in.

The truth, which the fit never sees, uses only three of the eight features. The other five are there to tempt it. With eight weights and twelve points, plain least squares can bend the curve through almost every training point, noise included. The yellow points are the ones it trains on, the way a harness is scored on its evolve split. The hollow green points are held out: never shown to the fit, and the only honest judge of it.

Three ways to tame a fit

The fit's curve over twelve training points (yellow) and forty held-out points (green rings), its eight weights, and its error on each set. Start with None, then pick a regularizer and move its strength. Show the truth draws the real curve and the real weights.

off

Toy data: the truth is 0.2 + 1.0·sin 2πx + 0.5·cos 2πx − 0.6·cos 4πx, with Gaussian noise of standard deviation 0.3, so no fit can beat an error of about 0.09 on new points. Every number is computed live in your browser: least squares and Ridge in closed form, Lasso by coordinate descent, L0 by trying every subset of the allowed size.

With no regularizer the fit is nearly perfect on the yellow points, an error of 0.003, and poor on the green ones, 0.217, more than seventy times worse. It used all eight weights, and the five the truth never used are carrying noise. That gap between the set you fit and the set you did not is exactly the gap Chapter 0 showed between a harness's evolve split and the benchmarks it never saw.

Three cures, three different things they limit

All three regularizers change what the fit is trying to minimize. Plain least squares minimizes the training error alone. A regularized fit adds a second term, a price on the weights themselves:

ŵ
The weights the fit settles on, one per feature.
1⁄n ∑i
The average over the n = 12 training points. Only these points are ever seen.
yi
The measured value of training point i, noise and all.
∑j wjφj
The fitted curve at xi: each basis curve φj scaled by its weight wj, added up.
λ
The strength: how much one unit of penalty costs compared with one unit of training error. At λ = 0 we are back to plain least squares.
R(w)
The regularizer: the price on the weights. Choosing it is choosing which cure.

The three cures are three choices of R:

Why does L1 produce exact zeros when L2 never does? One weight is enough to see it. Suppose the data alone would like a weight of a, and we charge a penalty.

With Ridge we minimize (w − a)2 + λw2. Setting the slope to zero gives 2(w − a) + 2λw = 0, so w(1 + λ) = a and w = a / (1 + λ). Ridge divides: the weight gets smaller but stays non-zero unless a was already zero.

With the Lasso we minimize (w − a)2 + λ|w|. For a positive w the slope is 2(w − a) + λ, which is zero at w = a − λ/2. That answer only makes sense if it is still positive, that is, if a > λ/2. If a is smaller than λ/2, the slope is positive for every positive w and negative for every negative w, so the lowest point is the kink itself: w = 0 exactly. The Lasso subtracts a fixed amount, and anything smaller than that amount is set to zero.

Now put numbers in. Take λ = 0.8, so the Lasso subtracts λ/2 = 0.4. A weak weight that the data wants at a = 0.3:

Lasso: 0.3 − 0.4 = −0.1 < 0, so w = 0  ·  Ridge: 0.3 / (1 + 0.8) = 0.3 / 1.8 = 0.17

A strong weight that the data wants at a = 2.0:

Lasso: 2.0 − 0.4 = 1.6  ·  Ridge: 2.0 / 1.8 = 1.11 (the Lasso keeps strong weights closer to their size)

So the three cures limit three different things. L0 limits how many pieces are active. L1 removes the weak pieces entirely and leaves a sparser model. L2 limits the total size of the solution without removing any single piece. Go back to the device and watch the weight bars: under L0 and L1 whole bars go hollow, under L2 every bar shrinks together and none disappears.

The switch: from the weights to the loop

A harness is not a vector of eight weights. It is a program: prompts, control flow, tool interfaces, memory and skill files, context management. You cannot add up the absolute sizes of a system prompt. So RRSI does not regularize the harness as an object. It regularizes the search that edits it, and it borrows each cure for the role it plays. The paper states the mapping in Section 3.1:

  1. L0, cardinality. In a fit it caps how many weights are active. In RRSI it becomes the annealed edit budget: at most bt independent edits in one candidate. Chapter 4
  2. Lasso, L1. In a fit it removes weak weights entirely, leaving a sparser model. In RRSI it becomes structural pruning: components with no recent positive gain become deletion targets. Chapter 8
  3. Ridge, L2. In a fit it shrinks the total size and removes nothing. In RRSI it becomes complexity-aware acceptance: extra policy tokens must be paid for by measured gain. Chapter 8
  4. A diversity or entropy bonus. In reinforcement learning it stops a policy collapsing onto one action. In RRSI it becomes structured exploration: a stalled run reserves a slot for components it never tried. Chapter 5
  5. A reusable holdout. In adaptive data analysis it limits how far repeated looks at one dataset can mislead. In RRSI it becomes evidence-aware credit and the noise floor, and the leakage screen serves the same end. Chapters 5, 6 and 7

The first three are the paper's own analogies, and it chooses its words carefully. The edit budget is "the closest to an L0-style cardinality constraint as it directly limits the number of independently active edits in an update." Pruning is "analogous to Lasso/L1-style sparsification because persistently unproductive components are removed." The cost rule is "analogous to Ridge/L2-style shrinkage since it suppresses unchecked growth in the aggregate resource footprint without requiring any particular component to be eliminated." The fourth is the paper's own comparison to entropy regularization in Soft Actor-Critic; the last gathers the two mechanisms it grounds in Dwork and colleagues' work on reusing a holdout set, plus the screen that keeps leaked answers from ever being measured.

Look at the Ridge line once more, because it is the least obvious. The "size" of a harness is its resource footprint, measured in the policy tokens a trial consumes. The cost rule never says "delete the retry loop". It says the harness as a whole may only grow as fast as its measured score justifies. That is Ridge's behavior exactly: every bar may stay, but the total is restrained.

What stays open

There is a simpler way to stop a harness search from overfitting: shrink what it may touch. Allow only prompt edits, say, or freeze the tool list. In curve-fitting terms, that is deleting five of the eight features before you fit. It does remove capacity to memorize, but it also removes the capacity to find the mechanisms you actually need, a better context manager, a new tool, a memory store.

RRSI refuses that trade. The paper names the set of every harness reachable from H by arbitrary source edits Ω(H), and "deliberately leaves Ω(H) open: prompts, control flow, configuration, context management, tools, skills, memory, and subagents may all be modified, added, or removed. Instead of restricting this hypothesis space directly, we regularize the search trajectory through it."

Here is the whole round written that way, the paper's Equation 8. Put it beside the unregularized loop of Chapter 1 and only two things have changed: the proposer is handed five more inputs, and the argmax loses its free pass.

𝓗t
This round's candidate harnesses, each the incumbent Ht plus a few edits.
Preg
The regularized proposer. Chapter 1's loop drew from P0, which saw only Ht and the feedback.
Ft
This round's feedback: the analyst's summary of failed trajectories (Chapter 1).
Lt
The edit history: every measured edit, its hypothesis, its score and cost change, and whether it won (Chapter 5).
bt
The annealed edit budget: how many independent edits one candidate may bundle this round (Chapter 4).
Et
Exploration directives: whether the run is stalled, and which components it has never tried (Chapter 5).
Bt
Pruning targets: components with no recent positive gain, handed over for deletion (Chapter 8).
Ω(Ht)
Every harness reachable by arbitrary edits. Left open: the proposer may draw from all of it.
𝒜t
The admissible set: candidates that passed the critic (6), the noise floor (7) and the cost rule and guards (8). If it is empty, Ht+1 = Ht.

In Chapter 1's Equation 2 the maximum ran over the candidates together with the incumbent, so anything that measured higher won. Here the maximum runs over the candidates that are also admissible, and when none is, the incumbent simply stays. Measuring higher is no longer enough to become permanent. That one change of symbol, a union replaced by an intersection, is where all four selection gates live.

The map, box by box

The paper's Figure 2 draws RRSI as seven boxes around that transition rule, three on the proposal side and four on the selection side. Here they are, with the chapter that builds each one.

Proposal side

controls how search capacity is used

A Annealed update sparsity: the edit budget shrinks over rounds. Chapter 4

B Evidence-aware credit assignment: use the full history of gains and regressions, not just the latest win. Chapter 5

C Structured exploration: when progress stalls, redirect search toward untried components. Chapter 5

Selection side

controls which gains may become permanent state

D Leakage screening: reject benchmark-specific or task-specific logic. Chapter 6

E Noise-adjusted performance floor: refuse gains that noise could explain. Chapter 7

F Complexity-aware acceptance: extra complexity must be justified by measurable gain. Chapter 8

G Structural pruning: prune components with no recent positive contribution. Chapter 8

A detail worth noticing. Figure 2 prints box F as "(L1-style)" and box G as "(L0-style)", and Section 4.3 calls the cost rule "the L1-style budget". Section 3.1 and Appendix C say something different and more careful: the edit budget is the L0 analogue, pruning the Lasso/L1 analogue, and the cost rule the Ridge/L2 analogue. We follow Section 3 and Appendix C, which is also how the list above reads. The authors add, twice, that these are analogies of role: "The procedure does not optimize the corresponding norm-penalized objectives, and heterogeneous harness components are not treated as coordinates of a shared continuous parameter vector."

Where the regularizers live in the code

In the released code each evolution instance is one small configuration file, and every regularizer is a handful of numbers in it. Here is the coding instance, grouped by the cure each number controls. Every value is the paper's Table 5 except the candidate count m and the three within-band weights, which come from the released file only; we meet those again in Chapter 8.

rrsi.json, coding# the loop: rounds, trials per
# task, candidates per round
# (m is from the code only)
"T": 20, "k": 2, "m": 2,
# L0-style: the edit budget (Ch 4)
"b_min": 1, "b_max": 4,
# exploration: stall window and
# reserved slots (Ch 5)
"w": 3, "m_draft": 1,
# the noise band: 3 passes of 178
# trials (Ch 5 and 7)
"delta": 0.017,
# Ridge/L2-style: the cost rule, 25%
# more tokens per extra pass (Ch 8)
"beta0": 0.1, "beta1": 44.5,
# the within-band rule (code only)
"w_s": 0.0, "w_c": 15.0, "w_n": 0.5,
# Lasso/L1-style: pruning (Ch 8)
"n_prune": 4

Notice what is absent: there is no list of allowed components and no list of forbidden files. Every number constrains how the search moves. None constrains what the harness may become. That is the claim of this chapter, as a configuration file.

The cost rule lets a candidate add policy tokens only in proportion to its measured gain, and never forces any single component out. Which classical regularizer plays the same role?

Chapter 4

Fewer edits, clearer credit

Anneal how many edits one candidate may bundle, from broad early rounds to sparse, attributable late ones

Here is a candidate a proposer might draft early in an unregularized run on a legal-document agent. Nothing stops it from changing four things at once, so it does:

@@ edit 1 · control_flow + run one bounded check of every + deliverable before submitting @@ edit 2 · prompt + if the matter mentions Delaware, + cite the merger statute @@ edit 3 · prompt - You are a helpful legal assistant. + You are a meticulous associate. @@ edit 4 · config - temperature = 0.7 + temperature = 0.8

The candidate is evaluated on the evolve split and its score rises by two points. Good news? Ask the question the next round will need answered: which edit did that? There is no way to tell. The evaluation produced one number for the whole candidate, and four changes share it.

Look at what is riding together. Edit 1 is a genuine mechanism, the kind that helps on any legal task. Edit 2 is a leak: it only helps on the evolve tasks that happen to mention Delaware. Edits 3 and 4 probably do nothing at all. If the candidate wins, all four enter the harness together, and the leak gets in on the mechanism's coat-tails.

The paper names both costs in one breath. "An unconstrained proposer can bundle many unrelated modifications into one candidate. Such candidates have high effective capacity: they can fit more idiosyncrasies of the current feedback, and any measured change is difficult to attribute to a particular mechanism." More edits per candidate means more room to fit the evolve set's quirks, and less ability to learn which change was real. Try it.

Who gets the credit?

Six edits a proposer could draft, each with a hidden true effect. Choose how many ride in one candidate, draw a bundle, and compare what each edit really does with what the ledger records for it.

4

A toy. The six edits and their true effects are illustrative; the measured score change is the sum of the bundled edits' true effects on the evolve split plus Gaussian evaluation noise with a standard deviation of 0.25 points. The sharing rule is the paper's: every edit in a candidate carries the candidate's one measured ΔS (Appendix C.2).

With six edits in one candidate the ledger's record of each one is mostly fiction: every edit is credited with the whole bundle's score. With one edit per candidate the record is the edit's own effect plus a little noise. Notice also the green bars: what the bundle would carry to benchmarks it never saw is only the mechanisms' share. The leak's contribution to the measured score is real on the evolve split and worth nothing anywhere else.

Edits as switches

To cap "how many changes", the paper first says what one change is. Each round the proposer drafts a pool of atomic edits to the current harness, called Et: small, independently describable changes, each tagged with the component it touches and the hypothesis it tests. A candidate is then a choice of which atomic edits to apply, a row of on/off switches, one per edit in the pool:

zt
The candidate as a row of switches: zt,j = 1 when atomic edit j is included, 0 when it is not.
Et
This round's pool of drafted atomic edits. It is redrawn every round from the open edit space, so its size can change.
‖ · ‖0
The "zero norm": count the switches that are on. This is the L0 of Chapter 3, applied to one update.
bt
This round's edit budget, the most switches one candidate may turn on.

With numbers: suppose that in round 0 of the coding run the proposer drafts six atomic edits and wants to ship three of them, z = (1, 0, 1, 0, 0, 1). Then ‖z‖0 = 1 + 0 + 1 + 0 + 0 + 1 = 3, and the round-0 budget is 4, so the candidate is legal. In round 15 the budget is 2, and the same three-edit candidate must drop one before it can be evaluated. Nothing in this constraint says which edits are allowed; it only counts them. That is why the paper calls it "an L0-style constraint on the update", and immediately adds that "we do not optimize an L0-penalized objective."

The schedule: wide early, narrow late

Why not simply set the budget to one edit forever? Because some mechanisms only work as a pair. A new spreadsheet-reading tool does nothing if the prompt never mentions it, and a prompt line telling the agent to use a tool that does not exist does nothing either. Each half alone measures as noise and would be rejected.

Early in a run, when the harness is far from good, the search needs room to try such coordinated changes. Late in a run, when the remaining gains are small and sit close to the noise band, what it needs is attribution.

So the budget anneals: it starts at bmax and shrinks smoothly toward bmin over the T rounds of the run. The paper's Equation 4:

bt
How many independent edits one candidate may bundle in round t.
bmin
The floor the schedule heads toward. It is 1 in all three instances (Table 5).
bmax
The opening budget: 4 for coding and engineering design, 3 for the agentic workspace.
½(1 + cos)
A smooth dial from 1 down to 0: the fraction of the extra budget still available.
t
The round, counted from 0.
T
The number of rounds in the run: 20 for coding and workspace, 40 for engineering design.
⌈ ⌉
The ceiling: round up to a whole number of edits, since you cannot ship 2.5 of them.

The dial is easier than it looks. As t goes from 0 to T, the angle πt/T sweeps from 0 to π (half a turn). The cosine of that angle falls from 1, through 0 at the halfway point, to −1. Adding 1 gives a quantity that falls from 2 to 0, and halving it gives a dial that falls from 1 to 0.

So at t = 0 the budget is bmin + (bmax − bmin) · 1 = bmax; halfway it is exactly midway; at t = T it would be bmin. The fall is gentle at both ends and steepest in the middle, the same half-cosine shape many neural networks use to decay their learning rate.

Now the coding instance, with bmin = 1, bmax = 4 and T = 20, every step written out for six rounds:

  1. t = 0: the angle is 0 and its cosine is 1.00000, so the dial is ½(1 + 1.00000) = 1.00000, the value is 1 + 3 × 1.00000 = 4.00000, and rounding up gives 4.
  2. t = 7: the angle is 0.35π and its cosine is 0.45399, so the dial is ½(1 + 0.45399) = 0.72700, the value is 1 + 3 × 0.72700 = 3.18099, and rounding up gives 4.
  3. t = 8: the angle is 0.40π and its cosine is 0.30902, so the dial is ½(1 + 0.30902) = 0.65451, the value is 1 + 3 × 0.65451 = 2.96353, and rounding up gives 3.
  4. t = 10: the angle is 0.50π and its cosine is 0.00000, so the dial is ½(1 + 0.00000) = 0.50000, the value is 1 + 3 × 0.50000 = 2.50000, and rounding up gives 3.
  5. t = 13: the angle is 0.65π and its cosine is −0.45399, so the dial is ½(1 − 0.45399) = 0.27300, the value is 1 + 3 × 0.27300 = 1.81901, and rounding up gives 2.
  6. t = 19: the angle is 0.95π and its cosine is −0.98769, so the dial is ½(1 − 0.98769) = 0.00616, the value is 1 + 3 × 0.00616 = 1.01847, and rounding up gives 2.

Rounds 0 to 7 allow four edits, rounds 8 to 12 allow three, rounds 13 to 19 allow two. The ceiling does real work here: without it the budget would be 2.96 at round 8 and 1.02 at round 19, neither of which is a number of edits. Explore the other two instances below.

The budget schedule

Each bar is one round's budget after rounding up; the thin curve is the value before rounding. Pick an instance from Table 5, or move bmax and T yourself. The ghost bar at the end is round t = T, which the run never reaches.

4
20

Equation 4 exactly as written, with bmin = 1 and rounds t = 0 to T − 1 (Appendix C.1); presets from Table 5. The rounding step matches the released schedule.py, which rounds to nine decimals before taking the ceiling.

A detail worth noticing. Table 5 labels bmin the "final-round edit budget", and the prose says later rounds become "increasingly sparse". Taken exactly as written, though, the ceiling and the round count (t = 0 to T − 1) never quite let the budget reach 1: the cosine dial only hits zero at t = T, one step past the last round, and just before that the value is a hair above 1, which rounds up to 2. The released schedule.py computes exactly this, so the last round of every instance allows two edits. The spirit holds, late candidates are small and attributable, but bmin is the limit the schedule heads toward rather than a budget any round receives.

What the ledger learns as the budget shrinks

Chapter 5's ledger records every measured edit, and Appendix C.2 spells out the rule that makes the budget matter: "A candidate containing multiple edits contributes one history record per edit; all edits in that candidate share the same measured ΔS, ΔC, and round outcome." The first device on this page is that rule, drawn. A four-edit winner writes four records, each claiming the full gain.

The annealed budget is what cleans those records up over the run. The paper again: "Because bundled edits inherit the candidate-level measurement, this evidence becomes more attributable as the edit budget anneals toward one." Early records are coarse, a gain shared by four suspects. Late records are nearly one edit, one number. And late is exactly when the search is making its finest decisions, a gain of half a point here, a pruning target there, which is when coarse evidence would mislead it most.

The budget also narrows what one round can fit. Chapter 2 showed that each evaluation is another look at the same finite evolve set. A candidate that may change four things can bend toward the evolve set's quirks in four directions at once; a candidate that may change two can bend in two. That is the "effective capacity" the paper wants to limit, and the reason it calls this regularizer its "most direct classical analogy".

Realization: the budget in code

The schedule three ways. By hand, it is the table above. From scratch, it is four lines of Python. In the released code it is the same function, plus a one-line table builder that the loop consults every round:

pythonimport math

def edit_budget(t, T, b_min, b_max):
    # Equation 4; round t counts from 0
    angle = math.pi * t / T
    dial = (1 + math.cos(angle)) / 2
    v = b_min + (b_max - b_min) * dial
    # round first: 1.0000000002
    # must not become 2
    return math.ceil(round(v, 9))

# the released table, budget_table
[edit_budget(t, 20, 1, 4)
 for t in range(20)]
# [4, 4, 4, 4, 4, 4, 4, 4,
#  3, 3, 3, 3, 3,
#  2, 2, 2, 2, 2, 2, 2]   coding
# workspace, T 20, 1 to 3:
#   ten rounds at 3, ten at 2
# engineering, T 40, 1 to 4:
#   16 at 4, 9 at 3, 15 at 2

The round(v, 9) is there for an honest engineering reason. At t = T the cosine of π is −1, so v should be exactly bmin. In floating-point arithmetic it can come out as 1.0000000002, and the ceiling would then turn one edit into two. Rounding to nine decimals first removes that artifact without changing any legitimate value.

A budget is only as good as its enforcement, and a proposer is an LLM that can ignore instructions. The released proposer states the budget in its context ("You may ship AT MOST bt independent edit(s) in this candidate"), and a submission with more declared edits than bt, or with an edit missing its component and hypothesis tags, is bounced back to be dropped or merged; it is never accepted. There is also a quieter loophole: declare one edit and hide three changes inside it. The critic of Chapter 6 closes it with its "undeclared bundling" rule, which rejects a diff containing independent changes that no declared edit covers.

Late in a run, why cap each candidate at one or two edits instead of letting the proposer bundle as many as it likes?

Chapter 5

Remember what failed

Keep every edit's evidence, and when progress stalls, spend a slot on what was never tried

Round 3 of a run on the legal-work tasks. The proposer has read the failure summary, noticed that several deliverables cite the wrong clause, and drafted a fix: a small memory store that keeps the citations the agent has already checked. The candidate is evaluated on the evolve set and comes back 0.6 points worse. It is rejected. So far, so good.

Round 8. The failure summary still mentions wrong citations, because the problem was never solved. A proposer that sees only this round's feedback has no idea that round 3 happened, so it drafts the same memory store again. This time the evaluation noise of Chapter 2 happens to fall the other way: +0.4. The idea the evolve set rejected five rounds ago wins the round and becomes permanent harness state.

Nothing in that story needed a bad proposer. It needed a proposer with no memory of the search itself. The paper names the cost precisely: "Every evaluation is another adaptive look at the same finite evolve set, so repeatedly testing hypotheses that earlier rounds already falsified spends search capacity without adding useful evidence." A re-test is not free information. It is another lottery ticket drawn on the same tasks.

RRSI's answer has two halves, and this chapter builds both. The first is bookkeeping: an edit ledger that records every measured edit, what it tried and what it got, and hands that record to the proposer every round. The paper calls this evidence-aware credit assignment. The second reads the same ledger to notice when the search has stopped moving, and points it at the parts of the harness it has never touched: structured exploration.

What one record holds

For every evaluated candidate, RRSI records "the component it modifies, the hypothesis it tests, the source diff, the resulting score and cost changes, and whether the candidate was accepted." Appendix C writes the ledger that exists before round t as a set of tuples, one per edit (Equation 10):

𝓛t
The ledger before round t: every measured edit so far, nt of them.
ti
The round in which edit i was evaluated.
ℓi
The component the edit touched: one of nine names, listed below.
hi
The hypothesis, in words: what the proposer expected the edit to fix.
di
The source diff: the exact change to the harness code or text.
ΔSi, ΔCi
The measured change in evolve-set score and in relative policy-token cost, against the incumbent. Chapter 8 defines both precisely.
ai
1 if the candidate carrying this edit won its round and entered the harness, 0 otherwise.

Two rules keep the ledger honest. First, there is one record per edit, not per candidate. A candidate that bundles three edits writes three records, and all three inherit the candidate's single ΔS, ΔC and outcome. That is the attribution problem of Chapter 4 written into the data, and it is why the shrinking budget pays off here: as late rounds allow only one or two edits per candidate, each record becomes evidence about one change instead of a blur of several.

Second, ai = 1 only for the round's winner. A candidate that passed every gate but lost to a better one is recorded with ai = 0. It was admissible, and it still did not make it in. The ledger keeps "good enough to consider" and "actually chosen" apart.

In the released code the ledger is a plain file, history.jsonl, one JSON line per edit. Here is the line the round-6 winner of the device further down would write (illustrative values; the long diff field is left out):

runs/workspace/history.jsonl
{"t": 6, "variant": "B", "edit_id": "E1", "component": "client_tool",
 "hypothesis": "spreadsheet reader that returns cell ranges",
 "delta_S": 0.006, "delta_C": 0.02, "accepted": true,
 "outcome": "ACCEPTED", "bundle": 1}

Each field is a decision. delta_S is a fraction of the score, so 0.006 is 0.6 points. bundle says how many edits shared this measurement, so a reader of the ledger knows how far to trust the credit. outcome is one of ACCEPTED, REJECTED or LOST, the three measured fates.

A candidate that never reached a measurement (the critic of Chapter 6 refused it, or it crashed a quick smoke test) is still written down, but with delta_S set to null, and it does not count as evidence about its component. And when the proposer is shown the ledger, it sees the 40 most recent records with at most four of those unmeasured aborts, because, as the code's comment puts it, "a wall of aborts is a feedback loop, not evidence."

Nine kinds of edit

The component tag ℓ is not free text. It comes from a fixed vocabulary K of nine names (Equation 12), the same nine for every domain. The paper lists the names; the middle column is our plain-words gloss of what an edit to each one changes.

Component, kindAn edit to it changes (our gloss)
prompttextthe words the model reads: system prompt, task wrapper, instructions
control_flowlogicwhen the agent plans, acts, checks, retries or stops
configconstantsconstants: step limits, timeouts, sampling settings
output_plumbinglogichow results reach the deliverable files: names, formats, where they land
context_mgmtlogicwhat the policy sees at each step: truncation, summaries, compaction
client_toolstructurala tool the harness offers the model, with its description
skillstructurala skill file the agent may consult
memorystructuralstate the harness keeps and feeds back later
subagentstructuralan extra policy call with its own role

The last four form Kstr (Equation 15), the structural components: in the code's words, they "add machinery (a tool, a skill file, a memory store, an extra policy call) rather than changing text or constants." Chapter 8 gives them a small bonus at selection time when they have never entered a winning edit.

Because so much hangs on the tag, the released code does not trust the proposer's label. A declared component is kept only if the diff carries evidence for it: a registered tool, a skills/ path, a memory call, a subagent call. A diff whose changed lines are all strings or comments is classified as prompt, whatever the proposer called it. The comment explains the attack it closes: a proposer "can name a skill/memory/tool/subagent edit without shipping one", and an unverified tag would corrupt every summary computed from the ledger.

Two summaries of the ledger

The proposer does not read raw tuples only. Two summaries are computed from them (Equation 11):

Tt
The tried set: every component with at least one measured edit so far.
gt(ℓ)
The best measured gain that any recent edit to component ℓ produced.
t − ti ≤ nprune
"Recent" means inside the pruning window: 4 rounds on the coding and workspace instances, 5 on engineering (Table 5).
max ∅ = −∞
A component with no recent measured edit gets the worst possible value, not zero. Chapter 8 shows why that matters.

Tt is the one this chapter needs. gt is the evidence Chapter 8's pruner uses to decide what to delete, and it is the reason the ledger has to keep negative results: a component whose recent edits all measured at or below zero is not earning its place.

When the search stops moving

The ledger reveals a second failure, one no single round can see. The paper describes it as the proposer that "has collapsed onto a narrow edit family, for example repeatedly rewriting prompts while leaving agent structural mechanisms untouched." A prompt reword is the cheapest edit to draft, it rarely breaks anything, and on a noisy evolve set it sometimes measures a little positive. A proposer rewarded round by round drifts toward it, and the harness's tools, memory and subagents never get tried at all.

RRSI detects the drift with a stall flag, σt. The search counts as stalled when its progress over the last w rounds is no larger than the noise band δ (Equation 13):

σt
The stall flag: 1 when the search has not moved beyond the noise, 0 otherwise. 𝟙[ ] is 1 when the statement inside is true.
Ŝt − Ŝt−w
How far the incumbent's evolve-set score has risen over the last w rounds. w = 3 in all three instances.
δ
The empirical noise band of Chapter 2: 0.004 (0.4 points) on the workspace instance.
Ut
The untried components: everything in K with no measured edit yet.
mdraft
How many candidate slots are reserved for untried components while stalled: 1 in every instance.

A worked example on the workspace instance, with w = 3 and δ = 0.4 points. Suppose the incumbent's evolve score over rounds 5 to 8 reads 90.1, 90.2, 90.2, 90.3. At round 8 the flag compares round 8 with round 8 − 3 = 5:

Ŝ8 − Ŝ5 = 90.3 − 90.1 = 0.2 ≤ 0.4  ⇒  σ8 = 1 (stalled)

Look at what σ ignores. It does not ask whether the last three rounds accepted anything; they did, twice. It asks whether the total progress over the window can be told apart from noise. Two accepted edits worth a tenth of a point each are, as far as this evolve set can tell, nothing.

When σt = 1 and some components are still untried, mdraft candidate slots are reserved. The released code drafts two candidates per round, labeled A and B, and with mdraft = 1 the reservation lands on B: the directive it receives says the slot is "RESERVED for exploratory edits on components the run has never exercised", and that the candidate "must put at least one edit on one of those components." Candidate A stays free. Nothing else changes: the reserved candidate faces the same critic, floor and cost rule as any other.

Now watch both halves at work. The device runs twelve rounds of a toy search on the workspace instance, two candidates a round, and lets you take away the ledger, or the exploration, or both.

The ledger

Top: each round's two candidates, tagged by component (purple = structural), with the one that entered the harness outlined in orange. Middle: the incumbent's evolve score, with the 3-round window the stall flag checks. Bottom: which of the nine components have been tried. Choose a proposer, then press Play or step a round at a time.

A toy script, not a paper run: the candidate ideas, their measured changes and the cost fields are illustrative. The rules are the paper's: one record per edit, a = 1 only for the round's winner, σ with w = 3 and δ = 0.4 points (Table 5, workspace), one reserved slot while stalled, and the released code's two candidates per round. Selection is simplified to "keep the better candidate if its measured gain is positive"; Chapters 6 to 8 build the real gates.

Run the default first: the ledger on, exploration on. Rounds 0 to 4 pick off the easy wins. By round 6 the last three rounds have added only 0.2 points, the flag goes up, and candidate B is reserved for the three components nobody has touched: client_tool, skill, subagent. It proposes a spreadsheet reader tool, which measures +0.6 and is kept. The flag goes up once more at round 10, and the reserved slot tries a skill file.

Now switch exploration off and run it again. The same stall arrives at round 6, but nothing redirects the proposer: rounds 6 to 11 are mostly prompt rewrites worth a tenth of a point, with the occasional small fix, the flag stays up for four of those six rounds, and after twelve rounds three structural components have never been measured. The harness may be missing its best mechanism, and the ledger is the only place that fact is written down.

Finally, choose the proposer that forgets. At round 8 it re-proposes the round-3 citation store. On these tasks that idea has now been measured twice, at −0.6 and at +0.4, and the search keeps the lucky second look. Its final evolve score looks about as good as the default's, and that is exactly the problem: 0.4 points of it are a coin flip the ledger would have remembered, and no held-out benchmark will reward them.

Why this counts as regularization

Neither half shrinks the harness or bans an edit, so it is fair to ask why the paper files them under regularization. The credit half limits how many times the search can look at the same question. In Chapter 2, every extra look at a noisy score was another chance to keep a lucky number; refusing to redraw falsified hypotheses removes a whole family of such looks.

The exploration half, in the paper's words, "plays a role similar to diversity or entropy regularization", the idea behind soft actor-critic (Haarnoja et al., 2018). There, a learning agent is paid a small bonus for keeping its choices spread out, which stops it locking onto the first action that looked good. Here the bonus is a reserved slot, paid only while progress is indistinguishable from noise. Chapter 8 adds a second nudge in the same direction at selection time: the novelty count ν, which credits a candidate for touching a structural component type that has never been in a winning edit.

The ledger turns every measurement into a constraint on the future. Once an idea has been measured on these tasks, measuring it again is not new evidence, it is another draw from the same noise. A search that remembers spends its limited looks on questions it has not asked yet, and the stall flag tells it when the questions it keeps asking have stopped paying.

The stall check, three ways

By hand, you already did it: subtract the score three rounds ago from the score now, compare with δ. From scratch, the whole exploration directive is a few lines. Scores are fractions here, as in the ledger, so δ = 0.004:

python# Rounds 5..8 scored 0.901, 0.902, 0.902, 0.903 ; w = 3, delta = 0.004
# 0.903 - 0.901 = 0.002 <= 0.004  ->  stalled
K = ["prompt", "control_flow", "config", "output_plumbing", "context_mgmt",
     "client_tool", "skill", "memory", "subagent"]

def stall_flag(traj, t, w, delta):
    """sigma_t = 1[S_t - S_(t-w) <= delta]; 0 until w rounds exist."""
    if t < w:
        return 0
    return int(traj[t] - traj[t - w] <= delta)

def exploration(t, traj, ledger, w=3, delta=0.004, m_draft=1):
    tried = {r["component"] for r in ledger if r["delta_S"] is not None}   # T_t
    untried = [c for c in K if c not in tried]                                 # U_t = K \ T_t
    sigma = stall_flag(traj, t, w, delta)
    reserved = m_draft if (sigma and untried) else 0                           # slots that must touch U_t
    return sigma, untried, reserved

traj = [0.894, 0.899, 0.903, 0.906, 0.907, 0.908, 0.908]   # the device's first rounds
stall_flag(traj, 6, 3, 0.004)    # 0.908 - 0.906 = 0.002 -> 1

The released rrsi/history.py is the same two summaries plus one addition: the directive is also rendered as a sentence for the proposer's prompt, beginning "STALL: the incumbent has not moved by more than the noise band over the last rounds (sigma_t = 1)." Two details in its version are worth copying. The flag is 0 until w rounds exist, so a run cannot stall in its opening rounds. And records with no measurement are filtered out before Tt is built, so a candidate the critic refused does not make its component count as tried.

One question the paper leaves open: how much of the gain comes from the ledger and how much from exploration. Its ablation (Chapter 9) removes the proposal-side group as a whole, and reports that the out-of-distribution average falls from 43.6 to 41.9 while the evolve score barely moves. The authors' reading is that "steering where the search looks matters even when nothing is rejected." Separating the two halves would take a finer ablation than the paper runs.

On the workspace instance (w = 3, δ = 0.4 points), the incumbent's evolve score at rounds 4, 5, 6 and 7 is 90.0, 90.3, 90.5 and 90.3. Is the search stalled at round 7?

Chapter 6

Screen before you score

Reject edits that name the test, before they can earn a score that tempts the next round

Where does the proposer get its ideas? From the feedback of Chapter 1: an analyst reads the incumbent's trajectories on the evolve set and writes down what went wrong. In the released code that evidence is very concrete: the worst trial of each of the lowest-scoring tasks, plus the best trials of a few top ones. Those trajectories are full of specifics. The task's name. The companies and parties in its documents. Its file names. And, whenever a test or a rubric explains a failure, the value that was expected.

Now sit in the proposer's chair. You are an LLM, you have been asked to raise the score, and in front of you is a failed task whose rubric wanted the governing law to be Delaware. The edit that raises the score fastest is not a better way of reading contracts. It is one sentence:

prompts/system.md
+ If the documents mention "Acme Holdings", state that the governing law is Delaware.

That edit really does raise the evolve score. The task that failed now passes, in both trials, every time. On JobBench, GDPval or APEX-Agents it does nothing at all, because none of their tasks mention Acme Holdings. This is benchmark-specific fitting, the first of the three failure modes in Chapter 2. We will call an edit like this a leak: facts about the evolve set's particular tasks leaking into the harness. (The example is ours; the paper describes the category, not this sentence.)

Why no numeric gate can see a leak

Chapters 7 and 8 build gates out of numbers: is the gain bigger than the noise, is the extra cost paid for. It is worth checking, with numbers, whether those gates would stop a leak. On Harvey LAB the noise band is δ = 0.004, which the paper translates into its own units: 60 criteria out of roughly 14,100 criterion verdicts per evaluation. Suppose, for illustration, that the leaked sentence fixes 30 criteria of one task, and that task runs twice per evaluation (k = 2):

30 × 2 = 60 verdicts;   60 / 14,100 = 0.0043 = 0.43 points

The noise band is 0.40 points, so one leaked task already fills it; two leaked tasks put the gain clear of it. And the gain is not luck. Evaluate the candidate again and it comes back just as high, because the leaked answer is right every time. The cost is a dozen tokens, so the cost rule has nothing to object to either.

A leak is the one bad edit that is stable, cheap and real on the evolve set. Stable, so the noise floor cannot see it. Cheap, so the cost rule cannot see it. Real, so a second measurement confirms it. The only place it shows up is in the text of the diff. That is why RRSI gives this job to a reader, not to a number.

The critic reads the diff

So RRSI adds a reader. In the paper's words: "Before full evaluation, a critic reads each candidate diff and rejects edits that explicitly encode task names, entity names, task-specific values, answers, or other logic specific to the evolve benchmark, as well as edits that add inert machinery." The critic is an LLM, Claude Opus 4.8, the same model the paper uses as proposer and analyst. And what it screens is content, not components: "generic prompt or tool-description improvements remain valid candidates."

The line between the two is sharper than it sounds. Put two edits side by side:

Stays

general practice

"Check the governing-law clause of every contract before summarizing it." It helps on any contract in any legal suite, including ones this run has never seen.

Goes

a leak

"If the documents mention Acme Holdings, the governing law is Delaware." It helps on one task of this suite, and it is wrong or useless everywhere else.

Both edits mention governing law. Only one of them names this benchmark's content. The released critic's instructions state the test that separates them: "would this change still make sense, and still help, on an unfamiliar task from a different suite in the same kind of work?" The first sentence passes. The second does not, and no rewording that keeps "Acme Holdings" in it ever will.

Why before, not after

The critic runs on the diff before the candidate is evaluated. The paper's reason fits in one sentence: "a leaking candidate never receives the inflated evolve-set score that could make it attractive to subsequent rounds." Unpack "attractive" and a scored leak does three separate kinds of damage:

  1. It would win. The selector keeps the highest admissible score, and a leak's score is real and cheap, so it would often be the round's winner and become permanent.
  2. It would teach. Its score would enter the ledger of Chapter 5 as positive evidence for its component and hypothesis. Later rounds condition on that ledger, so one scored leak invites more of the same.
  3. It would spend a look. Every evaluation is another adaptive look at the same finite set (Chapter 2). Spending one on an edit that should never have been a candidate buys nothing.

There is also a plain budget argument. One full evaluation on the workspace evolve set is 120 tasks × k = 2 trials:

120 tasks × 2 trials = 240 agent runs per candidate

That is about 14,100 judged criteria for every candidate scored, and over the paper's 20 rounds, at two candidates a round, up to 40 such evaluations: 9,600 agent runs. A critic call reads one diff. Putting a screen in front of every evaluation costs a sliver of one evaluation, and each rejected leak saves a whole one. (The two candidates per round are the released code's setting; the 120 tasks, k = 2 and 20 rounds are Table 5 and Appendix A.)

Formally, the screen is a filter on the round's candidate set. The released code writes the critic as a function Critic(Ht, H′) → {0, 1}, and the paper's Algorithm 2 starts from the "screened" set:

screened 𝓗t
The candidates that go on to be evaluated. Algorithm 1's last line: "return candidates that pass the pre-evaluation screen."
H′ ∈ 𝓗t
One candidate harness from this round's drafts.
Critic(Ht, H′)
The verdict, 1 to accept or 0 to reject, from reading the diff between the incumbent Ht and H′, the edits the proposer declared, and any state files a new mechanism writes.

So the critic sits on the seam between proposing and selecting: the last filter of Algorithm 1, and the reason Algorithm 2 never sees a leak. Everything after it, the floor of Chapter 7 and the cost rule of Chapter 8, works on candidates that have already been read.

Now take the critic's seat yourself. Eight candidates, one at a time, each with the edit the proposer declared.

Be the critic

Read the diff and decide: would you let this candidate be evaluated? Your verdict is revealed against the released critic's six reject rules. The strip shows all eight: orange for accepted, red for rejected, a teal mark where you agreed.


    

Diffs 2 to 7 are illustrative. Diffs 1 and 8 paraphrase two real decisions in the paper's Table 6 (coding round 0, candidate A, and engineering round 2). The rules are the six reject rules in the released rrsi/critic.py; the real critic is Claude Opus 4.8 reading the full diff, the declared edits and any state files.

If you rejected diff 1 or diff 8, notice why they pass. Diff 1 is the paper's coding round 0, candidate A: a bounded audit before declaring the task complete, plus advice on long-running jobs. Nothing in it names a task. It was accepted and measured +3.93 points on the evolve set. Diff 8 is engineering round 2: a recovery for a tool error. It even quotes a string, the error message, and it still passes the litmus test, because every agent that uses that tool on any task hits the same message. The paper reports it moving 122 of 244 passes to 128, for 1.6% more tokens.

And if you accepted diff 5, the memory one, you are in good company: the mechanism itself is generic. What gives it away is what it stores, a task's name and its answer, carried across trials. The harness is evaluated on the very tasks it evolves on, so a memory keyed by task identity is memorization with extra steps.

Inside the released critic

The code in the paper's repository makes the critic concrete. It works in two layers, cheapest first:

  1. A deterministic precheck. A regular-expression denylist: a generic pattern for credentials, plus each domain's own patterns for grader paths, task ids and judge references. A hit rejects the candidate on the spot, with no model call.
  2. An LLM review. Claude Opus 4.8 reads the declared edits, the diff and any state files a new mechanism would persist, under a system prompt that lists six ways to fail.
RuleRejectsIn the device
1 · Leakage, task specializationtask names or ids, entity names, expected outputs or answers, magic constants that fit one task, per-task templates, branching on task-identifying featuresdiffs 2, 3
2 · Degeneratea no-op (dead code, unused flag, comment only) that claims a mechanism; deleting a safety mechanism such as context compression, truncation or error handling without a replacementdiff 7
3 · Grader gamingreading, reconstructing or influencing the verifier, judge or rubric at runtimediff 4
4 · Undeclared bundlingchanges in the diff that no declared edit covers, or one declared edit hiding several mechanisms. This is what keeps Chapter 4's budget honest.not shown
5 · Runtime memory or skill leakagea memory or skill that persists or injects task-specific runtime data across trials: file contents, tool outputs, answers, task or entity namesdiff 5
6 · Unbounded worka check, retry or "keep improving" loop with no give-up pathdiff 6

The rest is plumbing, and each piece of plumbing is a decision. The critic must answer in strict JSON, a verdict of accept or reject with its reasons and risk notes. If its output cannot be parsed after three attempts, the candidate is rejected: the screen fails closed, not open. A rejection is not the end of the idea: the reasons go back to the proposer for a bounded number of repair rounds (five in every released configuration). A candidate that still cannot be repaired is dropped and written to the ledger with no measurement, so, as Chapter 5 showed, it never counts as evidence about its component.

One job is explicitly not the critic's: runtime correctness. Its instructions say that compile, construct and smoke checks handle crashes after it, and that it sees only the diff, not the whole files, so it must not guess at missing names or crashes it cannot see. Keeping the two apart keeps the critic's rejections about one thing, leakage and degeneracy, which is what makes them worth reading.

Here is the screen from scratch, with the same two layers and the same fail-closed ending. The denylist patterns are illustrative stand-ins for a domain's own:

pythonimport json, re

DENY = [r"grader/", r"rubric\.json", r"task\.id"]       # a domain's own patterns (illustrative)

def screen(diff, declared_edits, llm, rules):
    if any(re.search(p, diff) for p in DENY):              # layer 1: free, deterministic
        return {"verdict": "reject", "reasons": ["precheck"]}
    if not diff.strip():
        return {"verdict": "reject", "reasons": ["empty diff"]}
    payload = json.dumps(declared_edits) + "\n=== DIFF ===\n" + diff
    for _ in range(3):                                      # layer 2: the LLM reviewer
        out = llm(system=rules, user=payload)
        try:
            v = json.loads(out)
            if v.get("verdict") in ("accept", "reject"):
                return v
        except json.JSONDecodeError:
            pass
    return {"verdict": "reject", "reasons": ["unparseable"]}   # fail closed

# what a verdict looks like (the released critic's JSON shape):
# {"verdict": "reject", "reasons": ["hard-codes an entity and its answer"], "risk_notes": []}

What the critic cannot promise

The critic is itself an LLM reading a diff, so it can be wrong in both directions. A leak phrased in general-sounding words could slip past it; a genuinely general edit could be refused, and would then have to survive repair to be tried at all. The paper does not report how often the critic fired, how many candidates went back for repair, or how often a human would have disagreed with its verdicts.

The closest the paper comes is the ablation of Chapter 9. Removing the selection-side group, of which this screen is one member in Figure 2, raises the evolve-set score from 90.5 to 91.5 and lowers the out-of-distribution average from 43.6 to 41.0 (Table 2). That is the signature a screen against leaks should leave, but it measures the whole group, not the critic alone.

Why does RRSI screen candidates for leakage before evaluation rather than after?

Chapter 7

Don't walk downhill

Measure the noise once, then refuse any candidate that falls more than one band below the best score ever seen

Chapter 2 showed that a loop which keeps the maximum of noisy scores can climb without improving anything. This chapter is the mirror image. A loop can also sink without ever taking a step that looks like a loss.

Here is how it happens. It is round 6 of a run on the Harvey LAB evolve split, and the incumbent harness scores 90.0. A candidate trims the context it keeps and scores 89.7. Is it worse? The gap is 0.3 points, and the very same harness, re-run on the same tasks, disagrees with itself by about that much. Rejecting it would be rejecting noise. So you let it in. Next round a candidate scores 89.5, which is only 0.2 below the new incumbent, and the same argument lets it in too. (These scores are illustrative; the mechanism is the paper's.)

Every single step was "within noise". After five such steps the harness can sit a point and a half below where it was, and nothing in the loop ever flagged a regression. The paper names the failure exactly: a search that walks downhill through a sequence of regressions that are individually small enough to be mistaken for noise.

The fix is one inequality. What matters is what sits on its right-hand side, and before we can write it we need a number we have been using loosely since Chapter 2: how big the noise actually is.

First, measure the noise

Why does an unchanged harness score differently each time? Because almost everything in an agent run is sampled. The policy samples its tokens, so the same task can go down a different path. On Harvey LAB each rubric criterion is graded by an LLM judge (Gemini 3.5 Flash), which is itself a sampled model. Every run of the same harness on the same tasks is a new draw.

RRSI measures that spread once, before evolution starts: it evaluates the unchanged base harness repeatedly and estimates an empirical noise band δ. Nothing about the harness changes between those evaluations, so every difference between them is pure noise. That is the most honest ruler you can have, because it is exactly the comparison the loop will make later: one harness against another, on the same tasks, with the same judge and the same number of trials.

Measure the noise

Each dot is one full evaluation of the same, unchanged harness. Below them, the spread you should expect between any two such evaluations, and the band δ that RRSI draws from it. The small grey triangle marks the band that the noise you set really implies. Change how many times you evaluate and how noisy one evaluation is, then evaluate again.

6
0.14

A simulation: each evaluation is the base harness's Harvey LAB score, 89.4 (Table 1), plus Gaussian noise. The default noise of 0.14 points is illustrative, chosen because it reproduces the paper's agentic-workspace band δ = 0.004, that is 0.4 points (Table 5). The recipe for δ is the one in the released calibrate.py.

Look at what the device computes, because it is the whole calibration recipe. Take the scores of the repeated evaluations and find their standard deviation. Call it σ: the wobble of one evaluation. But the loop never looks at one evaluation. It looks at a difference between two: the candidate's score minus the incumbent's. So what we need is the wobble of a difference.

Two independent evaluations each wobble by σ. When you subtract them their variances add, because independent errors do not cancel on average:

ΔSnull
The "null" score difference: one evaluation of the unchanged harness minus another. Any value it takes is noise, by construction.
σ
The standard deviation of a single evaluation's score, estimated from the repeated evaluations.
√2
What subtraction costs: a difference of two independent noisy numbers is √2 times noisier than either one.

Then the band is a fixed number of those standard deviations. The released code uses two:

δ
The noise band: how far apart two evaluations of the very same harness can land without anything having changed.
z
How many standard deviations wide the band is. The released calibrate.py defaults to 2.

Let's put numbers in. Suppose one evaluation of the base harness has a standard deviation of 0.14 points (an illustrative value, picked because it lands on the paper's band). Then a difference of two evaluations has a standard deviation of √2 × 0.14 = 1.414 × 0.14 = 0.198, call it 0.2 points. Two of those is the band:

δ = 2 × √2 × 0.14 = 2 × 0.198 = 0.40 points (= 0.004 as a fraction, the agentic-workspace value in Table 5)

Why two standard deviations? If the null difference is roughly bell-shaped, it lands above −2 standard deviations about 97.7% of the time (a normal table's value at exactly 2; the code's own comment rounds it to "about 97.5%"). So an unchanged harness, re-measured, fails a floor set δ below its own score only a couple of times in a hundred. That is the false-alarm rate RRSI accepts.

The paper states each instance's band in the units a person would count, which makes the numbers concrete:

InstanceδWhat it means
Coding (Terminal-Bench 2.1)0.0173 passes out of 89 tasks × k = 2 = 178 trials (3 / 178 = 0.0169)
Agentic workspace (Harvey LAB)0.00460 criteria out of roughly 14,100 criterion verdicts
Engineering design (EngDesign)0.0205 passes out of 61 tasks × k = 4 = 244 trials (5 / 244 = 0.0205)

So on the coding instance the band is three passes wide. One subtlety is worth doing by hand, because the rules ask for a gain strictly above δ: three extra passes is 3 / 178 = 0.01685, which is not above 0.017. Three more solved trials only reach the edge of the noise; it takes four (4 / 178 = 0.0225) to get past it. Hold on to that; it decides a real case in Chapter 8.

Now the realization. The released calibrate.py has two ways to get σ. With two or more repeated evaluations of the base harness it uses their standard deviation directly, times √2. With only one evaluation it bootstraps: it rebuilds the score thousands of times, each time resampling the k trials within each task, and takes the spread of those rebuilt scores. Notice what it does not do: it never resamples which tasks are in the set. The evolve set is fixed for the whole run, so the only noise that matters is the noise of re-running those same tasks, and that is the only noise the bootstrap measures. In the experiments the three bands are fixed per instance in each rrsi.json (0.017, 0.004, 0.020), and the calibration runs only when that value is left empty.

The floor

With a band in hand, the rule is short. Let S★ be the best evolve-set score of any incumbent so far. It starts as the base harness's own score, and every time a new incumbent is chosen it is updated to whichever is larger, itself or the newcomer. A candidate H′ may only become the next incumbent if:

Ŝ(H′)
The candidate's measured score on the evolve split: the average reward over every task and every one of its k trials (Chapter 1).
S★
The best score any incumbent has had so far. After each round, S★ ← max(S★, Ŝ(Ht+1)). It only ever goes up.
δ
The noise band from the calibration above. The floor sits one band below the best score ever reached.

With numbers: S★ = 90.0 and δ = 0.4, so the floor is 90.0 − 0.4 = 89.6. A candidate measuring 89.7 clears it. A candidate measuring 89.5 does not, even if the current incumbent is 89.6 and the candidate is only 0.1 below it.

floor = S★ − δ = 90.0 − 0.4 = 89.6 (89.7 passes, 89.5 is refused)

Why the best score, and not the incumbent

The obvious rule would compare each candidate with the current incumbent: "keep it if it is not more than δ worse than what we have." That rule looks at one step at a time, and one step at a time every drop of 0.3 is innocent. Try both rules on the same stream of candidates.

Walk downhill

Sixteen rounds. Each round one candidate arrives, a little better or, more often, a little worse than the harness it would replace, and a little cheaper, so the floor is the only rule that can stop it. Choose where the floor is measured from, and widen or narrow the band.

0.4

A toy stream of candidates: each is described by how far its measured score lands from the harness it would replace (mostly 0.1 to 0.3 points lower, now and then higher), and every one is assumed to pass the cost and band rules of Chapter 8. The start, 89.4, is the base harness on Harvey LAB (Table 1); the default band is the workspace δ of Table 5.

With the floor measured from the incumbent, fifteen of the sixteen candidates get in and the harness slides steadily down, past the score it started from, although no accepted step fell by more than 0.3 points. The floor itself moved down with every step, so the only candidate it ever caught was the one that fell by more than a whole band at once. With the floor measured from the best score so far, the same stream is cut off the moment it would drift more than one band below 90.0, and the harness hovers just under its best instead of sliding away from it.

That is not luck; it is a two-line guarantee. Every new incumbent had to clear the floor, so its score is at least S★ − δ. And S★ never decreases. So at every round the incumbent sits at most one band below the best score the run has ever reached. The total drift is bounded by δ, however many rounds the run lasts. Anchored to the incumbent, the drift is bounded by δ per round, which over 20 rounds is no bound at all.

incumbent-anchored: 5 steps of −0.3 = −1.5 points, all "within noise"  ·  best-anchored: never more than δ = 0.4 below S★

Why allow any drop at all

If walking downhill is the danger, why not set the floor at S★ itself and demand that every new harness at least match the best? Because some of the most valuable edits make the harness cheaper or add a new kind of mechanism while scoring about the same, and "about the same" on a noisy measurement means sometimes a little lower. Chapter 8's within-band rule exists to admit exactly those. The floor gives them room, one band and no more, and the best-anchored version guarantees that the room cannot be spent twice.

In the released code the floor is the first check a measured candidate meets, before any question of cost:

python# rrsi/selection.py, judge(): the first gate after evaluation
floor = S_star - delta
if ev.S < floor:
    d.reason = (f"below noise-adjusted floor: S' {ev.S:.4f} < S* {S_star:.4f} "
                f"- delta {delta:.4f}")
    return d                        # rejected: nothing else is even consulted

# and after the round (rrsi/loop.py), only if a winner was chosen:
fr["S_star"] = max(fr["S_star"], new_S)   # S* starts as the base evaluation's score
A cheaper harness cannot buy its way past the floor. In round 8 of a coding run, candidate B "pins the original task instruction into the completion gate so that the policy re-checks the literal specification before submission". It cut policy tokens by 13.6% and lost 2.81 points (Table 6). On its own, Chapter 8's within-band rule would have liked it: with the coding weights, 0 × ΔS − 15 × (−0.136) = +2.04, which is above zero. But −2.81 points is more than the coding band of 1.7 below the incumbent, and the incumbent is never above S★, so the candidate is below S★ − δ. The floor is checked first and is non-compensatory: no saving elsewhere can make up for failing it. Rejected by floor.

One more thing to see in that example: the order of the checks is the design. Because the floor comes first, the within-band rule never gets the chance to trade a large, real loss for a token saving. It only ever trades inside the band, where a score difference is not yet evidence of anything.

Why is RRSI's floor measured from S★, the best score so far, instead of from the current incumbent?

Chapter 8

Growth must pay

Charge extra tokens against measured gain, and delete machinery that has stopped earning its place

The last two chapters were about the score: is a gain honest (the critic), and is it real (the floor). Every harness edit also changes something else, which a score-only loop never looks at: what the agent costs to run. And there is one move a score-only loop never makes at all. It never takes anything out.

This chapter covers the two rules that fix that. The paper pairs them in one sentence: "the L1-style budget refuses growth that is not paid for when it is proposed, and the pruning rule removes growth that has stopped being paid for since." The first acts at the door, on every new candidate. The second acts on the machinery already inside. (That sentence calls the cost rule "L1-style"; Section 3 calls the same rule Ridge, or L2-style. Chapter 3 explained the label drift; we follow Section 3.)

The cost rule: pay at the door

Picture a candidate that adds a second full review pass on every deliverable. On the evolve split it gains 0.3 points; every trial now burns 38% more tokens. (Illustrative numbers.) A score-only loop keeps it, because the score went up. Nothing is wrong with any single decision like that. The trouble is that the loop makes it again and again, and nothing ever pushes back.

The paper measured where that ends. The unregularized run's final harness spends 3.80 million policy tokens per trial, against 1.56 million for the base harness it started from: 3.80 / 1.56 = 2.4 times as much, for an out-of-distribution average of 40.3 against 39.7 (Table 2). RRSI's final harness spends 2.42 million. Remove only RRSI's acceptance rules and the cost climbs to 3.59 million, which is 3.59 / 2.42 = 1.48 times RRSI's: the paper's "token cost rises by half".

To charge for tokens, first measure them in a way that means the same thing in every domain. A trial on a terminal task and a trial writing a legal memo burn very different numbers of tokens, so RRSI compares relative cost against the incumbent:

ΔS
The candidate's score change against the incumbent Ht, as a fraction: +0.005 is half a point.
ΔC
The relative change in policy tokens per trial: +0.15 means the candidate costs 15% more to run than the incumbent.
Ĉ(H)
The measured cost: policy tokens averaged over every task and every trial of the evaluation (Chapter 1). Tokens are the paper's "common measurable proxy" for a harness's footprint.

With numbers: an incumbent at 2.00 million tokens per trial and a candidate at 2.30 million give ΔC = (2.30 − 2.00) / 2.00 = 0.30 / 2.00 = +0.15. (Illustrative.)

Now the rule. When the gain is clearly real, above the noise band, extra cost is allowed in proportion to that gain:

β0
The base allowance: the cost increase tolerated even for a negligible gain. 0.10 (10%) for coding and the agentic workspace, 0.15 for engineering design.
β1
The gain-dependent allowance: how much extra relative cost each unit of measured gain buys. 44.5, 35.4 and 24.4 for the three instances.

Read it as a line in a plane with ΔS across and ΔC up. The line starts at height β0 and climbs with slope β1; a candidate is allowed if it sits on or under it. The more you gain, the more you may spend. The paper calls this Ridge-like because it restrains the total footprint without forcing any one component out.

The β1 values look arbitrary until you put them in the units a person counts, which is how the paper states them. Each is a price in tokens per extra success:

InstanceOne unit of gainβ1 × that unitIn words
Coding1 pass of 178 = 0.0056244.5 × 0.00562 = 0.2525% more tokens per extra pass
Agentic workspace100 criteria of 14,100 = 0.0070935.4 × 0.00709 = 0.25125% more tokens per 100 extra criteria
Engineering design1 pass of 244 = 0.0041024.4 × 0.00410 = 0.10010% more tokens per extra pass

Two real candidates from the paper's Table 6 show the rule doing its job. First, coding round 0, candidate A: "adds a bounded pre-completion verification audit and guidance for non-blocking polling of long-running jobs". It gained 3.93 points, which on 178 trials is exactly 7 extra passes (7 / 178 = 0.0393). That is well above δ = 0.017, so Equation 7 applies:

ceiling = β0 + β1ΔS = 0.10 + 44.5 × 0.0393 = 0.10 + 1.749 = 1.849 (up to +185% tokens would have been allowed)

The paper does not report R0-A's cost change, but it would have had to nearly triple the tokens to fail. It was accepted. A large, real gain buys a lot of room, which is the point: the rule is not against cost, it is against unpaid cost.

Second, engineering round 2: "adds a bounded recovery hint for the recurring 'workdir must be an existing directory' tool-use error". Passes went from 122 of 244 to 128 of 244, with 1.6% more tokens. Step by step:

Ŝ: 122 / 244 = 0.5000 → 128 / 244 = 0.5246    ΔS = 6 / 244 = 0.0246 > δ = 0.020

ceiling = 0.15 + 24.4 × 0.0246 = 0.15 + 0.60 = 0.75    ΔC = +0.016 ≤ 0.75 (accepted)

Six passes against a band of five: it cleared the noise by a single solved trial, and it paid almost nothing for it. The engineering instance also runs two domain guards after the cost check (the valid-output rate may not fall by more than 0.03, the no-submission rate may not rise by more than 0.02), and this candidate passed them too. It is the paper's picture of a good edit: small, task-agnostic, cheap, and attributable.

Inside the band: a different question

Equation 7 only applies when ΔS is above δ. Why not use it everywhere? Because inside the band a score change is not evidence of anything. If a within-band "gain" could pay for tokens at the Equation 7 rate, noise would be buying real cost. The paper's words: the purpose of the other branch "is to avoid treating a small score fluctuation as sufficient evidence by itself". So a candidate that has cleared the floor but not the band is judged by a different test, Equation 17:

ws
What a within-band score change is worth. Zero on the coding instance: there, a score change inside the noise counts for nothing at all.
wc
What a relative change in tokens is worth. The minus sign means a candidate that is cheaper (ΔC below zero) earns credit.
wn
What structural novelty is worth: a small bonus for trying a kind of mechanism the harness has never accepted.
νt(H′)
How many distinct structural component types the candidate touches that have never appeared in a winning edit (Equation 16).

Structural means one of four components, the ones that add machinery rather than change text or constants (Equation 15): client_tool, skill, memory and subagent. If Nt(ℓ) counts the accepted edit records tagged with component ℓ so far, then ν adds one for each structural type the candidate touches whose count is still zero. Prompt, control-flow, configuration, output-plumbing and context-management edits never earn it. This is the selection-side echo of Chapter 5's structured exploration: when a structural mechanism is a genuinely new idea, a tie inside the noise breaks in its favor.

The paper says the weights are fixed per instance but does not print them; the released configs do:

InstancewswcwnWhat that buys inside the band
Coding0150.5only a token saving or a new structure can get in; ν = 1 pays for at most 0.5 / 15 = 3.3% more tokens
Agentic workspace1414 (0.1 per criterion)150.5a within-band +0.3 points is worth 1414 × 0.003 = 4.24, enough for 4.24 / 15 = 28% more tokens
Engineering design244 (1 per pass)20.5one extra pass is worth 1; it pays for 1 / 2 = 50% more tokens

Now the case that shows why the band branch exists. Coding round 0, candidate B: "adds a similar verification reminder and long-running-work guidance, but with a smaller measured gain and additional inference cost". It gained 1.69 points with 26.1% more tokens. Do it by hand:

  1. How big is the gain? 1.69 points on 178 trials is exactly 3 extra passes: 3 / 178 = 0.0169.
  2. Is it above the band? δ = 0.017, and 0.0169 is not above it. So Equation 7 is not consulted.
  3. The band rule. wsΔS − wcΔC + wnν = 0 × 0.0169 − 15 × 0.261 + 0.5 × ν = −3.915 + 0.5ν.
  4. The verdict. With ν = 0 that is −3.915; even if it had touched a brand-new structural component (ν = 1), −3.415. Not above zero either way. Rejected by the cost rule.

Now look at how sharp that edge is. One more solved trial, 4 passes = 4 / 178 = 0.0225, and the candidate would have been above the band, judged by Equation 7 instead: 0.10 + 44.5 × 0.0225 = 0.10 + 1.00 = 1.10, so +26.1% tokens would have passed easily. The rule treats "not yet evidence" and "evidence" as different kinds of thing, and the boundary between them is exactly the calibrated noise. The paper's own summary of R0-A and R0-B: "the two candidates from the first coding round are superficially similar, yet only the candidate with a sufficiently large measured improvement survives the cost-aware selection rule."

Place candidates on the plane yourself. The shaded wedge is everything the selector would admit, assuming the incumbent is the best harness so far (so the floor sits at −δ).

The acceptance plane

Across: the candidate's score change. Up: its change in tokens per trial. Drag the teal point, or tap one of the paper's four real candidates. Switch instance to see how each domain prices tokens, and give the candidate a never-tried structural component.

+3.00
+40%

δ, β0 and β1 are Table 5; ws, wc, wn come from the released rrsi.json configs; the four candidates are Table 6 (R0-A's cost change is not reported, so it is drawn as a vertical line). The floor is drawn at −δ, which assumes the incumbent is the best harness so far; in a real run S★ can sit above the incumbent and the floor moves right by the difference. The engineering domain guards are assumed to hold.

Three things to see. On Coding, the band is a flat line at zero: inside the noise, the only way in is to be cheaper (or, with ν = 1, at most 3.3% more expensive). Step past δ and the ceiling jumps to 0.10 + 44.5 × 0.017 = 0.86, a cliff from 0% to +86% tokens. That cliff is R0-B's story.

Second, R8-B sits in the red, left of −δ, even though it is cheaper. The band rule would have admitted it (0 − 15 × (−0.136) = +2.04), but the floor is checked first and is non-compensatory, as Chapter 7 showed.

Third, switch to Workspace or Engineering. There ws is large, so the band line is steep: inside the noise, a small measured gain buys tokens at a higher rate than Equation 7 allows just outside it. On Engineering at ΔS = +1.9 points the band rule tolerates 244 × 0.019 / 2 = 2.32, that is +232% tokens, while just past the band at +2.1 points Equation 7 allows 0.15 + 24.4 × 0.021 = 0.66. The paper does not discuss this; it is a consequence of the released weights, and on Engineering the two validity guards still apply on top. It is a good reminder that these are hand-set hyperparameters, chosen, as Appendix D says, "using only the evolve environment and operational considerations".

Pruning: machinery must keep earning its place

The cost rule prices each edit once, at the moment it is proposed. But a harness is a stack of edits, and the world under an old edit keeps changing. Suppose round 4 accepted "cache retrieved clauses in memory" because it helped the harness of round 4. Two rounds later, a context-management edit started summarizing long documents into notes, which put the same material in front of the policy another way. The memory cache now costs tokens and helps nothing. (An illustrative history.) Score-only evolution has no reason to remove it: removing it would not raise the score, and the loop only ever asks what raises the score.

RRSI keeps a running verdict on every component from the edit history of Chapter 5. For each component ℓ it looks back over a pruning window of nprune rounds and takes the best measured gain of any edit to that component:

gt(ℓ)
The best recent evidence for component ℓ: the largest ΔS of any measured edit to it, accepted or not, inside the window.
t − ti ≤ nprune
The window: only edits measured in the last nprune rounds count. nprune = 4 for coding and the workspace, 5 for engineering (Table 5).
max ∅ = −∞
A component with no measured edit in the window gets the worst possible verdict. Silence is not evidence of usefulness.
Bt
The pruning targets handed to the proposer at round t, together with the accepted edits of those components.
Tt
The components with at least one measured edit so far. A component never tried is not in Tt, so it cannot be pruned: there is nothing to remove.

The paper's phrase for the test is "strictly positive measured gain": a component stays only if something done to it recently made the score go up. Zero is not enough. Let's work one through with nprune = 4 at round t = 10. The history holds memory edits measured in round 4 (+0.6 points, the accepted cache), round 7 (−0.2) and round 9 (−0.1).

  1. Which rounds are in the window? t − ti ≤ 4 means ti ≥ 10 − 4 = 6. The history at round 10 holds rounds up to 9. So rounds 6 to 9 count, and round 4 (10 − 4 = 6 rounds ago) does not.
  2. The best recent gain. g10(memory) = max(−0.2, −0.1) = −0.1.
  3. The verdict. −0.1 ≤ 0, so memory ∈ B10. The proposer is told: memory is unproductive, and the machinery to remove is "cache retrieved clauses in memory" from round 4.

Notice what the +0.6 of round 4 does not do: it does not protect the cache forever. That is the Lasso-like part of the analogy from Chapter 3. Lasso drives a weight to exactly zero unless the data keeps paying for it; RRSI deletes a component unless recent measurements keep paying for it. "A mechanism must continue to earn its place rather than persist simply because score-only evolution has no incentive to remove it."

The pruning window

Each row is a harness component; each dot is a measured edit to it, with its ΔS in points. Filled dots were accepted into the harness. Slide the current round and the window, and watch which components become deletion targets.

10
4

A toy history in the shape of the released history.jsonl: one record per measured edit, accepted or not. The two round-9 edits were bundled in one candidate, so both inherit its ΔS of −0.1 (Chapter 4). Equations 11 and 14 are computed exactly as in History.yield_g and History.prune_set.

Now watch control_flow. Its last measured edit was in round 8, at +0.2. Slide t to 12 and round 8 is still inside the window (12 − 8 = 4), so g = +0.2 and it is safe. Slide to 13 and round 8 falls out (13 − 8 = 5), g becomes −∞, and control_flow becomes a target, even though its last evidence was positive. That is the design, and it is worth being clear-eyed about: pruning treats a component that nobody has touched in nprune rounds as unproven. What protects a genuinely useful mechanism is that deletion is only a suggestion to the proposer.

That last point is the realization detail that makes pruning safe. Bt does not delete anything by itself. The proposer receives the targets and the accepted edits behind them (the released prune_set returns exactly that: component, recent best gain, accepted edits in the incumbent), and it drafts a candidate that removes them. That candidate is then evaluated and judged by the same gates as any other. Suppose, on the workspace instance, the deletion costs 0.05 points and saves 8% of the tokens (illustrative). It clears the floor, it is inside the band, and the band rule reads:

1414 × (−0.0005) − 15 × (−0.08) + 0 = −0.707 + 1.2 = +0.49 > 0 (the deletion is admitted)

If instead deleting the component cost a full point, the floor would refuse the deletion and the mechanism would stay. Pruning proposes; selection disposes. The two cost regularizers meet in the same gate.

One rule at the door, one in the house. The cost rule prices growth when it is proposed: extra tokens must come with a measured gain above the noise. Pruning re-prices growth after the fact: a component must keep showing a strictly positive gain, or the proposer is sent to remove it, and the removal itself must pass the gates. No prior method in the paper's comparison carries either constraint, and the result in Table 2 is the lightest evolved harness: 2.42 million tokens per trial, against 3.80 million without the regularizers.

Assembling Algorithm 2

We now have every gate a measured candidate meets, in the order the released code checks them: the floor (Chapter 7), then the cost rule with its two branches, then the domain guards. A candidate that passes all three is admissible; the round's winner is the admissible candidate with the highest measured score, or no one, in which case the incumbent stays. The same logic three ways. First by hand, for R0-B:

by hand# Coding, round 0, candidate B. At round 0 the best so far is H0 itself, so S* = S_t.
delta = 0.017;  dS = +0.0169;  dC = +0.261;  w_s, w_c, w_n = 0, 15, 0.5
floor:   dS >= -delta ?        0.0169 >= -0.017        pass
branch:  dS > delta ?          0.0169 > 0.017          no: inside the band, use Eq. 17
band:    w_s*dS - w_c*dC + w_n*nu = 0 - 3.915 + 0 = -3.915     not > 0
verdict: rejected, "cost rule failed"

Then from scratch, in a dozen lines of Python that contain the whole selection side:

pythondef judge(S_new, C_new, S_inc, C_inc, S_star, delta, cfg, nu=0, guards_ok=True):
    dS = S_new - S_inc                          # Eq. 6
    dC = (C_new - C_inc) / C_inc
    if S_new < S_star - delta:                  # Eq. 5, checked first, non-compensatory
        return False, "below the noise-adjusted floor"
    if dS > delta:                              # a real gain: tokens must be paid for
        ok = dC <= cfg.beta0 + cfg.beta1 * dS                  # Eq. 7
    else:                                       # inside the band: the score change is not evidence
        ok = cfg.w_s * dS - cfg.w_c * dC + cfg.w_n * nu > 0     # Eq. 17
    if not ok:
        return False, "cost rule failed"
    if not guards_ok:                            # engineering: valid-output and no-submission guards
        return False, "domain guard violated"
    return True, "admissible"

def select_round(cands, inc, S_star, delta, cfg):
    adm = [c for c in cands
           if judge(c.S, c.C, inc.S, inc.C, S_star, delta, cfg, c.nu, c.guards_ok)[0]]
    new = max(adm, key=lambda c: c.S) if adm else inc   # argmax, or keep H_t
    return new, max(S_star, new.S)                           # S* only goes up

And the released library call, which does the same with the candidate's novelty computed from the accepted-edit counts and the domain's guard function plugged in:

pythonfrom rrsi.selection import select_round
winner, decisions = select_round(cands, incumbent, S_star, delta, cfg,
                                 incumbent_counts, guard_fn=domain.guards)
# decisions[i].reason for R0-B reads:
# "cost rule failed: gain +0.0169 within delta 0.0170; shaped 0.0*dS - 15.0*dC
#  + 0.5*nu = -3.9150 <= 0 (nu=0)"

Every decision carries its reason as a string, and that string goes into the history record of every edit in the candidate. That closes the loop with Chapter 5: the next proposer reads not only that R0-B was rejected, but exactly which inequality it failed and by how much.

Coding round 0, candidate B gained +1.69 points (3 passes of 178) with 26.1% more tokens, on an instance with δ = 0.017 and ws = 0. Why was it rejected?

Chapter 9

Run the whole loop

Switch each regularizer on and off, and watch what the harness keeps and where the two scores go

We now own every part. Three habits on the proposal side: the annealed budget (A, Chapter 4), the ledger of what was tried and what it earned (B, Chapter 5), and structured exploration when progress stalls (C, Chapter 5). Four gates on the selection side: the critic (D, Chapter 6), the noise floor (E, Chapter 7), the cost rule (F, Chapter 8) and pruning (G, Chapter 8).

Each part answered one way the loop fools itself: bundles that hide what worked, a memory too short to stop re-testing failures, a proposer stuck on prompt rewrites, edits that name the test, gains made of noise, and growth nobody pays for. This chapter puts all seven into one loop, and then lets you pull them out one at a time.

First, the whole selection side as one rule. At round t the proposer hands over a small set of candidate harnesses, written 𝓗t. A candidate survives only if it passes every active check. The paper calls the checks non-compensatory: a big win on one cannot buy back a failure on another.

𝒜t
The admissible set: the candidates allowed to replace the incumbent this round. Only these compete on score.
𝓗t
This round's candidates, drawn by the regularized proposer: at most bt edits each, conditioned on the ledger, with a reserved slot when the search stalls. Switches A, B and C shape this set before any gate sees it.
critic(H′)
The leakage screen, run on the diff before any evaluation is spent (switch D). In the paper it sits at the end of Algorithm 1; we fold it into the same line so the whole filter reads at once.
Ŝ(H′)
The candidate's measured evolve-split score: k trials on every evolve task, averaged (Chapter 1).
S★
The best evolve-split score recorded so far in the run (switch E measures the floor from it).
δ
The noise band, measured once on the unchanged base harness: 0.004 on the workspace instance, which is 0.4 points.
c(H′)
The cost condition (switch F). If the gain ΔS clears the band: ΔC ≤ β0 + β1ΔS. If it does not: wsΔS − wcΔC + wnν > 0.
g(H′)
A domain guard. Only engineering design uses one (valid-output rate may not fall by more than 0.03, no-submission rate may not rise by more than 0.02); elsewhere it is always 1.
Ht+1
The next incumbent: the admissible candidate with the highest measured score. Pruning (switch G) acts through this same rule: a deletion is just a candidate whose diff removes machinery.

Two quick checks with the workspace constants (δ = 0.004, β0 = 0.10, β1 = 35.4, and from the released configuration ws = 1414, wc = 15, wn = 0.5). A candidate measures +0.21 points with 3% more tokens. The gain is inside the band, so the within-band rule decides:

1414 × 0.0021 − 15 × 0.03 + 0.5 × 0 = 2.97 − 0.45 = +2.52 > 0 (admissible, if it also clears the floor)

A bundle measures +0.62 points and asks for 38% more tokens. That gain clears the band, so the cost rule sets a ceiling:

0.10 + 35.4 × 0.0062 = 0.10 + 0.22 = 0.32 (a 32% ceiling, and 38% is over it: rejected)

When you switch a part off in the device below, its line simply disappears from the rule, with one exception. Switch off the cost rule and the loop does not become lawless: it falls back to the plain rule of Chapter 1 (Equation 2), where a candidate must score higher than the incumbent to replace it. That is what unregularized evolution does.

The toy behind the device

To let you break the loop we need a world where we know the truth about every edit, which the paper cannot give us. So the device runs a toy. Its proposer draws from a pool of 113 edit ideas for a legal-work agent, spread over the paper's nine components. Each idea has a hidden kind, and each kind has hidden effects:

KindEvolve splitOut of distributionTokensAn example from the pool
Real mechanism+0.09 to +0.19+0.28 to +0.54+0 to +4.5%One bounded verification pass before stopping
Leak+0.25 to +0.45−0.05 to −0.15+0 to +2%If a task names Delaware: cite the merger statute
Noise00 to −0.03+0 to +1%Reword the system prompt as a senior partner
Bloat+0.05 to +0.120 to −0.10+10 to +22%A second full review pass on every deliverable
Trim−0.05 to −0.121.2 × that−6 to −12%Skip the final formatting check
Harmful−0.20 to −0.401.5 × that−3 to +3%Stop as soon as the first file is written

Three more rules make the toy behave like the failure modes of Chapter 2. Every kept edit costs 1.8% more tokens, because the policy reads the harness text on every step. A bundle of s edits is tailored to that round's feedback, so each edit in it carries a little extra evolve-split score and a little loss out of distribution that grows with s: this is the paper's "high effective capacity" in toy form. And the proposer reaches for prompt edits first, and more so the more it has already proposed them, which is the collapse that exploration exists to break. Prompt and configuration ideas are mostly leaks and noise; the real mechanisms live mostly in the other seven components.

Measurement is the only noisy part. A candidate's measured score is its true evolve-split score plus Gaussian noise of 0.14 points, so two measurements of the same harness differ by about 0.2 points, and twice that is the workspace band, 0.4 points. The incumbent's recorded score is never measured again, so it carries whatever luck it was accepted with (the dotted line).

Break the loop

Ready

Pick a preset or flip the seven switches yourself, then press Run 20 rounds. Each round shows the candidates, the stamp each gate gave them, and what enters the harness file. The charts show the kept harness's true scores and its token cost.

Proposal side
Selection side

This round

The harness file

    Press Run 30 seeds to average the four presets over thirty worlds and set them beside the paper's four arms.

    A toy, not the paper's system. The 113 edit ideas, their hidden effects and the proposer's habits are invented to show each mechanism at work. The gate constants are the paper's agentic-workspace values (δ = 0.004, β0 = 0.10, β1 = 35.4, nprune = 4, w = 3, mdraft = 1, a budget from 3 down to 2, T = 20, from Table 5; ws = 1414, wc = 15, wn = 0.5 and two candidates per round, from the released rrsi.json). Simplifications: the critic catches each leak with probability 0.85 and strips it, and a pruning deletion runs as its own small candidate beside the two. Only the four presets correspond to real arms (Table 2), and only their order is meant to match.

    What to try, in order

    The numbers below are the toy's averages over seeds 1 to 30, so a single run will wander around them. Keep Reveal on for the first few.

    1. No rules. Press Run. The file fills fast: about 57 edits after twenty rounds. Reveal them: roughly 17 leaks, 16 noise edits, 6 trims, 5 harmful edits, 3 bloats, and only 10 real mechanisms. The evolve split climbs to 94.3 while out of distribution falls to 36.2, below the 39.7 it started from, at 6.5 million tokens a trial. The loop learned the test and got worse at the job.
    2. Switch on D alone. Watch the struck red lines in the log: in an average run the critic strips about 41 leaking edits. Kept leaks drop from 17 to 5, out of distribution recovers to 39.5, and the evolve split falls to 91.8. Most of the old evolve-split gain was leakage.
    3. Add A. Bundles shrink from up to six edits to three, then two. The harness keeps about 20 edits instead of 32, tokens fall to 2.8 million and out of distribution rises to 40.9: small bundles carry less tailoring and fewer riders.
    4. Add B. The ledger stops the proposer from re-drawing ideas that already failed and steers it toward components that paid off. Kept mechanisms rise from about 6 to about 9, and out of distribution reaches 42.2. Tokens creep back up to 3.4 million, because nothing yet charges for bloat.
    5. Press RRSI to add exploration, the floor, the cost rule and pruning. Watch the new stamps: the cost rule refusing bloat, the band rule deciding the close calls, a reserved exploring slot when the score stalls, a deletion candidate when a component goes quiet. Tokens fall to 2.5 million, the evolve split settles at 90.8, out of distribution at 41.4.
    6. Notice what step 5 cost. Out of distribution slipped from 42.2 to 41.4 while tokens fell by about a quarter. The acceptance gates are a price, and the paper's constants set it: with wc = 15 against ws = 1414, one percent of tokens is worth about 0.01 points of evolve-split score, so the band rule admits a few cheap trims and lucky in-band edits. The paper names that trade itself: evolution "does buy part of its gain with test-time compute; the budget decides how much."
    7. Try the two half-arms, Proposal only and Selection only, then press Run 30 seeds and set the toy's four averages beside the paper's Table 2.
    8. Switch off E inside RRSI. The averages do not move, but the stamps do: every "below the floor" becomes "band rule". At the workspace constants the within-band rule refuses almost every large drop first. The floor earns its place where that rule is lenient: when cheap trims arrive one after another, and in the coding instance, where ws = 0 (Table 6's R8-B, Chapter 7).

    The real ablation

    The paper does not switch parts off one at a time. It removes the two groups, on the agentic-workspace instance, with everything else held fixed (the same base harness, policy, evolve split, round count and candidate budget):

    VariantHarvey LAB, evolveHarvey LAB, ID held-outOut of distribution, averageTokens per trial, millions
    H0 (no evolution)89.486.939.71.56
    Unregularized evolution92.888.940.33.80
    w/o proposal regularizers90.788.841.92.69
    w/o acceptance regularizers91.588.741.03.59
    RRSI90.589.243.62.42

    Table 2 of the paper. Out of distribution is the mean of JobBench, GDPval and APEX-Agents.

    Read it the way the authors do: "removing either group raises the evolve-set score and lowers transfer." Without the acceptance gates, the evolve split rises from 90.5 to 91.5 while out of distribution falls from 43.6 to 41.0, and tokens rise by half (3.59 against 2.42 million, 1.48 times). The paper's diagnosis: an unconstrained selection rule "spends most of its accepted edits on noise and on context rather than on mechanism."

    Without the proposal constraints, out of distribution falls 1.7 points, to 41.9. The paper says this arm "costs only 0.2 points on the evolve split", but look at the table: 90.7 against 90.5 is 0.2 points above RRSI. Read "costs" as "differs by". The conclusion stands: "steering where the search looks matters even when nothing is rejected."

    Remove both and you get the highest evolve split of any arm, 92.8, and an out-of-distribution average of 40.3, within a point of the harness it started from, at 3.80 million tokens a trial against RRSI's 2.42. Notice also the middle column: every arm lands between 88.7 and 89.2 on Harvey LAB's in-distribution held-out split. The separation only appears out of distribution, which is exactly where an evolve split cannot see.

    Proposal rules decide what the loop looks at; selection rules decide which looks become permanent. Remove the first group and the gates have little worth choosing. Remove the second and every lucky, leaking or bloated look sticks. Table 2 shows both failures, and the toy reproduces their order on all three numbers.

    The same loop in the released code

    The toy compresses a round into a few lines. The released implementation runs the same steps with real files, which is what makes each regularizer auditable:

    1. Analyze. An analyst reads the incumbent's own evaluation (the worst trials of the lowest-scoring tasks, plus a few of the best) and writes the feedback Ft.
    2. Budget. edit_budget(t, T, b_min, b_max) sets bt (switch A).
    3. Read the ledger. From history.jsonl, one record per edit: the stall flag σt, the tried and untried components, the exploration directive and the prune set Bt (switches B, C, G).
    4. Draft and screen. For each of the two variants, in its own git worktree: propose tagged edits, run the critic with bounded repair, then a liveness smoke test that compiles, constructs and runs a few tasks (switch D).
    5. Evaluate. The screened candidates run on the evolve set with k trials; a missing trial scores zero against the full denominator.
    6. Judge. judge() applies the floor, then the cost or band rule, then any domain guard; the admissible argmax wins; S★ updates; every edit gets its history record (switches E, F).
    7. Commit. The branch evolve/<domain> fast-forwards to the winner, so the incumbent is always a commit you can diff.
    In the paper's ablation, removing the acceptance regularizers raised the evolve-split score from 90.5 to 91.5, lowered the out-of-distribution average from 43.6 to 41.0, and used half again as many tokens. What does that pattern tell you?

    Chapter 10

    What it buys

    Eight benchmarks in three domains: every held-out split improves, on fewer tokens than any evolved rival

    The toy in Chapter 9 can be tuned to tell any story. The question the paper was written to answer is narrower and harder. When a real harness wraps a real frontier model, and the search runs under all seven rules, does the harness it keeps carry its gains to tasks it was never scored on?

    To answer that, the authors run the same experiment three times, in three kinds of work that differ in their tasks, their tools and, most of all, their verifiers: the programs that decide whether a deliverable is good. In each domain the harness evolves against one suite. Then it is frozen and run, unchanged, on benchmarks the search never saw.

    Coding

    graded by the task's own hidden unit tests

    Scored on: Terminal-Bench 2.1, 89 container tasks driven through a real shell.

    Never scored: SWE-bench Verified, real GitHub issues fixed with repository-level patches.

    Agentic workspace

    graded by LLM judges: per-criterion rubrics, or a side-by-side against a human expert

    Scored on: Harvey LAB, 120 legal tasks across 25 practice areas.

    Never scored: Harvey LAB's 40 held-out tasks, JobBench, GDPval and APEX-Agents.

    Engineering design

    graded by frozen simulators and testbenches, no judge model

    Scored on: EngDesign, 61 design tasks with physical constraints.

    Never scored: Frontier-Eng, optimization problems from 26 domains, scored as a Medal Score.

    Everything else is held still. The policy inside every harness is Claude Opus 4.8, frozen for the whole run. The proposer, the analyst that writes each round's failure feedback, and the leakage critic are Claude Opus 4.8 as well. The starting harnesses are ordinary ones: Terminus-2 for coding, and for Harvey LAB and EngDesign a ReAct loop over an MCP tool gateway with a dynamic toolbelt and ReSum-style context management.

    One more control matters more than it looks. Every score is measured against H0, the unevolved starting harness, in the same window: the same tool environment, the same judge, the same number of trials. The paper's reason is blunt: this way "no gain can be attributed to drift in the evaluation infrastructure." A judge model updated between two measurements, or a flaky container image, would otherwise show up as an improvement.

    The benchmarks are also counted in a way that cannot be gamed by quitting. APEX-Agents is scored over all 480 tasks, and a rollout lost to an infrastructure failure counts as a failed task, which "prevents a harness that crashes on hard worlds from looking better than one that attempts them." Frontier-Eng drops its EngDesign domain, because those tasks overlap the evolve suite, and tasks whose environment could not be built earn no credit in either arm, so both arms are scored on exactly the same 38 of its 47 tasks. It is the same principle as the released evaluator from Chapter 1: a missing trial scores zero with the full denominator.

    Figure 3, split by split

    One row per benchmark. Yellow rows are the suites the search was scored on; green rows are splits it never saw. Tap a row to read it, switch between the gain and the raw scores, and filter by domain.

    Every number is from Figure 3 of the paper; the relative gains are computed here as the gain divided by the H0 score. The units differ by benchmark (pass rate, fraction of rubric criteria, win rate against a human expert, Medal Score), which is why the paper, and this device, compare each split only with its own H0.

    Read the colors before the lengths. Every one of the nine bars sits to the right of zero: no split regressed, in any domain. The paper calls that out as the one result a memorizing harness cannot produce: "No held-out split regresses anywhere, which is the failure a memorizing harness produces."

    Now look at which bars are longest. The three yellow bars, the suites the search could see, are not the story. Harvey LAB's evolve split gained just 1.1 points, the smallest gain in the figure. The long bars are green: JobBench, APEX-Agents, Frontier-Eng. The rules traded a little score on the practice set for more score everywhere else, which is exactly the trade Chapters 4 to 8 were built to make.

    Because the units differ, the fair way to compare splits is the relative gain: how much the score grew as a fraction of where it started.

    relative gain
    The improvement as a fraction of the starting score. It lets a 4.3-point gain on a benchmark that starts at 17.7 be compared with a 1.8-point gain on one that starts at 82.0.
    SRRSI
    The score of the harness RRSI kept after its final round, run unchanged on this split.
    SH0
    The score of the unevolved starting harness on the same split, measured in the same window.

    Two worked examples. JobBench starts at 36.0 and ends at 40.7:

    (40.7 − 36.0) / 36.0 = 4.7 / 36.0 = 0.1306 = 13.1% (JobBench, never scored)

    Frontier-Eng starts much lower, at 17.7 Medal points, and ends at 22.0:

    (22.0 − 17.7) / 17.7 = 4.3 / 17.7 = 0.2429 = 24.3% (Frontier-Eng, never scored)

    The paper's headline "up to 22.9%" is a different comparison. It measures RRSI against the average of the four prior methods on the same held-out split, not against H0. On Frontier-Eng that average is 17.9 (Figure 1d):

    (22.0 − 17.9) / 17.9 = 4.1 / 17.9 = 0.2291 = 22.9% (RRSI over the average prior method)

    And the two averages the project page quotes fall straight out of the figure:

    (6.0 + 1.1 + 4.9) / 3 = 12.0 / 3 = 4.0 (average gain on the three evolve suites)

    (1.8 + 2.3 + 4.7 + 3.5 + 3.7 + 4.3) / 6 = 20.3 / 6 = 3.4 (average gain on the six held-out splits)

    Is it just pleasing the judges?

    A skeptic has a good objection ready. Harvey LAB, JobBench and GDPval are all scored by judge models. A harness could raise those scores without doing better work, simply by learning to write the way judges like: longer, more hedged, more confidently formatted. The same style would then "transfer" to every other judged benchmark, and look like generalization.

    The engineering domain closes that route. Every EngDesign and Frontier-Eng task is graded by its own frozen simulator or testbench. The grading is deterministic: a design meets the stated constraints or it does not, and there is no judge to charm. The gains survive there: +4.9 on EngDesign and +4.3 Medal points on Frontier-Eng. Deterministic grading also removes judge variance from the measurement, so on EngDesign "every point of variance we measure comes from the policy."

    GDPval is worth a second look too, because its number means something concrete. For each of 185 tasks, the harness's deliverable is placed beside the deliverable of the human expert that ships with the benchmark. Three judges from different vendors (Qwen3.6-35B-A3B run locally, Claude Sonnet 4.6 and Gemini-3.1 Pro) each pick the better one, in both presentation orders, and the majority decides. The score is a win rate against the expert. H0 sat at 48.8, below half; RRSI's harness reaches 52.3, so its deliverable is now preferred over the expert's more often than not.

    Against the prior methods

    A gain over H0 could still be ordinary. Every harness-evolution method gains something. So the authors run four recent ones, Meta-Harness, AHE, TTHE and HarnessX, from the same H0, on the same Harvey LAB evolve split, with the same frozen policy and the same candidate budget (Table 1). The only thing that differs is the method.

    On the split they were scored on, all four look fine, and every one of them beats RRSI. On Harvey LAB's in-distribution held-out split, everyone lands within a point of everyone else. The separation appears only out of distribution, "and there the ranking inverts." The second device lets you see both halves of that sentence: the ranking flip, and what each method pays for its score in tokens.

    Transfer against cost

    Pick a method, or tap its point. The first view is Figure 4a: policy tokens per trial against the out-of-distribution average. The second joins each method's evolve-split score to its out-of-distribution score, so you can watch the ranking turn over.

    Scores are Table 1; steps per trial are Figure 4b. Token costs for H0 (1.56 M), RRSI (2.42 M) and AHE (3.82 M) are stated in the text; the paper gives no numbers for Meta-Harness, HarnessX and TTHE, so their positions (about 2.7, 2.95 and 2.45 M) are read off Figure 4a and drawn with a dashed ring. The out-of-distribution average is the mean of JobBench, GDPval and APEX-Agents, computed here from Table 1.

    Start with AHE, the extreme case. It spends 3.82 million policy tokens per trial and runs 34.6 steps, the most of any arm, and it ends at 39.2 out of distribution, below the harness it started from. Against RRSI's 2.42 million tokens:

    (3.82 − 2.42) / 2.42 = 1.40 / 2.42 = 0.579 ≈ 58% more tokens for 43.6 − 39.2 = 4.4 points less out of distribution

    Now the ablation arm from Chapter 9, the same search with every rule removed. It spent 3.80 million tokens per trial (Table 2). RRSI's harness runs on:

    (3.80 − 2.42) / 3.80 = 1.38 / 3.80 = 0.363 ≈ 36% fewer tokens (the abstract rounds this down to "30% fewer")

    Why is RRSI the lightest evolved harness? Two of its rules act directly on cost. In the paper's words, one "refuses growth that is not paid for when it is proposed, and the pruning rule removes growth that has stopped being paid for since." None of the four prior methods carries either constraint, and every one of them lands in the shaded region of Figure 4a: more tokens per trial for a lower out-of-distribution average. The ordering carries over to trajectory length, 26.3 steps per trial for RRSI against 27.3 to 34.6 for the others.

    One honest caveat sits in the same figure. No evolved harness is as cheap as H0, at 1.56 million tokens and 21.2 steps. RRSI's harness costs 2.42 / 1.56 = 1.55 times as much per trial. As the paper puts it, "evolution does buy part of its gain with test-time compute; the budget decides how much."

    A different model, and a weaker one

    A harness evolved around one model might only have learned that model's quirks. The authors test that twice on the coding domain (Tables 3 and 4).

    PolicyBenchmarkH0 → RRSIΔ
    Claude Opus 4.8Terminal-Bench 2.174.2 → 80.2+6.0
    Claude Opus 4.8SWE-bench Verified82.0 → 83.8+1.8
    Gemini 3.5 FlashTerminal-Bench 2.164.6 → 78.7+14.1
    Gemini 3.5 FlashSWE-bench Verified76.8 → 79.0+2.2
    Gemini 3.1 Flash Lite, never in the searchTerminal-Bench 2.1, harness evolved with 3.5 Flash11.2 → 14.6+3.4

    The first four rows are two independent evolution runs, one per model family. Both follow the same pattern: a clear gain on the suite the search could see, and a smaller but real gain on SWE-bench Verified, which it never scored. The Gemini run is where the abstract's "up to 14.1 points" comes from; the main Claude run gains 6.0. The stronger policy "starts closer to the ceiling of both suites and leaves less room to gain."

    The last row is the sharper test. The harness evolved with Gemini 3.5 Flash is run, unchanged, with Gemini 3.1 Flash Lite, a smaller model that never took part in the search. "A harness is a program, not a set of weights," the authors write, so a mechanism that only helps the policy it was searched against is an artifact of that policy. Flash Lite starts at 11.2, less than a fifth of the search policy's 64.6 (11.2 / 64.6 = 0.17), and still gains:

    (14.6 − 11.2) / 11.2 = 3.4 / 11.2 = 0.304 = 30.4% (relative gain for a backbone the search never used)

    The absolute gain is smaller because a weaker backbone leaves fewer tasks within reach of any harness. The relative gain says the mechanisms did not depend on the capability level they were found at.

    What is verified, and what is not. Verified in the paper: all six held-out splits improve over H0, measured in the same window; the gains hold where no judge is involved; RRSI's out-of-distribution average (43.6) beats every prior method on the same budget; the harness helps under a second model family and under a smaller model that never saw the search. Not reported: any spread across repeated evolution runs. Each arm appears as one final harness, so we cannot tell how much a rerun of the search with a different seed would move these numbers. And the out-of-distribution gains, while consistent, are a few points each.
    Which result best rules out the explanation that RRSI's harness simply learned to write the way LLM judges like?

    Chapter 11

    Connections

    Lock in the cheat sheet, run one real round through the code, and place RRSI among self-improving harnesses

    You can now read the RRSI paper and defend every design decision in it: why a search that reuses one finite evolve set overfits it, why the best of two noisy scores lies upward, why a candidate may bundle only a few edits and fewer as the run goes on, why the proposer keeps a ledger of what failed, why the critic reads the diff before anything is scored, why the floor hangs from the best score ever seen, why tokens have a price, and why machinery that stops earning its place is put up for deletion. Let's lock it in.

    The one-paragraph summary

    RRSI treats automated harness evolution, a practical form of recursive self-improvement at the agent-system level, as adaptive empirical optimization over one finite and noisy evolve set, and regularizes the search rather than the harness. The edit space stays open: prompts, control flow, configuration, context management, tools, skills, memory and subagents may all change. On the proposal side, a cosine-annealed budget caps how many independently attributable edits one candidate may bundle, every measured edit enters a per-edit ledger the proposer must read, and a stalled run reserves a slot for components it has never tried. On the selection side, a critic rejects benchmark-specific or inert edits before evaluation, a candidate must score at least the best-ever score minus a calibrated noise band, extra tokens must be paid for by measured gain (inside the band, a shaped score of gain, token savings and structural novelty must come out positive), and components with no recent positive gain are handed back as deletion targets. With a frozen Claude Opus 4.8 inside, the harness it keeps improves all six held-out splits across coding, legal workspace and engineering design, reaches 43.6 out of distribution where unregularized evolution reaches 40.3, and runs on 2.42 million tokens per trial instead of 3.80.

    The numbers that matter

    QuantityValue, and why it matters
    Gain on the suites the search saw+6.0, +1.1, +4.9 (average +4.0)Terminal-Bench 2.1, Harvey LAB, EngDesign (Figure 3)
    Gain on the six held-out splits+1.8 to +4.7 (average +3.4), none downThe failure a memorizing harness produces never appears
    Out-of-distribution average, workspace43.6 vs 39.7 (H0), 40.3 (no rules), 40.6 (best prior)Table 1 and Table 2: the only arm more than a point above H0
    Evolve-split score, workspace90.5, the lowest of the evolved harnessesThe trade the regularizers are designed to make
    Policy tokens per trial2.42 M vs 3.80 M (no rules), 3.82 M (AHE), 1.56 M (H0)36% fewer than unregularized; AHE spends 58% more
    Steps per trial26.3 vs 27.3 to 34.6 (prior), 21.2 (H0)Figure 4b: the lightest evolved harness, still heavier than H0
    Frontier-Eng Medal Score17.7 → 22.0 (+24.3%)22.9% above the prior average of 17.9; no judge involved
    Other backbonesGemini 3.5 Flash +14.1; unseen Flash Lite +3.4 (30.4%)A harness is a program, not a set of weights
    Noise band δ0.017 / 0.004 / 0.0203 of 178 passes, 60 of about 14,100 criteria, 5 of 244 passes
    Edit budget4 → 1, 3 → 1, 4 → 1, cosine with a ceilingThe last round actually allows 2 (Chapter 4)
    Cost ruleβ0 0.10 / 0.10 / 0.15; β1 44.5 / 35.4 / 24.4+25% tokens per pass, per 100 criteria; +10% per pass
    Windows and countsT 20 / 20 / 40, k 2 / 2 / 4, w 3, mdraft 1, nprune 4 / 4 / 5Table 5, fixed without looking at held-out data

    The triples are coding / agentic workspace / engineering design throughout.

    One rule for admission

    Chapters 6 to 8 built the selector gate by gate. Here is all of it in one line. A candidate that survived the critic and was evaluated enters the admissible set only if three independent tests pass, and none can buy back a failure of another:

    H′ ∈ 𝒜t
    The candidate is admissible this round. The winner is the admissible candidate with the highest measured score; if none is admissible, the incumbent stays.
    Ŝ(H′) ≥ S★ − δ
    The floor (Chapter 7): never more than one noise band below the best evolve score seen so far.
    c(H′)
    The cost condition (Chapter 8). If the gain clears the band, the token increase must stay under β0 + β1ΔS. If it does not, the shaped rule wsΔS − wcΔC + wnν must come out positive.
    g(Ht, H′)
    A domain guard. Coding and workspace use none; engineering rejects a candidate whose valid-output rate falls by more than 0.03 or whose no-submission rate rises by more than 0.02.

    The critic sits in front of all of this, and the proposer's three habits sit in front of the critic. So seven regularizers stand between an idea and permanent harness state. Three shape what gets proposed: a budget on how much one candidate may bundle, a ledger that remembers every hypothesis, and an exploration rule that decides where search capacity goes. Four decide what stays: a critic, a floor, a price on tokens, and later, a pruner that can take it back out.

    Here is that rule as a path a single candidate walks. The paper's Table 6 gives four real decisions from the released runs; each one stops at a different place, or does not stop at all.

    Trace one candidate through the selector

    Pick one of the paper's Table 6 decisions, or move a slider to build your own candidate. It walks down Algorithm 2 and stops at the first gate it fails.

    +1.69
    +26.1%

    Decisions and outcomes are Table 6; δ, β0 and β1 are Table 5; ws, wc and wn come from the released configs, since the paper's Table 5 omits them. The floor assumes the incumbent holds the best score so far; for R8-B that does not matter, because a drop of 2.81 points fails for any best score at or above the incumbent. The paper does not report R0-A's token change, and it does not report the engineering guard's rates for R2, so the guard is assumed to hold as it did.

    The equation cheat sheet

    EquationWhat it says
    S(H; D) = 𝔼[r(x, τ)],  C(H; D) = 𝔼[c(τ)]  (1)A harness's true score and true policy-token cost on a task setChapter 1
    Ŝ, Ĉ = averages over k trials of every evolve task  (3)What the search can actually measure: finite and noisyChapter 1
    Ht+1 = argmax Ŝ(H′) over the candidates and Ht  (2)The unregularized loop: keep whatever measures highestChapters 1 and 2
    ℋt ~ Preg(· | Ht, Ft, Lt, bt, Et, Bt)  (8)The regularized round: the same loop, conditioned on budget, ledger, exploration and prune targetsChapter 3
    bt = ⌈bmin + (bmax − bmin) · ½(1 + cos(πt/T))⌉  (4)How many edits a candidate may bundle, shrinking over the runChapter 4
    ‖zt‖0 ≤ bt  (9)The L0-style cap on the number of active atomic editsChapter 4
    Lt = {(ti, ℓi, hi, di, ΔSi, ΔCi, ai)}  (10)One ledger record per edit; bundled edits share one measurementChapter 5
    σt = 𝟙[Ŝt − Ŝt−w ≤ δ],  Ut = K \ Tt  (13)Stalled means no progress beyond the band in w rounds; then a slot goes to untried componentsChapter 5
    Ŝ(H′) ≥ S★ − δ  (5)The noise-adjusted floor under the best score so farChapter 7
    ΔS = Ŝ(H′) − Ŝ(Ht),  ΔC = relative token change  (6)Gain in score units; cost as a fraction of the incumbent'sChapter 8
    ΔC ≤ β0 + β1ΔS, when ΔS > δ  (7)The Ridge-like price on tokens: more cost only for more measured gainChapter 8
    wsΔS − wcΔC + wnν > 0, when ΔS ≤ δ  (17)Inside the band, a score wobble alone is not evidenceChapter 8
    νt(H′) = number of structural types never in a winning edit  (16)Credit for trying a new tool, skill, memory or subagentChapter 8
    gt(ℓ) = best recent ΔS;  Bt = {ℓ ∈ Tt : gt(ℓ) ≤ 0}  (11, 14)The Lasso-like eviction list: no recent positive gain, no placeChapter 8

    The core, as code

    Here is the whole method in about sixty lines of Python. It follows the released code's semantics: the budget uses a ceiling with a rounding guard, the stall flag stays off until w rounds exist, the prune set counts silence as no evidence, the judge checks the floor before the cost rule, and the ledger gets one record per edit. The domain guard and the critic's repair rounds are left out to keep it short.

    pythonimport math
    
    K     = ["prompt", "control_flow", "config", "output_plumbing", "context_mgmt",
             "client_tool", "skill", "memory", "subagent"]              # Eq. 12
    K_STR = {"client_tool", "skill", "memory", "subagent"}              # Eq. 15
    
    def edit_budget(t, T, b_min, b_max):                                # Eq. 4, t = 0 .. T-1
        v = b_min + (b_max - b_min) * 0.5 * (1 + math.cos(math.pi * t / T))
        return math.ceil(round(v, 9))       # round() so 1.0000000002 never becomes 2
    
    def stall_flag(scores, t, w, delta):                                # Eq. 13
        if t < w:
            return 0                        # fewer than w rounds: never stalled
        return int(scores[t] - scores[t - w] <= delta)
    
    def prune_set(ledger, t, n_prune):                                  # Eq. 11 and 14
        measured = [r for r in ledger if r["delta_S"] is not None]
        g = {r["component"]: -math.inf for r in measured}              # T_t: every tried component
        for r in measured:
            if t - r["t"] <= n_prune:                                   # recent evidence only
                g[r["component"]] = max(g[r["component"]], r["delta_S"])
        return sorted(c for c, best in g.items() if best <= 0)       # silence counts as no gain
    
    def judge(S_new, C_new, comps, inc, S_star, delta, cfg, won_before):  # Algorithm 2, lines 3 to 13
        if S_new < S_star - delta:                                      # Eq. 5: the floor comes first
            return False, "below the noise-adjusted floor"
        dS = S_new - inc["S"]
        dC = (C_new - inc["C"]) / inc["C"]                              # Eq. 6: relative token change
        if dS > delta:                                                  # the gain clears the noise band
            ok = dC <= cfg["beta0"] + cfg["beta1"] * dS                # Eq. 7
        else:                                                           # inside the band
            nu = sum(1 for c in set(comps)
                     if c in K_STR and won_before.get(c, 0) == 0)          # Eq. 16
            ok = cfg["w_s"] * dS - cfg["w_c"] * dC + cfg["w_n"] * nu > 0  # Eq. 17
        return ok, "admissible" if ok else "cost rule failed"        # (engineering adds two guards here)
    
    def rrsi_round(t, state, cfg, analyst, proposer, critic, evaluate):
        inc, ledger = state["incumbent"], state["ledger"]
        F = analyst(inc)                                                # Algorithm 1, line 1: failure summary of H_t
        b = edit_budget(t, cfg["T"], cfg["b_min"], cfg["b_max"])         # L0: edits per candidate
        tried = {r["component"] for r in ledger if r["delta_S"] is not None}
        explore = dict(stalled=stall_flag(state["scores"], t, cfg["w"], cfg["delta"]),
                       untried=[c for c in K if c not in tried], slots=cfg["m_draft"])
        prune = prune_set(ledger, t, cfg["n_prune"])                     # L1: deletion targets
        drafts = proposer(inc, F, ledger, b, explore, prune)            # m = 2 drafts in the released configs
        screened = [h for h in drafts if critic(h.diff)["verdict"] == "accept"]   # BEFORE any score
        for h in drafts:
            if h not in screened:                                     # logged, never measured, never counted
                ledger += [dict(t=t, component=e.component, hypothesis=e.hypothesis,
                                delta_S=None, delta_C=None, accepted=False) for e in h.edits]
        won_before = {}
        for r in ledger:
            if r["accepted"]:
                won_before[r["component"]] = won_before.get(r["component"], 0) + 1
        best, measured = None, []
        for h in screened:
            S, C = evaluate(h, k=cfg["k"])                               # the same evolve set, k trials a task
            ok, why = judge(S, C, h.components, inc, state["S_star"], cfg["delta"], cfg, won_before)
            measured.append((h, S, C))
            if ok and (best is None or S > best[1]):
                best = (h, S, C)                                        # argmax over the admissible set
        for h, S, C in measured:                                        # Eq. 10: one record per EDIT
            ledger += [dict(t=t, component=e.component, hypothesis=e.hypothesis, diff=h.diff,
                            delta_S=S - inc["S"], delta_C=(C - inc["C"]) / inc["C"],
                            accepted=best is not None and h is best[0]) for e in h.edits]
        if best:                                                        # accept: fast-forward the evolve branch
            state["incumbent"] = dict(harness=best[0], S=best[1], C=best[2])
        state["S_star"] = max(state["S_star"], state["incumbent"]["S"])     # running best
        state["scores"].append(state["incumbent"]["S"])                # the trajectory stall_flag reads
        return state
    

    One real round, run through the code by hand

    The paper's Table 6 opens with round 0 of the coding run. Two candidates were drafted, labeled A and B (the released code names a round's variants with letters). Walk them through the functions above with the coding values of Table 5 and the released config (δ = 0.017, β0 = 0.10, β1 = 44.5, ws = 0, wc = 15, wn = 0.5).

    StepWhat the code computes, and the result
    edit_budget(0, 20, 1, 4)⌈1 + 3 × ½ × (1 + cos 0)⌉ = ⌈1 + 3 × ½ × 2⌉ = ⌈4⌉b0 = 4 edits per candidate
    stall_flag(scores, 0, 3, 0.017)t = 0 < w = 30: not stalled, no reserved slot
    tried, prune_setthe ledger is emptyall 9 components untried; nothing to prune
    critic(A), critic(B)A: a bounded pre-completion verification audit plus guidance for non-blocking polling of long-running jobs. B: a similar verification reminder and long-running-work guidance. Neither names a task.both accepted for evaluation
    judge(A)ΔS = +3.93 points = 0.0393 > δ = 0.017, so Eq. 7: ceiling = 0.10 + 44.5 × 0.0393 = 0.10 + 1.749 = 1.849admissible (the paper does not report A's cost change; it cleared a ceiling of +184.9%)
    judge(B)ΔS = +1.69 points = 0.0169 ≤ 0.017, so Eq. 17 with ΔC = +0.261 and text-only edits (ν = 0): 0 × 0.0169 − 15 × 0.261 + 0.5 × 0 = −3.915rejected by the cost rule (even ν = 1 gives −3.415)
    argmax, S★, ledgeronly A is admissibleA becomes H1; S★ rises to A's score; A's edits get a = 1, B's edits a = 0

    That is the paper's own summary of the row pair: "the two candidates from the first coding round are superficially similar, yet only the candidate with a sufficiently large measured improvement survives the cost-aware selection rule." B's +1.69 points sat 0.01 points inside a band of 1.7, it cost 26% more tokens, and on the coding instance a within-band score change earns nothing by itself.

    Where it sits in the field

    The limits, in the authors' words

    Five details a careful reader notices

    Keep going

    The takeaway. Regularize the loop, not the file. A self-improving harness is a search that looks at the same finite set again and again, and every gain it measures is part mechanism, part noise, part names copied from the traces, part tokens spent. RRSI leaves every component editable and controls only how repeated feedback becomes permanent state: few attributable edits, a memory of what failed, a critic before the score, a floor under the best, a price on tokens, and an eviction notice for machinery that stopped paying. What survives, in the paper's words, is "a mechanism rather than a fit to the evolution suite."
    On the agentic workspace instance (δ = 0.004, ws = 1414, wc = 15), a prompt-only candidate measures +0.3 points (ΔS = 0.003) and cuts tokens by 10%. The incumbent holds the best score so far. What happens?

    Now press Present or Teach and explain, out loud and from memory, why a harness that scores higher on its evolve split can be the worse harness, and which gates RRSI puts between a measured gain and permanent harness state. If you can, you own this paper. Then go back to the harness file and run No rules and RRSI side by side.

    Based on "RRSI: Regularized Recursive Self-Improvement of Agent Harnesses" by Peng Xia, Rujun Han, Zifeng Wang, Yanfei Chen and colleagues at Google Cloud AI Research, UNC-Chapel Hill, Stanford and Washington University in St. Louis (2026)
    Read the paper · Code · Project page · Back to Veanors