An AI writes research ideas faster than anyone can test them, and most of them sound good; this paper builds a lab that runs every idea for real and feeds the scores back. In ten rounds that loop lifted a math model from 48.0% to 69.4%, and taught a small GPT to reach its target in 19.7 minutes instead of 35.9.
Learn how an automated lab turns AI-written research ideas into GPU experiments, and why searching over the results beats retraining the AI on them.
Pick a way to learn and a problem, then watch the rounds go by: ideas written, run and scored. Search keeps the winning ideas in the AI's prompt; reinforcement learning retrains the AI on the scores. Then we build, piece by piece, the lab and both loops.
You need what training a neural network means and the idea of a test score for a model. We build the rest from zero.
Each dot is one idea, written by an AI and run on a GPU (the chip that trains it). Watch the best score climb, round by round.
The lines are the paper's: search from Figure 3 and Table 1 (the best idea of each round, read off the plot, so approximate), reinforcement learning from Figure 5 and Section 5.2 (average and best of each training round, read off the plot), the pink count from Figure 7. The dots are illustrative samples around those lines, and the crashed share is illustrative: near each search writer's run rate in Figure 2, and near the reinforcement-learning model's execution rates in Figure 6. Search rounds are numbered from 1 here; the paper counts epochs from 0. Search used strong commercial models (Claude-4.5-Sonnet for math, Claude-4.5-Opus for the GPT) as the idea writer, while reinforcement learning retrained a smaller open model (Qwen3-30B-A3B), so compare the shapes of the curves, not their heights.
Chapter 0
Why an idea that reads well can still lose, and why the only honest judge is running it
It is Friday evening. You are training a small model to solve school math problems, and on a test set it gets 48 out of every 100 right. You ask a chatbot how to do better, and within a minute it hands you fifty ideas. Every one of them reads well. Every one comes with a reason. Which of them will actually move the 48?
You cannot tell by reading. That is not a figure of speech; it has been measured. Earlier studies led by this paper's first author asked expert researchers to review research ideas written by language models, and had human researchers carry some of the ideas out. The ideas often looked convincing on paper and turned out ineffective once someone ran them. A good pitch and a good result are different things.
Try it yourself. Below are eight ideas that AI models wrote for exactly the situation above: improving the training recipe of a 1.5-billion-parameter math model whose starting recipe scores 48.0% on a held-out set (problems the model never trained on) of competition math problems. Every one was turned into code and run. For each, guess whether it beat the starting recipe, then see what happened.
Read the idea, then tell us whether it will beat the starting recipe's 48.0%. The bar shows what the idea really scored when it was run.
No idea has been run yet. Make your first guess.
Every idea and every score is from the paper: Table 3 and Table 5, each run by the paper's automated executor on the math environment (baseline 48.0%). The ideas are paraphrased in plain words; the originals are longer and more technical. We picked these eight to show a range, including one whose code never ran; they are not a random sample.
How did you do? The attention idea sounds sophisticated and lands below the start, at 45.2%. The diversity bonus sounds like good scientific hygiene (explore more, don't all say the same thing) and scores 19.2%, less than half the starting point. A plain tweak to the reward, paying a little more for correct answers that show more reasoning steps, is among the best at 65.6%. And the value-function idea, which a textbook would call principled, never produced a score at all: the code to test it failed.
That is the whole motivation of the paper in one device. If you want an AI to do research, and not just write research proposals, it has to learn what works, and the only reliable teacher is the experiment. The authors call this execution grounding: tying idea generation to the measured result of actually carrying the idea out.
Two words need pinning down, because the whole lesson leans on them. An idea, here, is a short paragraph in plain English that describes a change to a training recipe: "add a buffer of the best past answers", "make the feed-forward layers five times wider". Executing an idea means turning that paragraph into code, running the full training job with the change, and measuring the resulting model on a fixed test. The output of execution is one number per idea, or a crash.
So why has nobody simply done this? Two reasons, and the paper is organised around them. First, running an idea is expensive and fiddly. Someone has to turn a paragraph into working code, find a free GPU (a graphics processing unit, the kind of chip neural networks are trained on), launch a training job, wait, and read back the number. For one idea that is a morning's work. For thousands of ideas, automatically, it is an engineering project.
Second, even with the scores in hand, it is not obvious that a language model can learn from them. A score is a single number; the idea was a paragraph. Can the model work out what made the good ideas good, and write better ones next time? Nothing guarantees it. The number could be too noisy, the space of ideas too large, the model too set in its ways.
The paper, by Chenglei Si, Zitong Yang, Yejin Choi, Emmanuel Candès, Diyi Yang and Tatsunori Hashimoto at Stanford, answers both questions in order. It builds an automated idea executor: a system that takes a batch of plain-English ideas and returns a benchmark number for each, running hundreds of experiments in parallel on GPUs. It points that executor at two real research problems that the whole field cares about: making a small GPT train faster (a pre-training problem, teaching a model language from scratch) and making a math model learn better from practice (a post-training problem, improving a model that already exists).
Then it tries two ways of learning from the numbers. The model that writes ideas, which the paper calls the ideator, stays at the centre of both. In evolutionary search the ideator is never changed. Instead, each round, the best ideas so far are pasted back into its prompt, and it is asked for variations on them and for fresh ideas unlike anything tried. Think of a plant breeder: keep the best plants, cross them, and sow a few wild seeds every season.
In reinforcement learning (RL) the ideator itself is retrained. Its weights are nudged so that ideas that scored well become more likely to be written next time, and ideas that scored badly become less likely. It is the same family of method that taught recent models to reason through math problems, with one difference: the reward is no longer "is the final answer correct" but "how well did this research idea do when we ran it".
The hero shows how the two compare. Search, in ten rounds, found a math recipe at 69.4% (the start was 48.0%, and the best student in a Stanford graduate class that ran the same assignment reached 68.8%) and a pre-training recipe that reaches its target in 19.7 minutes instead of 35.9. Reinforcement learning made the ideator better on average but not better at its best: it learned to play safe, and by the end almost every idea it wrote was one of the same two easy tricks.
Here is the road, in order. Each chapter explains one part of the instrument in the hero.
Chapter 1
How a paragraph becomes a code change, a GPU job and a score, with nobody in the loop
Picture the manager of a busy research lab. On Monday morning fifty proposals land on her desk, each a paragraph long. By Friday she wants a number next to every one: did it help, and by how much? She has a team of engineers who turn proposals into code, a queue that hands out the lab's GPUs, and a wall of machines that run the jobs. Her work is logistics: making sure every proposal gets built, queued, run and scored, and that the numbers come back in one place.
The paper replaces that whole lab with software. The authors describe it as a high-level interface, an API in programmer's language: you call it with a batch of plain-English ideas, and it returns the benchmark score of each one. Everything between those two ends happens without a person. That is the automated idea executor, and it is the foundation for everything else in the paper, because both learning methods later in the lesson treat it as their source of truth.
It has three parts, one for each job the lab manager's team did. The Implementer turns each idea into a code change. The Scheduler finds GPUs for the changed code. The Worker runs the experiment and reports the result. Let's walk one idea through all three, starting with the part that surprises most people: how a paragraph becomes code.
Every research problem in the paper comes with a baseline codebase: a working training program that already runs end to end and produces a score. Ideas never start from a blank page. They are changes to that program. So what the Implementer asks a language model for is not a whole program but a diff: a compact list of which lines to delete and which lines to add, the same format programmers use to review each other's changes. Here is a real one from the paper's appendix, part of the code for one of the math ideas you met in Chapter 0:
--- grpo.py (baseline, trimmed) +++ grpo.py (after the idea) -learning_rate=1e-5 +learning_rate=3e-5 cliprange=0.2 -loss_type=grpo_clip +loss_type=reinforce_with_baseline
Red lines go, green lines come in, and the unmarked line is context: an unchanged neighbour that tells the tool exactly where the change belongs. Applying a diff to a codebase is called patching. A patch can fail even when the idea is sensible, because the language model misremembered a line, got the indentation wrong, or pointed at a function that does not exist. When the context lines do not match the real file, the patch tool refuses.
The Implementer's recipe is built around that fragility. It runs on an ordinary CPU machine with fast disk and network access. For each idea it sends the code execution model (the language model that writes code) both the idea and the whole baseline codebase, and asks for 10 diffs in parallel. Each diff is tried. If one fails to patch, the model is shown the patch tool's error log and asked to revise that diff, at most 2 times. The Implementer keeps the first diff that patches cleanly, applies it, zips the changed codebase, and uploads the zip file to a cloud storage bucket.
Why ten at once instead of one careful attempt? The paper's stated reason is efficiency: the ten diffs are requested in parallel, so the wall-clock cost stays close to one round trip, and each extra sample is a fresh chance. Revising with the error log fixes the cheap mistakes, like a wrong line number. The device below lets you feel how quickly those chances add up, and where they stop helping.
Top: the Implementer's 10 diffs, each with a first try and up to two fixes. Bottom: the Scheduler and the Worker. Pick a kind of idea, adjust the two chances, and press Send the idea.
The recipe is the paper's (Section 2.2): 10 parallel diffs, up to 2 revisions each on a patch failure, keep the first diff that applies. The two chances are illustrative knobs, not measured values, and treating every attempt as an independent coin flip is a simplification. A "library" idea stands for the failure the paper names in Appendix A.2: code that needs a package the lab's environment does not have.
Turn the patch chance down to 10% and watch: most first tries fail, the fixes rescue a few, and one diff almost always gets through. That is the arithmetic of many independent chances. If every attempt patched with probability p, and the attempts were independent, the only way to end up with nothing is for all of them to fail:
Let's put a number on it, step by step. Say a hard idea patches on any single attempt only one time in ten, so p = 0.1. One attempt fails with probability 0.9. Thirty independent failures in a row happen with probability 0.930. Take it in pieces: 0.910 = 0.349, so 0.930 = 0.3493 = 0.349 × 0.349 × 0.349 = 0.0424. The idea is left with no diff about 4 times in 100, and gets one about 96 times in 100. With a single attempt and no revision it would have been 10 in 100.
1 − 0.930 = 1 − 0.042 = 0.958 (versus 0.1 for one attempt)
Now look at the second knob. A diff that patches is not code that works. The patched program can still crash while training: a shape mismatch, a missing import, a division by zero. The Implementer never sees those errors; it only ever sees the patch log. So the overall chance of getting a score is roughly the chance of patching times the chance the patched code trains to the end. With p = 0.1 and a 70% chance of training cleanly, that is 0.958 × 0.70 = 0.67. Past a handful of attempts, extra diffs stop helping, and whether the code actually trains is what decides the score.
The Scheduler is the queue. On a fixed clock, it looks in the bucket for codebases it has not run yet, downloads them, checks what the research environment needs (how many GPUs, of what kind), and writes a job configuration: the instructions a GPU cluster needs to start the run. When GPUs free up, it hands the job over.
The Worker is the machine that does the work: it starts the training run from the patched codebase, then evaluates the result with the environment's own test code. If the run succeeds, it uploads the full logs, every performance number, and the complete record of the idea (its text, its code change, its execution log) to wandb, short for Weights & Biases, a hosted service for tracking machine-learning experiments. If the run fails, the Worker simply halts. The idea writer, which the paper calls the user of this whole API, then downloads the results for its batch, with the training logs attached.
| Stage | Takes in | Hands on | Where it can fail |
|---|---|---|---|
| Implementer | An idea (plain English) and the baseline codebase | A patched codebase, zipped, in a cloud bucket | No diff patches after 10 samples × 3 attempts |
| Scheduler | New zip files, on a clock | A job configuration matched to free GPUs | Only waits; it does not judge the code |
| Worker | A job and its GPUs | Scores, logs, idea text and code change, in wandb | The code crashes: the Worker halts, no score |
Two design choices are worth pausing on, because they decide what the later learning can and cannot exploit. First, only the idea writer ever learns. The Implementer's language model, the Scheduler, the Worker and the scoring code are all fixed; the paper's overview figure says it plainly: only the ideator is updated. That keeps the scoreboard stable. If the grader could change, a rising score might mean the grader got easier.
Second, the executor is a reward function in the full sense a reinforcement-learning researcher means: a fixed procedure that maps an action (here, an idea) to a number. Chapter 4 plugs it into a search loop; Chapter 6 plugs it into a training loop. And because the executor treats every idea the same way, a failed build and a failed training run look identical from the outside: no score. Later, when reinforcement learning assigns such ideas a reward of 0, that detail will turn out to matter a great deal.
Chapter 2
The pre-training and post-training problems, how each one is scored, and how the authors stopped the AI from cheating
An executor needs something to execute against. Think of a running club that wants to test training plans. It needs a fixed track, a fixed distance, a stopwatch everyone trusts, and a current club record to beat. Change the track between runs and nothing can be compared. The paper builds the research version of that track twice, and calls each one a research environment.
Every environment has five parts: a research problem (what we are trying to improve), a baseline codebase (the current recipe, which runs and produces a score), a benchmark (the test the score comes from), fixed training and evaluation data, and a metric (the single number that says who won). The authors chose problems that pull in two directions at once. They are open-ended, so there is room for genuinely new algorithms. And they already have well-established baselines and metrics, so measuring whether an idea helped is straightforward.
Both problems are about training language models, which makes them realistic in a pointed way: they are the problems AI researchers actually work on, and progress on them could feed back into better AI models. Neither sets any limit on the kind of idea allowed. Anything from careful tuning of existing settings to a new architecture or a new training algorithm is in scope, as long as it can be written as a change to the baseline code.
Pre-training is the first stage of building a language model: a network reads a huge amount of text and learns to predict the next word, or more precisely the next token (a word or a piece of one). How good it has become is measured by the validation loss: the model's average surprise at the true next token on text it never trained on. Lower is better. A loss of 3.28, the number that matters here, means the model is still surprised a lot, but far less than when it started.
The environment is built on a public competition called the nanoGPT speedrun (the modded-nanogpt project by Keller Jordan and collaborators). The task: pre-train a GPT-2 sized model with 124 million parameters on FineWeb, a large collection of filtered web text, until its validation loss reaches 3.28, using 8 H100 GPUs (a common data-centre GPU), as fast as possible. People around the world compete on the wall-clock time. The paper's baseline codebase, adapted from the speedrun, takes 35.9 minutes. The best human entry on the leaderboard, as of December 2025, takes under 2.1 minutes. That gap is the room left for ideas.
The authors changed the rules in two ways. The first is about the clock. If every idea ran until it reached 3.28, a good idea would finish in 20 minutes and a bad one might never finish, which makes a batch of runs hard to schedule and impossible to compare fairly. So during search and reinforcement learning, every run gets the same fixed 25 minutes of training, and the score is the validation loss it reaches in that time. To make "higher is better" hold everywhere, the paper uses 1 ÷ loss as the score. Only the single best recipe at the end is timed the speedrun way, so it can be compared with the leaderboard.
Two toy loss curves: recipe A (the baseline) and recipe B (an idea). Lower is better. Move the fixed training budget and the target loss, and compare the two verdicts. Then let B cheat.
Both curves are invented shapes, in toy minutes, chosen to show how the two scoring rules behave; they are not the paper's measurements. The paper's real numbers: a fixed 25-minute budget for search and reinforcement learning, a 3.28 target for the final timing, 35.9 minutes for the baseline and 19.7 for the best recipe found.
With B faster all along, the two verdicts always agree: whichever recipe is lower at the budget also reaches the target first. That is why a fixed clock is a fair stand-in for a stopwatch. Switch B to "starts slow, ends fast" and pull the budget down to a few minutes: now the fixed clock prefers A, while the stopwatch at a 3.28 target prefers B. A short budget can hide ideas that pay off late. The paper's 25 minutes is long enough to reach the neighbourhood of the target, so the two rules are unlikely to disagree there.
Now the numbers the paper reports, worked through, because the 1 ÷ loss score hides how large the gains are. The baseline's loss in the fixed budget is 3.255, so its score is 1 ÷ 3.255 = 0.3072. The best recipe the search found reached 3.1407, a score of 1 ÷ 3.1407 = 0.3184. The score rose by 0.3184 − 0.3072 = 0.0112, which is 3.6% (3.255 ÷ 3.1407 = 1.036). That sounds small. Timed on 8 H100s the speedrun way, the same recipe reaches 3.28, a higher (easier) loss than 3.1407, in 19.7 minutes against the baseline's 35.9: 16.2 minutes saved, which is 45% of the training time, or 1.82 times as fast. These two numbers are measured differently on purpose: the 25-minute figures compare losses at a fixed clock, during search, while the 19.7 and 35.9 minute figures time how long the same recipes take, on 8 H100s, to reach the easier 3.28 target the speedrun uses. The paper does not say what hardware the 25-minute search runs used, so compare losses with losses and minutes with minutes, never mix the two across rows.
35.9 ÷ 19.7 = 1.82× faster (from a 3.6% higher score)
The second rule change is about cheating, which the field calls reward hacking: finding a way to raise the score without doing the task. A language model predicts each token from the tokens before it, never from the ones after; the attention mechanism, which lets each position look at other positions, is masked so it cannot see the future. An idea that quietly loosens that mask lets the model peek at the answer, and its loss falls through the floor. The paper reports that this happened multiple times during the authors' early development.
The fix has two parts. All evaluation settings are frozen, so an idea cannot, for example, shorten the validation text. And final validation runs through an inference function written by the authors that predicts one future token at a time: the model is fed the text up to position t and asked for token t + 1, then the next, so no change to the attention mechanism can smuggle in tokens it has not been given yet. Try the peek toggle in the device to see what an unguarded leak looks like.
Post-training is what happens after pre-training: the model already knows language, and now it is shaped for a job. Here the job is competition math. The starting model is Qwen2.5-Math-1.5B, a 1.5-billion-parameter model already specialised for mathematics. It practises on the MATH dataset of competition problems, and the baseline recipe that trains it is GRPO, short for group relative policy optimisation (introduced in the DeepSeekMath paper, Shao and colleagues, 2024).
You will meet GRPO twice in this lesson, once as the thing being improved and once as the tool doing the improving, so here is the plain version. For each practice problem the model writes a group of answers (8 in this codebase). Each answer gets a reward: did it reach the right final answer? Each answer is then pushed up or down according to how much better it did than the average of its own group. Answers that beat their siblings become more likely; answers that trail them become less likely. No separate critic network (a second model trained just to predict how good a partial answer is) is needed; the group is its own yardstick.
The metric is the model's accuracy on the MATH validation set, and the environment reports the highest validation accuracy reached at any point during a fixed training time. The baseline recipe scores 48.0%. To rule out cheating, all validation code lives in a separate file that the executor can neither read nor change. The same assignment, with the same time budget, was set to the students of Stanford's CS336 graduate course on language models; the best student solution reached 68.8%.
Chapter 3
Six AI models write fifty ideas each; we count how many turn into a score, and whether the best one beats the start
Imagine a cooking contest where each chef hands a written recipe to a line cook, and only dishes that make it to the judges' table get scored. If most recipes come back burnt or half-made, the judges learn almost nothing about which recipes were good. Before a loop can learn from a scoreboard, the scoreboard needs entries. So before any search or training, the paper asks a plain question: when a frontier AI writes research ideas and an AI implements them, how many of those ideas produce a score at all?
That question hides two roles. The ideator writes the idea. The executor, meaning the language model inside the Implementer, writes the code. The paper tests them in two arrangements. In self-execution each model implements its own ideas: Claude-4.5-Opus, Claude-4.5-Sonnet and GPT-5 each write 50 ideas and then write the diffs for those same 50. In the second arrangement, GPT-5 is the executor for everyone, and four other models write ideas: Claude-4.5-Sonnet, Gemini 3, and two models whose weights are public, Kimi-K2-Thinking and Qwen3-235B-A22B.
Three numbers are recorded for every model on every problem. The completion rate is the share of ideas that were executed successfully with a valid, non-zero result. The average performance is the mean score over the ideas that did execute. And the best performance is the single best score among them. That last one has a name worth learning: best-of-N, the simplest possible way to use a scoreboard. Write N ideas, run them all, keep the winner. No learning, no feedback, just a lottery with N tickets.
Top: the share of 50 ideas that produced a score. Bottom: the average and the best score among those that ran, against the starting recipe. Choose who writes the code and which problem.
All numbers from Figure 2 (50 ideas per model), with the more precise values the text of Section 3 gives where it gives them (60.4%, 50.2%, 3.237, 3.238). Averages count only ideas that ran. On the small GPT the score is validation loss, so lower is better and the chart is flipped: up means better. Kimi-K2-Thinking's average loss of 4.29 is off the chart.
Start with the top panel. When each model builds its own ideas, a large share run: on the math problem 80% of Claude-4.5-Opus's ideas, 84% of Claude-4.5-Sonnet's and 52% of GPT-5's; on the small GPT 96%, 90% and 64%. The pre-training problem runs more reliably for all three, and the two Claude models execute markedly more of their ideas than GPT-5 does. The paper singles out the pre-training environment: over 90% for both Claude models.
Now switch to GPT-5 building everyone's ideas. The rates are still decent, from 42% to 78%, but look at Claude-4.5-Sonnet. The same kind of ideas that ran 84% of the time on the math problem when Sonnet wrote its own code run only 42% of the time when GPT-5 writes the code (and 90% against 78% on the small GPT). An idea is not implementable or unimplementable in the abstract. It is implementable by a particular executor. One plausible reading, which the paper does not test, is that a model tends to write ideas in a form it knows how to build itself.
Here is how completion rates turn into a scoreboard, worked step by step. Sonnet on the math problem executes 84% of 50 ideas: 50 × 0.84 = 42 scores. Built by GPT-5, 50 × 0.42 = 21 scores. Now suppose, purely to illustrate, that any executed idea has a 1-in-10 chance of beating the baseline, independently. The chance that none of 42 does is 0.942. Since 0.910 = 0.349 and 0.92 = 0.81, we get 0.942 = 0.3494 × 0.81 = 0.0148 × 0.81 = 0.0120. So best-of-50 finds a winner 98.8% of the time. With 21 scores, 0.921 = 0.3492 × 0.9 = 0.122 × 0.9 = 0.109, and the winner appears 89.1% of the time.
1 − 0.942 = 0.988 versus 1 − 0.921 = 0.891
The completion rate is a tax on every later step: every idea that crashes is a lottery ticket thrown away. It matters even more for reinforcement learning, where, as Chapter 6 shows, a crashed idea is not merely wasted but actively punished.
Now the bottom panel, which holds the chapter's most important pattern. On every problem, in both arrangements, every model's best idea beats the starting recipe, and every model's average idea is worse than it. On the math problem the averages sit between 35% and 45%, all below 48.0%, while the best ideas reach 49% to 60.4%. On the small GPT the average losses sit between 3.30 and 4.29, all worse than the baseline's 3.255, while the best ideas reach 3.237 to 3.25.
Read that twice, because it is the thesis of Chapter 0 in numbers. The typical AI-written idea makes the recipe worse. A few make it better. Without execution you cannot tell which few. With execution and nothing else, best-of-50 already beats the baseline: Claude-4.5-Sonnet's best math idea scores 60.4% against 48.0%, and Claude-4.5-Opus's best pre-training idea reaches a loss of 3.237 against 3.255. Even the open-weight Qwen3-235B, with GPT-5 building its code, gets 50.2% and 3.238.
Two more readings repay a slower look, because "completion rate" hides real texture. First, what counts as a failed execution: a patch that never applies and a patched program that trains for a few steps and then crashes are recorded exactly the same way, as no score, even though one never touched a GPU and the other burned real compute. An idea can fail for reasons that have nothing to do with whether the idea itself was promising. Second, take Kimi-K2-Thinking on the small GPT, with GPT-5 building its code: its average loss is 4.29, far worse than the baseline's 3.255, and yet its best single idea still lands at 3.24, on par with everyone else's best. A high average like that usually means a handful of its ideas trained into a badly broken model, dragging the mean far from the pack, while the rest behaved normally; the average and the best are answering different questions, and for best-of-N only the best one matters. Third, GPT-5 has the lowest completion rate of the three self-execution models, 52% on the math problem and 64% on the small GPT, which reads like a real handicap. It costs less than it looks. Even a bare 52% of 50 ideas is 26 executed attempts, and by the same coin-flip arithmetic as before, if each independently had just a one-in-ten chance of beating the baseline, the chance that none of the 26 would is 0.926 ≈ 0.065, so a winner turns up about 93.5% of the time. A lower completion rate shrinks how many lottery tickets you hold; it does not shrink the odds each ticket carries.
One more detail to carry forward. Best-of-N uses the scoreboard only at the very end, to pick the winner; the ideas themselves are written blind. The obvious improvement is to show the ideator the scoreboard while it is still writing. That is exactly what the next chapter does, and Chapter 5 will compare it with best-of-N at the same number of ideas.
Chapter 4
The search loop keeps the winners, breeds them, and keeps sending out scouts; here it is line by line
A tomato grower wants a sweeter variety. Every season she plants a field, tastes the harvest, and keeps seeds from the sweetest plants. Some of next season's rows are crosses of those winners. A few rows, every year, are wild varieties she has never grown, in case sweetness is hiding somewhere she has not looked. Early on she plants lots of wild rows; once she has a promising line, she spends most of the field refining it. She never studies plant genetics. The harvest is her only teacher, and her notebook of past seasons is her only memory.
That is evolutionary search, one of the oldest ideas in optimisation: keep a population of candidates, score them, keep the best, make variations, repeat. It needs no gradients and no model of why anything works, only a way to score candidates. Genetic programming evolved computer programs this way in the 1990s (Koza, 1994). More recently, researchers realised that a language model makes an excellent source of variations, because it can take a good candidate and propose a sensible change to it (Lehman and colleagues, 2023). The search here is inspired by AlphaEvolve (Novikov and colleagues, 2025), which evolves code with a language model inside the loop.
What is new in this paper is what gets evolved. Not code, directly, but research ideas in plain English, each scored by the executor of Chapter 1. And crucially, the ideator never changes. Its weights are frozen for the whole search. All the learning lives in the prompt: each round, the model is shown a curated slice of the notebook and asked for new ideas. Learning from examples placed in the prompt, without changing any weights, is called in-context learning, and the prompt's size limit, the context window, is the notebook's page count.
The paper's Algorithm 1 takes a batch size N (ideas per round), a number of rounds T, and the baseline's score β. It also takes an exploitation rate a, the percentage of each batch spent refining winners, which changes over the rounds on a schedule.
The split is one line of arithmetic, rounded down so the counts are whole ideas:
Worked through for the small GPT, where N = 80. In round 2, the first split round, a = 50: 50 ÷ 100 × 80 = 40, so 40 ideas exploit and 80 − 40 = 40 explore. Suppose a later round has a = 65 (an illustrative value; the paper does not publish its exact schedule): 0.65 × 80 = 52, so 52 exploit and 28 explore. On the math problem, N = 50 and the same a = 65 gives 0.65 × 50 = 32.5, which rounds down to 32 exploiting and 18 exploring. The rounding always favours exploration by a fraction of an idea.
⌊0.65 × 50⌋ = ⌊32.5⌋ = 32 exploit, 50 − 32 = 18 explore
Why keep both? Because each fails alone, in opposite ways. Exploitation alone is a hill climber: it refines whatever won in round 1 and walks up the nearest hill, even if a mountain stands a valley away. Exploration alone is a scout who reports every new valley but never builds anything: it keeps finding new ground and never polishes the best spot it found. Starting half and half, then shifting toward exploitation, is the grower's strategy: scout widely while you know little, refine hard once you know where the good ground is. The same shape appears in annealing, the metalworker's slow cooling that gives the process its name.
A toy map of possible ideas. The AI's blind guesses cluster near its comfort zone; the warm patches, hidden from the AI, are where ideas beat the start. Pick a strategy and press Run 10 rounds. The chart compares its best idea so far with best-of-N on the same number of ideas.
A toy, not the paper's data. The map, the scores (the start is 0.50), the crash chance (higher far from the comfort zone, as the paper observes for complex ideas) and the late-round rate of 90% are invented. The loop is Algorithm 1: round 1 blind; then exploit (variants and crosses of the best past ideas that beat the start) and explore (a random sample of past ideas in the prompt, a new idea far from all of them). Best-of-N draws every round blind, 10 ideas a round, 100 in all, the same budget. Rounds are numbered from 1 here, as in the hero and Chapter 5: round r is the paper's epoch r − 1.
Run the paper-like schedule and watch the order of events. Round 1 scatters ideas around the comfort zone; a few land on a small warm patch. The scouts (hollow circles) fan out and one of them lands on the slope of a larger patch. From then on the exploiters (filled dots) swarm that slope and climb it. Best-of-N, with the same hundred ideas, keeps sampling the comfort zone and never leaves it. Now try "only exploit": it climbs the first hill it found and stays there. And "only explore" finds the big patch but never polishes it, because no one refines the scouts' finds.
The toy says yes; the paper checks it on the real thing. On the small GPT, with GPT-5 as both ideator and executor and 80 ideas per round, the authors compare the first three rounds of search against best-of-N with the same number of ideas: best of 80, of 160, of 240. Read off Figure 4, in the 1 ÷ loss score: round 1, the blind round, is nearly a tie (about 0.3087 for search and 0.3089 for best-of-N, different only by sampling luck). By round 2 search has already pulled ahead to about 0.3117, and by round 3 it reaches about 0.3128, while best-of-N has crept only to about 0.3091.
Convert those to losses to feel the gap: 1 ÷ 0.3128 = 3.197 for search after round 2, against 1 ÷ 0.3091 = 3.235 for the best of 240 blind ideas. Blind sampling has almost stopped improving, because the ideator keeps writing the same kind of idea; search has pulled clearly ahead after a single round of feedback. The paper's reading: the model is effectively using the trajectories from earlier rounds to write better ideas in later ones.
Here is the whole loop in Python, from scratch. The ideator is any function that turns a prompt into a list of idea strings, and execute is the executor of Chapter 1, returning a score or None for a crash.
pythonimport math, random CONTEXT_LIMIT = 40 # fits one prompt def positive(book, beta): return [ (i, s) for i, s in book if s is not None and s > beta ] def exploit_step(ideator, book, beta, n): wins = positive(book, beta) if not wins: return [] msg = f"Beats baseline: {wins}" return ideator(msg, n) def explore_step(ideator, book, n): lim = CONTEXT_LIMIT seen = random.sample( book, k=min(len(book), lim) ) msg = f"Tried: {seen}. Be different." return ideator(msg, n) def execution_guided_search( ideator, execute, N, T, beta, rate ): # round 1: blind, no feedback yet start = ideator("Propose ideas.", N) notebook = [ (i, execute(i)) for i in start ] for r in range(2, T + 1): a = rate(r) # 50 at round 2, up n_exp = math.floor(a / 100 * N) n_expl = N - n_exp exploit = exploit_step( ideator, notebook, beta, n_exp ) fill = n_exp - len(exploit) explore = explore_step( ideator, notebook, n_expl + fill ) batch = exploit + explore new = [ (i, execute(i)) for i in batch ] notebook += new return notebook
Notice what is not in the loop: no gradient, no loss function, no change to the ideator. The only state is the notebook, a list of (idea, score) pairs, and the only lever is which of them go into the next prompt. That simplicity is why it works with any model behind an API, and why it is cheap in ideas: the paper's searches run just ten rounds.
Chapter 5
Three AI models search for ten rounds on each problem; one keeps climbing, two level off, and all beat the starting recipes
Give three cooks the same kitchen, the notebook method from the last chapter, and ten evenings each to improve one dish. Does every cook keep improving night after night, or do some settle on a recipe after a few evenings and stop getting better? That is the experiment this chapter reports.
The three cooks are three frontier AI models, Claude-4.5-Opus, Claude-4.5-Sonnet and GPT-5. Each runs the search loop of Chapter 4 for ten rounds on each problem, with the same model writing the ideas and the code. That is 50 ideas a round on the math problem and 80 on the small GPT: 500 and 800 ideas per model, every one of them sent through the executor, and most of them run on GPUs. The question is simple to state. Does feeding scores back make the ideas better, and does it keep making them better the longer you search?
The paper plots the best idea of each round, and the device below lets you replay it. Two things to watch for. First, whether a model's best climbs above the dashed starting recipe (all of them do). Second, the shape: whether the line keeps rising across all ten rounds, which the paper calls a scaling trend, or rises quickly and then goes flat, which it calls saturating. A scaling trend is the one you want: it means more search buys more discovery.
Each line is one model searching on its own. Switch the problem, switch between each round's best idea and the best so far, and turn models on or off.
Figure 3, read off the plot (approximate), with each model's peak set to the exact value the paper prints (Table 2 and Appendix A.2: losses 3.1407, 3.2081 and 3.1697; accuracies 69.4%, 61.6% and 60.0%). Rounds are numbered 1 to 10 here; the paper counts epochs 0 to 9. Each model is both ideator and executor.
On the small GPT the three lines tell three different stories. Claude-4.5-Opus climbs round after round: in 1 ÷ loss, from about 0.3090 in round 1 to 0.3184 in round 10, still rising at the end. That is the scaling trend, and it is the only one of its kind in the figure. Claude-4.5-Sonnet climbs for a few rounds and then stops at about 0.3117, a loss of 3.208, from round 5 on. GPT-5 rises fastest at first, peaks in round 4 at a loss of 3.170, and then wanders below its own peak.
On the math problem the order flips. Claude-4.5-Sonnet reaches 69.4% in round 3 (the paper's epoch 2) and is never beaten, by itself or by anyone; Claude-4.5-Opus climbs slowly from about 51% to 61.6%; GPT-5 swings between about 49% and 60%. Opus still shows the upward trend on both problems, and Sonnet saturates on both. The authors' summary: models often generate meaningful algorithmic ideas during search, but they tend to saturate early and only occasionally show scaling trends.
What did Sonnet find? Something that surprises people who know reinforcement learning: it made the algorithm simpler. Recall GRPO from Chapter 2: each answer is pushed up or down by how much better it did than its group's average. The standard GRPO objective adds two safety devices on top of that. One is the importance ratio: the answers were sampled by a slightly older copy of the model, so each update is reweighted by how much more (or less) likely the current model finds that answer. The other is clipping: that ratio is capped to a narrow band, 0.8 to 1.2 in this codebase (a clip range of 0.2), so no single update can move the model too far.
Sonnet's idea removed both. It trains with plain policy gradient with a group-average baseline (the policy gradients lesson builds it from zero): raise the probability of each answer in proportion to its reward minus the group's average, with no reweighting and no clipping. In this codebase that almost certainly maps onto a single existing setting: Table 2 files the 69.4% idea under settings changes, not new algorithms, and every Sonnet diff in the appendix that reaches that neighbourhood flips the loss_type from grpo_clip to reinforce_with_baseline, the very line you saw change in Chapter 1's diff, and a setting many of Sonnet's later ideas carry. The paper states that this outperforms the standard GRPO objective in this particular experiment setup, and that Sonnet exploited the finding in every later round, combining it with precise tuning of settings such as the learning rate (how big a step each training update takes; raised from 1e-5 to 3e-5 in the diffs the paper shows). Many of Sonnet's later ideas in the appendix open by keeping "the proven 3e-5 learning rate" and the new loss type: the notebook at work.
Opus's winner on the small GPT, found in the tenth round, is the opposite of simple. The paper prints it in full, and it stacks a dozen changes. The feed-forward layers switch to a gated design called SwiGLU and become five times as wide as the model's width, with a learned scale on their output. Extra skip connections (direct wires that carry a signal past several blocks unchanged) are added at every 4th block and every 8th block, each one also receiving the output from 4 (or 8) layers earlier, with learnable weights starting at 0.52 and 0.31. The attention and feed-forward branches get separate learnable scales (starting at 0.98). The input and output word tables are no longer tied together.
Then come a dozen training settings stacked together. None of them matters much alone; what matters is that Opus tuned all of them at once. In plain words: the learning rate (how big a step each update takes) went from 0.0015 to 0.00168. The weight decay (a small constant pull back toward smaller weights, which keeps the model from memorising too hard) went from 0.1 to 0.065. The warm-up (a slow ramp-up of the learning rate over the first steps, instead of starting at full size right away) shortened from 256 to 173 steps. The learning rate then decays on a cosine curve, easing down to 3% of its peak by the end of training instead of stopping abruptly. And the optimiser's second-moment decay (how many recent updates it averages over when deciding how much to trust each one) rose from 0.95 to 0.99, so it reacts a little more slowly to any one noisy update. One trick applies only at evaluation: an exponential moving average (EMA) of the weights, a slowly updated running average (each step keeps 99.9% of the average and mixes in 0.1% of the current weights), is swapped in for validation and swapped back out afterwards. Averaging recent weights smooths out the noise of the last few updates.
That recipe reaches a loss of 3.1407 in the fixed budget. Rerun on 8 H100 GPUs the speedrun way, it reaches the 3.28 target in 19.7 minutes, against 35.9 for the baseline codebase. Worked out, round by round, in loss: Opus's best was about 1 ÷ 0.3090 = 3.236 in round 1, and 3.1407 in round 10, a drop of 3.236 − 3.141 = 0.095. Sonnet stopped at 3.208 and GPT-5 at 3.170, both still clearly better than the baseline's 3.255.
| Math (accuracy, higher is better) | Small GPT (minutes to loss 3.28, lower is better) | |
|---|---|---|
| Starting recipe | 48.0% | 35.9 min |
| Execution-guided search | 69.4% (Claude-4.5-Sonnet) | 19.7 min (Claude-4.5-Opus) |
| Best human expert | 68.8% (best CS336 student) | 2.1 min (speedrun record, December 2025) |
The humans put both results in perspective, and the two comparisons point in opposite directions. On the math problem, the search's 69.4% edges past the best of a graduate class that attacked the identical assignment under the identical time budget: 69.4 − 68.8 = 0.6 points. On the small GPT, the human record is 19.7 ÷ 2.1 = 9.4 times faster than what the search found. The authors read that gap as significant headroom, both for more capable models and for better search methods.
Chapter 6
Reinforcement learning uses the executor's scores to retrain the ideator itself; the average idea gets better, the best one does not
A writing coach cannot tell a student what to write, but she can hand back every essay with a mark. Over hundreds of essays the student absorbs what earns marks, without being told why. That is the bet behind reinforcement learning (RL): let a model act, score the result, and nudge the model's weights so that high-scoring behaviour becomes more likely. It is exactly the opposite of search. Search keeps the model fixed and changes what it reads. RL keeps the prompt fixed and changes the model.
RL has recently transformed how language models do math and write code, most visibly in DeepSeek-R1, because those domains come with a verifiable reward: the final answer is right or wrong, the tests pass or fail. The executor turns research ideas into something similar: every idea gets a number. So the paper asks, for what it describes as the first time: can the executor serve as the reward function that trains a model to write more effective research ideas?
The model being trained is Qwen3-30B-A3B, an open model with 30 billion parameters (the "A3B" marks that only about 3 billion of them are active for each token). It is fine-tuned with standard GRPO, the same group-based algorithm you met in Chapter 2 as the thing being improved on the math problem. Now it is the tool doing the improving. The training runs on the Tinker API from Thinking Machines Lab, a service for fine-tuning open models.
There is only one prompt per environment: the baseline codebase, plus a request for new ideas to improve it. Every training round samples a whole group of answers to that single prompt, so the prompt batch size is one. The paper notes this resembles earlier work that ran RL on a single training example. Each answer, called a rollout, is a thinking trace (the model reasoning to itself) followed by the idea, at most 8,192 tokens in all. Only the idea is sent to the executor; the thinking is thrown away.
The groups are large, to keep training stable: 256 ideas per round on the math problem and 128 on the small GPT. Each math idea trains on 1 GPU and each small-GPT idea on 8, so a single round of training occupies 256 × 1 = 256 GPUs or 128 × 8 = 1,024 GPUs at once, just to score one batch of ideas. The reward is the idea's validation accuracy on the math problem, and 1 ÷ loss on the small GPT. An idea whose execution fails gets a reward of 0. Hold on to that last rule; it is the hinge of the next chapter.
GRPO compares each idea with its own group. The advantage of an idea is how much better it scored than the group's average, measured in units of the group's spread (its standard deviation). Ideas with positive advantage are made more likely, ideas with negative advantage less likely, in proportion to the size of the advantage.
Let's run it by hand on a tiny illustrative group of four math ideas (the real groups hold 256). Three ran and scored 0.52, 0.47 and 0.49; the fourth crashed and scores 0. The mean is (0.52 + 0.47 + 0.49 + 0) ÷ 4 = 1.48 ÷ 4 = 0.37. The gaps from the mean are +0.15, +0.10, +0.12 and −0.37. Square them: 0.0225, 0.0100, 0.0144 and 0.1369, which sum to 0.1838; divided by 4 that is 0.04595, whose square root, the standard deviation, is 0.214. Divide each gap by 0.214: the advantages are +0.70, +0.47, +0.56 and −1.73.
Acrash = (0 − 0.37) ÷ 0.214 = −1.73 (the largest push in the group)
Look at what the update says. The crash produces by far the largest signal: "do not write ideas like that one". The differences among the three working ideas, the thing research actually cares about, produce smaller pushes. The best idea, at 0.52, is only a little ahead of the others. An RL learner reading these numbers learns first and hardest to avoid crashing. Remember this when you meet the collapse in the next chapter.
The first result is a success, and a first in the paper's words: for open-ended research environments, the average quality of the generated ideas can rise with enough training. On the math problem the average reward goes from 0.253 at the start to 0.343 after 40 training rounds. On the small GPT it goes from 0.194 to 0.246 after 68 rounds, which the paper translates into an average loss falling from 5.150 to 4.066 (that is, 1 ÷ 0.194 and 1 ÷ 0.246). The curves look like those of earlier single-example RL on math.
Top: the average idea and the best idea of each training round. Bottom: what one round's batch of ideas might look like, crashes at 0. Drag through training, or press play, and watch which edge of the batch moves.
The two lines are Figure 5, read off the plot (approximate), with the end points the paper states (0.253 to 0.343; 0.194 to 0.246). On the small GPT the paper's figure skips a few rounds; the slider steps through the rounds it plots. The batch at the bottom is illustrative: an invented spread of scores whose average and best match the lines at that round, with crashed ideas at 0.
Now the turn. For scientific discovery the average is the wrong thing to care about. A lab does not need a thousand safe ideas that are each a little better than the last; it needs one idea that beats everything. The paper plots the max reward, the best idea of each training round, and the picture is very different: it fluctuates throughout training with no clear upward trend. On the math problem the best idea of a round sits between about 50% and 54% from the first round to the fortieth. On the small GPT it hovers around 0.309, a loss of about 3.23 to 3.24, from start to finish.
There is a useful piece of arithmetic hiding in the small-GPT numbers. A crashed idea scores 0 and no idea scored above the round's best, about 0.309. So an average of 0.194 means at least 0.194 ÷ 0.309 = 63% of the round's ideas must have run, and an average of 0.246 means at least 0.246 ÷ 0.309 = 80% did. The average can climb a long way simply because fewer ideas crash, without any idea getting better at the top. That is the mechanism the next chapter confirms.
fraction that ran ≥ average ÷ best = 0.246 ÷ 0.309 = 0.80
Chapter 7
The trained writer thinks less, avoids anything that might crash, and converges on two safe tricks
A student learns that the teacher gives a solid B+ to any tidy five-paragraph essay on a safe topic, and a mark of zero to anything that goes off the rails. Ambitious essays sometimes earn an A, but often earn nothing. What does a student who wants the highest average mark do? Writes the same safe essay every week, and stops thinking hard about it. The teacher's average rises. The chance of a brilliant essay falls to nothing.
The paper finds its RL-trained ideator doing exactly that, and it shows the behaviour three ways: in how long the model thinks, in which ideas survive execution, and in how many different ideas it writes at all. Together they explain Chapter 6's puzzle, an average that rises while the best idea stays flat.
Each rollout is a thinking trace followed by an idea. As training goes on, the thinking traces get rapidly shorter while the ideas stay about the same length. Read off Figure 6: on the math problem, thinking falls from roughly 3,000 tokens to roughly 1,250 over 40 rounds, while the idea stays between roughly 230 and 310 tokens; on the small GPT, thinking falls from roughly 3,250 tokens to roughly 750 over 68 rounds, while the idea stays near 600 to 750. That is the reverse of what happened in DeepSeek-R1, where RL on math made the model think longer.
Why would thinking shrink? The authors sort every round's ideas by thinking length, over the first 20 rounds, and compare the 30% with the longest thinking against the 30% with the shortest. Ideas that came from longer thinking consistently execute less often. Read off the figure, roughly 85% against 95% on the math problem, and roughly 45% against 70% on the small GPT. Their hypothesis: longer thinking goes with more complex ideas, complex ideas crash more, a crash scores 0, and so the model learns to think less.
Put the two facts together, a crash scores 0 and complex ideas crash more, and do the arithmetic of what RL maximises. Take two illustrative ideas. A risky idea that would reach 60% if it ran, but runs only half the time, earns on average 0.5 × 0.60 + 0.5 × 0 = 0.30. A safe idea that reaches only 50% but runs 95% of the time earns 0.95 × 0.50 = 0.475. The safe idea wins by 0.475 − 0.30 = 0.175, even though the risky one is the only one of the two that could ever beat 50%.
expected reward = P(it runs) × score: 0.5 × 0.60 = 0.30 < 0.95 × 0.50 = 0.475
GRPO follows expected reward, so it moves probability from the risky idea to the safe one, round after round. Chapter 6's advantage arithmetic showed the same thing from the other side: a crash inside a group produces the largest push of all, away from whatever caused it. Nothing in the objective says "keep writing an occasional long shot in case it is the breakthrough".
Reading through every rollout by hand, the authors saw the diversity of ideas collapse. On the small GPT the model converged on two simple ideas that reliably earn a positive reward. One is replacing RMSNorm with LayerNorm: both are normalisation layers that rescale each vector of activations to a standard size, and LayerNorm additionally subtracts the mean, so the swap is a small, local change to the code. The other is an exponential moving average of the model's weights, the same trick you met in Opus's winning recipe in Chapter 5.
The counts, from Figure 7, out of a batch of 128 ideas per round: 51 of the 128 ideas the model wrote before training were one of those two. By round 68, 119 of 128 were. As fractions: 51 ÷ 128 = 39.8% at the start, and 119 ÷ 128 = 93.0% at the end. More than nine in ten ideas are one of two tricks. The best idea of a round cannot improve when almost every idea in the round is the same idea.
51 ÷ 128 = 39.8% → 119 ÷ 128 = 93.0% (one of the same two ideas)
A toy writer chooses among eight kinds of idea. Safe ones rarely crash but score modestly; ambitious ones crash often but sometimes score high. Each round it writes 64 ideas, they are scored (a crash scores 0), and the group-relative update of Chapter 6 retrains it. Press Train 68 rounds, then try it with a penalty for repeating last round's ideas.
A toy, built to show the mechanism, not to reproduce the paper's curves. The eight kinds of idea, their crash chances and score spreads are invented (the two pink ones are named after the ideas the paper saw RL converge on; "helper model" and "system-level" after the ideas the paper says tend to fail to execute). The starting 40% share of pink ideas and the 68 rounds echo Figure 7. The update is a group-normalised policy gradient on the writer's preferences, standing in for GRPO on a 30-billion-parameter model.
Run it and watch the order of events. The pink bars grow first, because their rewards are never zero. The ambitious kinds shrink, because every crash is a large push away from them. The average reward climbs as crashes disappear. And the best idea of each round, which early on sometimes came from an ambitious kind landing big, stops getting those lucky draws. The penalty for repetition slows the collapse without abolishing it.
The paper sets this beside a known effect in other RL work on language models. pass@k is the chance that at least one of k attempts at a problem succeeds; it measures the ceiling of what a model can reach with several tries. Studies of RL on verifiable tasks (Yue and colleagues, 2025; Wu and colleagues, 2025) found that pass@k stagnates or even decreases after RL, even as single-attempt accuracy rises. The max reward per round is the research version of pass@k, and it stagnates for the same reason: RL sharpens the model onto what already works and prunes the rest. The wider family of effects, language models drifting toward the same few answers, has its own lesson: mode collapse.
Avoiding this convergence, the authors write, is an open problem that likely needs new algorithms beyond standard GRPO. They share three preliminary attempts, each stopped early, on the math problem (Appendix A.1).
Jaccard similarity is simple enough to do by hand: the number of tokens two texts share, divided by the number of distinct tokens in either. Treat words as tokens for illustration. "replace rmsnorm with layernorm" has 4 distinct words; "replace rmsnorm with layernorm and add ema" has 7, including all 4 of the first. Shared: 4. Distinct in either: 7. Similarity: 4 ÷ 7 = 0.571, so that rollout would lose 0.571 of reward (times whatever weight the penalty carries). A genuinely new idea sharing, say, 1 word out of 10 distinct would lose only 0.1. The diversity reward, amusingly, is close to an idea that Claude-4.5-Sonnet itself proposed for the math problem during search: pay answers for being unlike their siblings. As an idea for training a math model it scored 19.2%; as a patch for the idea writer it helped keep ideas varied.
Jaccard = |A ∩ B| ÷ |A ∪ B| = 4 ÷ 7 = 0.571
Chapter 8
Sorting 3,900 searched ideas into settings changes and new algorithms, and finding recent papers rediscovered along the way
Imagine auditing a lab after a busy month. You pull every notebook entry and sort them into two piles. In one pile, somebody turned a dial that already existed: a higher learning rate, a longer warm-up. In the other, somebody built something the lab did not have before: a new loss, a new layer, a new way of choosing training data. The first pile is cheap and safe; the second is where new methods come from. A skeptic's first question about automated research is which pile the AI fills.
The skeptic has a point. If an AI "researcher" only ever tuned knobs, it would be an expensive version of hyper-parameter search, the decades-old practice of trying many settings of a training recipe's dials (learning rate, batch size, weight decay) and keeping the best. That is useful, but it is not research in the sense of inventing methods. So the paper classifies every idea generated during the searches of Chapter 5: 500 per model on the math problem and 800 per model on the small GPT, 3,900 in all across the three models.
The rule is simple. A hyper-parameter idea is any idea that can be implemented by changing configuration settings that already exist in the baseline codebase. An algorithmic idea is any idea that needs new code, something the baseline does not already support. A language model does the sorting, acting as an LLM judge (LLM is short for large language model): it reads each idea and assigns the label. That is a sensible way to label thousands of ideas, and it means the labels are themselves a model's judgement.
For each model, the bar splits its searched ideas into settings changes (light) and new algorithms (solid). Below, the average and the best idea of each kind, against the starting recipe. Right is better on both problems.
Table 2: the share, average and best of each kind of idea for each model over the whole search, classified by an LLM judge. Averages count ideas that ran. On the small GPT the score is loss, drawn with lower losses to the right so that right means better in both views.
The first answer to the skeptic is clear: every model writes mostly algorithmic ideas. On the math problem GPT-5 wrote 95.0% algorithmic ideas and Claude-4.5-Opus 96.3%; on the small GPT, 84.6% and 91.3%. Claude-4.5-Sonnet is the outlier, with 41.1% hyper-parameter ideas on the math problem and 31.3% on the small GPT, far more knob-turning than the others. That fits what Chapter 5 found: Sonnet's winning math recipe was a small change of loss followed by precise tuning.
The second answer is about where the wins come from. In most cases the best idea is algorithmic. On the math problem GPT-5's best algorithmic idea scores 60.0% against 50.2% for its best settings change, and Opus's 61.6% against 50.4%. On the small GPT, GPT-5's best algorithmic idea reaches a loss of 3.170 against 3.195, and Opus's 3.141 against 3.147. The exception, again, is Sonnet: its best math idea, the 69.4%, is counted as a settings change, ahead of its best algorithmic idea at 67.4%; on the small GPT its two kinds tie at 3.208.
Now look at the averages, because they carry a quieter lesson. Take GPT-5 on the small GPT, worked out. Of 800 ideas, 15.4% were settings changes: 800 × 0.154 = 123 ideas. The other 84.6%, 800 × 0.846 = 677 ideas, were algorithmic. The settings changes average a loss of 3.254, almost exactly the baseline's 3.255. The algorithmic ideas average 3.894, worse by 3.894 − 3.254 = 0.640. Yet the best algorithmic idea, 3.170, beats the best settings change, 3.195, by 0.025.
average: 3.894 vs 3.254 (algorithmic worse by 0.640) · best: 3.170 vs 3.195 (algorithmic better by 0.025)
Worse on average, better at the top: that is the signature of a risky bet. It is the same shape as the risky idea in Chapter 7's arithmetic, and it explains the two methods' opposite fortunes. Search keeps only the winners, so risky ideas cost it nothing but GPU time; they can only add to the best. RL maximises the average, and on average the algorithmic ideas lose. The two learning methods are looking at the same shape of evidence and drawing opposite conclusions.
Reading the ideas themselves, the authors notice distinct personalities. Claude-4.5-Sonnet writes more intuitive ideas: balance the difficulty of practice problems as the model improves (scored 64.0%), keep a working-memory buffer of facts while solving (58.0%). Claude-4.5-Opus and GPT-5 are more mathematically inclined: Opus splits the update's ratio into a drifting average and a bounded residual (61.6%) or weights samples by how far their rank is from the expected rank (59.2%); GPT-5 averages the log-ratio over chunks of the response to smooth noisy token spikes (58.2%), or cools or heats each group's update according to how spread out its rewards are (49.4%). Each number in parentheses is that idea's own score.
On the small GPT the winning recipes are heavily tuned bundles, as Chapter 5 showed, but the appendix also lists "atomic" algorithmic ideas from Opus that ran on their own, each a single new mechanism: learnable scales for each attention head's output (loss 3.2386), a learned mix of token and position embeddings (3.2497), tied embeddings with a small learned twist on the output side (3.2499), a gate on the final normalisation (3.2503). Each beats the 3.255 baseline by a little, which is what a single sound change usually does.
Without any retrieval of papers (no search engine, no retrieval-augmented generation, the practice of pasting relevant documents into the prompt), several generated ideas turned out to be close to research papers released within the three months before the paper was written. Claude-4.5-Sonnet proposed rewarding responses for being dissimilar to the other responses in their group, similar to Li and colleagues (2025) on jointly reinforcing diversity and quality. Claude-4.5-Opus proposed "causal context compression", a learned layer that mixes the previous two or three tokens into the current representation before each attention layer, similar to the "canon layers" of Allen-Zhu (2025).
Two honest caveats come with that. The authors do not claim to measure novelty; they say only that rediscovering recent ideas suggests automated researchers could plausibly support work at the frontier. And Opus's canon-like idea appears in the appendix among the interesting ideas that did not execute. The executor could not build it, so it never got a score. The most forward-looking idea in the section is also a reminder of Chapter 1: an idea is only as good as the executor's ability to run it.
Chapter 9
Four limits the authors name, and how each one bends the numbers you have seen
A medicine that cures mice is good news and not yet a cure. It worked in one species, at one dose, measured one way. Every result in this paper has the same shape: it holds for a particular model size, a particular dataset, a particular training budget and a particular executor. The authors are unusually direct about this, and their discussion section lists four limits. Each one changes how you should read a number from an earlier chapter, so let's take them in turn.
Every idea was scored on one small setting: a 124-million-parameter GPT trained for 25 minutes on one dataset, or a 1.5-billion-parameter math model trained for a fixed time on one benchmark. The procedure never tests whether the best ideas still help at a larger scale or on other data. A trick that speeds up a tiny model's first 25 minutes might do nothing for a model a thousand times larger trained for weeks.
This matters more for automated search than for human research, for a sharp reason: a search that sees only one score will exploit everything that raises that score, including quirks of the setting. Sonnet's 69.4% recipe is the clean example. The paper says plainly that dropping the ratio and the clipping outperformed standard GRPO in this particular experiment setup. The authors suggest that future work test generalisation explicitly, and even build it into the objective the search optimises. (One way to picture that, our example rather than the paper's: score every idea at two model sizes and reward only ideas that help at both.)
The executor is a language model writing diffs, with no tools and no ability to install new libraries (Chapter 1). Ideas beyond its reach never get a score, however good they are. The paper's own examples include the canon-like "causal context compression" that echoed a recent paper, and ideas that need extra helper models or system-level changes. The authors call the result noise in the reward signal: the score an idea receives mixes how good the idea is with how buildable it is. The device shows what that does to a leaderboard.
Eight invented ideas, from easy to hard to build. The outline shows how much each would really help; the fill shows what the executor measures. Slide the executor's skill and watch which idea "wins".
A toy with invented ideas, difficulties and effects. The only piece taken from the paper is the rule: an idea the executor cannot build gets no score (in RL, a reward of 0), and the paper reports that complicated changes and unsupported packages are what usually fail.
At low skill, the measured winner is a modest, easy idea, and the ideas that would really help show up as crashes. Raise the skill and the true order emerges. Two things follow. For search, a weak executor hides good regions of the idea space, so the search never climbs toward them. For RL it is worse: the hidden good ideas are actively punished, which feeds straight into the collapse of Chapter 7. A stronger executor is not just a convenience. It changes what the whole loop can learn. The authors point to coding agents with external tools and the ability to install libraries as the next step.
Chapter 7 showed the collapse; the authors add that its causes are not settled. It could be a lack of diversity in the base model to begin with, or the missing incentive to explore in the standard RL objective. Their three repairs were early attempts, not solutions. They also point at a waste in the current setup: RL uses a single number per idea, while the executor produces full training logs, code changes and execution traces. Richer learning signals from those trajectories, beyond a scalar reward, are an open direction.
The only reward in the paper is effectiveness on a benchmark. Research is also judged on whether an idea is new, whether it is interesting, whether it explains something. The authors name novelty and interestingness as more subjective metrics that could complement effectiveness, if they can be measured computationally and added to the training objective. They note that assessing the novelty of the generated ideas was beyond the scope of this paper.
A few more cautions come from the results themselves rather than the discussion. The paper shows one curve per model and problem in Figure 3 and reports no variance across repeats, and Chapter 4's toy showed how much one lucky scout can change a run. The search ran for ten rounds only, so "saturates early" means within ten rounds. The comparison between search and RL is not like for like: search used large commercial models as ideators, while RL trained the smaller open Qwen3-30B-A3B, so the right conclusion is about the shapes (the best idea climbs under search and stays flat under RL), not the heights.
And the scale is real: one RL round on the small GPT keeps 1,024 GPUs busy. Worked out over the paper's run: 68 rounds × 128 ideas = 8,704 training runs, each on 8 GPUs for a fixed 25 minutes of training. That is 8,704 × 8 = 69,632 GPU-runs of 25 minutes, or 69,632 × 25 ÷ 60 = 29,013 GPU-hours if every run trained for its full budget. Crashed runs stop sooner, and building, queueing and evaluating add their own time, so treat it as an order of magnitude: tens of thousands of GPU-hours for one RL experiment. Execution grounding is honest, and it is expensive.
68 × 128 × 8 GPUs × 25 min ÷ 60 = 29,013 GPU-hours (if every run used its full 25 minutes)
Chapter 10
The cheat sheet, the whole system as code, and where this paper sits in the course and the field
You can now read this paper and explain every moving part: why plausible ideas need running, how a paragraph becomes a diff and a GPU job, how the two labs are scored and guarded, how many ideas run, how the search loop balances refining and scouting, what it found, why reinforcement learning raised the average but not the best, what the ideas look like, and what remains unproven. Let's lock it in.
The authors build an automated idea executor: an Implementer that asks a language model for 10 parallel diffs per idea (revising each at most twice on a patch failure), a Scheduler that queues the patched codebases, and Workers that train and evaluate them on GPUs, returning one score per plain-English idea. They turn two research problems into environments: speeding up the pre-training of a 124M GPT (the nanoGPT speedrun, scored as 1 ÷ loss after 25 minutes) and improving GRPO post-training of a 1.5B math model (scored by validation accuracy), both guarded against reward hacking. Frontier models' ideas execute at high rates (up to 96%), and even best-of-50 beats both baselines. Execution-guided evolutionary search, which keeps the ideator fixed and feeds winners and past attempts back through the prompt, reaches 69.4% against 48.0% and a 19.7-minute recipe against 35.9 within ten rounds, but only Claude-4.5-Opus keeps improving with more rounds. Reinforcement learning with the executor as the reward raises the average idea (0.253 to 0.343; 0.194 to 0.246) but not the best, because crashes score 0 and complex ideas crash more: thinking shrinks and 119 of 128 ideas end up as the same two tricks.
| Quantity | Value | Why it matters |
|---|---|---|
| Diffs per idea | 10 in parallel, at most 2 revisions each | Patching is rarely the bottleneck; running is |
| Pre-training lab | 124M GPT on FineWeb, target loss 3.28, 25-minute budget, score 1 ÷ loss | A fixed clock makes runs comparable |
| Post-training lab | GRPO on Qwen2.5-Math-1.5B, MATH, baseline 48.0% | Validation code sealed off from the executor |
| Ideas that run (self-execution) | 52% to 84% (math), 64% to 96% (small GPT) | Enough scores to learn from |
| With GPT-5 as the executor | 42% to 78% | Implementability depends on the executor |
| Search | 50 or 80 ideas a round, 10 rounds, exploit 50% from round 2 and rising | Refine winners, keep scouting |
| Search results | 69.4% (best student 68.8%); 19.7 min (record 2.1) | Past the class, far from the record |
| Search vs best-of-N | about 0.3128 vs 0.3091 after 240 ideas (1 ÷ loss) | The notebook helps from round 1 |
| RL setup | Qwen3-30B-A3B, GRPO, groups of 256 and 128 (256 and 1,024 GPUs) | Crashes score 0 |
| RL average | 0.253 → 0.343 (40 rounds); 0.194 → 0.246 (68 rounds) | It works on average |
| RL best | flat: about 50% to 54%; about 0.309 | It does not raise the ceiling |
| Collapse | 51 → 119 of 128 ideas are one of two tricks | Diversity dies first |
| Algorithmic share of searched ideas | 58.9% to 96.3% (math), 68.7% to 91.3% (small GPT) | Not just knob-turning |
Chapter 4 gave the search loop. Here is the rest from scratch: the executor as the paper describes it, and one round of reinforcement learning that uses it as the reward. Helpers such as try_patch and policy.update stand for a patch tool and a GRPO optimiser step.
pythonimport statistics from concurrent.futures import ( ThreadPoolExecutor, ) # executor: idea in, score out # (None means it crashed) def implement( idea, code, llm, tries=10, revs=2 ): def attempt(_): diff = llm( f"Base:\n{code}\n" f"Idea: {idea}\ndiff?" ) for n in range(revs + 1): ok, log = try_patch( code, diff ) if ok: return diff if n == revs: break diff = llm( f"Idea: {idea}\n" f"Diff:\n{diff}\n" f"Failed:\n{log}\n" "Revise it." ) return None with ThreadPoolExecutor( tries ) as pool: outs = pool.map( attempt, range(tries) ) diffs = [d for d in outs if d] return diffs[0] if diffs else None def execute(idea, env): diff = implement( idea, env.code, env.llm ) if diff is None: return None # never reaches a GPU patched = apply_diff(env.code, diff) job = env.scheduler.submit( patched, gpus=env.gpus ) result = job.wait() # fixed budget if result.crashed: return None return env.metric(result) def execute_batch(ideas, env): n = len(ideas) with ThreadPoolExecutor(n) as p: run = lambda i: execute(i, env) return list( p.map(run, ideas) ) # one round of RL from execution reward def rl_round( policy, prompt, env, size ): rollouts = [ policy.sample( prompt, max_tokens=8192 ) for _ in range(size) ] # only the idea is scored, # not the thinking behind it ideas = [ extract_idea(r) for r in rollouts ] scores = execute_batch(ideas, env) # a crash earns 0: the # root of the collapse rewards = [ 0.0 if s is None else s for s in scores ] mu = statistics.mean(rewards) sd = ( statistics.pstdev(rewards) or 1.0 ) advs = [ (r - mu) / sd for r in rewards ] # only the ideator changes policy.update(rollouts, advs) # mu rises; watch max instead return mu, max(rewards)
This paper is an additional reading for Lecture 1, "Intro to Agentic Systems", of Stanford's CS 329Z: Engineering AI Agents, taught by Diyi Yang (one of this paper's authors), Michael Ryan and John Yang. Two ideas from that lecture fit it closely. The lecture separates workflows, where code fixes the steps, from agents, where the language model decides the steps. The executor is a workflow in exactly that sense: Implementer, then Scheduler, then Worker, always in that order, with language models doing the writing inside fixed steps. The lecture also splits agent engineering into the system, the data, and the evals with their metric, and names training, evaluation and safety as the key challenges. This paper touches all of them: an evaluation harness turned into a reward, two ways of training on it, and guards against reward hacking.
Now press Present or Teach and explain, out loud and from memory, why a zero reward for crashed ideas makes reinforcement learning give up on long shots while search does not. If you can, you own this paper. Then go back to the research loop and switch between search and reinforcement learning on both problems.