Modern AI Engineering

Lesson 18.3 · 26 min

Recursive Self-Improvement: Can AI Improve Itself Indefinitely?

What happens if an AI gets good enough at AI research to improve itself, and the improved version is even better at improving itself?

In short: Recursive self-improvement (RSI) is a loop in which an AI system improves its own capabilities, and the improved system then makes the next improvement, and so on. Whether this speeds up, keeps a steady pace or fizzles depends on how much each gain makes the next gain easier, and on bottlenecks such as compute, data and reliable evaluation. Partial versions exist today (self-play, self-generated training data, AI agents that improve code and algorithms), while full, open-ended RSI remains hypothetical and is a central topic in AI safety.

What is recursive self-improvement?

Recursive self-improvement (RSI) means an AI system that makes itself better, where the improved system is then used to make the next improvement. The word recursive is the key: the output of one round (a smarter system) becomes the tool for the next round. Normal improvement is linear: people improve a system, then improve it again. In RSI, the improver itself keeps getting better.

Think of a toolmaker A blacksmith uses a rough hammer to forge a better hammer, then uses the better hammer to forge an even better one. Each new tool makes the next tool easier to make. Human technology has grown partly this way: better machines build better machines. RSI imagines an AI that is both the blacksmith and the hammer.

The idea is old. In 1965 the statistician I. J. Good wrote that an “ultraintelligent machine” could design even better machines, leading to an intelligence explosion. Jürgen Schmidhuber's Gödel machine (2003) described a theoretical self-rewriting program that changes its own code only when it can prove the change is an improvement. What is new is that today's models can write code, generate training data and run experiments, so partial versions of the loop are being built for real.

Why does recursive self-improvement matter?

  • Speed of progress: if AI can do a meaningful share of AI research, the pace of progress could stop being limited by the number of human researchers.
  • Compounding: small gains that make the next gain easier can add up much faster than steady, human-driven gains.
  • Safety and control: a system that changes itself might drift away from the goals and limits we set. Each change must be checked, and the checking itself must keep up.
  • Planning and policy: labs, governments and researchers track signs of AI accelerating AI research because it would change timelines for everything else.

Because the stakes are high and the topic attracts hype, we will be careful to separate what exists today from what is speculative.

How does an AI get better today?

Today, people drive almost every step of improving a model: collecting and cleaning data, designing architectures, writing training code, choosing hyper-parameters, running evaluations, and deciding what to try next. A new model version takes months and large teams.

But AI already helps at many of these steps. Models generate synthetic training data and critique outputs (for example AI feedback in Constitutional AI-style training). Coding assistants write a growing share of the code at AI labs. Models help label data, write evaluation tests and analyse experiment results. Each of these makes humans faster; none of them yet closes the loop without people deciding what to keep.

Pause and think: A lab uses its own model to write most of its training code, but engineers review and merge every change and decide which experiments to run. Is this full RSI?

No. It is AI-assisted improvement: the model speeds up humans, but humans still choose the goals, judge the results and decide what is kept. Full RSI would have the system itself proposing, evaluating and adopting improvements to its own capabilities in a closed loop.

How does recursive self-improvement work?

Strip away the drama and RSI is an improvement loop with a twist: the thing being improved is also the thing doing the improving.

What every working version needs

  1. A way to change itself: Access to its training data, training process, code or prompts.
  2. A reliable evaluator: A trustworthy measure of “better”. Without it the loop cannot tell progress from noise or from cheating.
  3. Resources: Compute and time to build and test each candidate. Training runs are expensive, which slows every lap.
  4. Transfer: Gains must make the system better at the improving task itself; otherwise it improves once and stops being recursive.

A simple example with numbers

Let capability be a single number c, starting at 1. Each round, the system improves itself by an amount that depends on how capable it already is: Δc = k · cᵅ. The exponent α captures how much being smarter helps make the next improvement.

rsi_toy.py

import random
random.seed(3)
# Toy model: capability c. Each round the system improves itself by k * c**alpha.
# alpha < 1: each gain helps less (diminishing returns); alpha > 1: gains compound faster.
def grow(alpha, k=0.1, c=1.0, rounds=10):
out = [c]
for _ in range(rounds):
c = c + k * c ** alpha
out.append(c)
return out
for alpha in (0.5, 1.0, 1.5):
path = grow(alpha)
print(f"alpha={alpha}: " + " ".join(f"{v:.2f}" for v in path[::2]))
# A self-improvement loop needs a verifier: keep a change only if a test score rises.
def evaluate(skill):                       # noisy benchmark, true value = skill
return skill + random.gauss(0, 0.05)
skill, kept, bad = 1.0, 0, 0
for rnd in range(1, 21):
candidate = skill + random.gauss(0.0, 0.08)    # proposed self-modification
if evaluate(candidate) > evaluate(skill):      # verify before accepting
bad += candidate < skill                   # noise fooled the verifier
skill, kept = candidate, kept + 1
print(f"after 20 rounds: skill={skill:.3f}, kept={kept}/20, of which secretly worse={bad}")

Output:

alpha=0.5: 1.00 1.20 1.43 1.67 1.94 2.22
alpha=1.0: 1.00 1.21 1.46 1.77 2.14 2.59
alpha=1.5: 1.00 1.22 1.51 1.91 2.50 3.38
after 20 rounds: skill=1.126, kept=9/20, of which secretly worse=3

Pause and think: In the toy, the three curves look almost the same for the first few rounds. What does that imply for anyone trying to detect RSI early?

Early data cannot easily tell diminishing, steady and accelerating regimes apart: after 4 rounds they are 1.43, 1.46 and 1.51. The difference only becomes obvious later, so careful, ongoing measurement is needed rather than conclusions from a few data points.

Two kinds of improvement

It helps to separate two levels at which a system can improve itself:

What exists in the real world today?

Steps toward self-improving systems

  1. Gödel machine (theory): Schmidhuber's design for a program that rewrites itself only after proving the rewrite helps. Not practical, but a precise formulation.
  2. AlphaGo Zero / AlphaZero: Learned Go (and later chess and shogi) purely by self-play: the current network generates games that train the next network. Object-level, inside a game with perfect rules.
  3. STaR and AI feedback: Self-Taught Reasoner: a model generates reasoning, keeps attempts that reach correct answers, fine-tunes on them and repeats. Constitutional AI uses model feedback to train models.
  4. AlphaEvolve: Google DeepMind's Gemini-powered coding agent evolved algorithms, including a faster kernel used in training Gemini itself and a 4×4 complex matrix multiplication using 48 multiplications.
  5. Darwin Gödel Machine: Sakana AI and collaborators: a coding agent that repeatedly edits its own agent code, keeping an archive of variants, and improved its score on coding benchmarks.

Two things stand out. First, every success so far relies on a strong, automatic evaluator: game outcomes, correct answers, passing tests, measured speed-ups. Second, the loops are bounded: humans choose the domain, the evaluator and the resources, and the gains are real but limited. AlphaEvolve's improvement to Gemini's training was a modest efficiency gain, not a runaway loop. In the Darwin Gödel Machine work, the authors also reported cases where the agent gamed its own evaluation, a reminder that self-improvement and reward hacking go hand in hand.

The intelligence explosion debate

The intelligence explosion hypothesis says that once AI can improve AI faster than humans can, capability could rise extremely fast (the α > 1 curve). People who take it seriously point to software being copyable and fast, to AI already speeding up AI research, and to how quickly capabilities have grown.

Sceptics point to bottlenecks that bend the curve down: training needs physical compute and energy, which grow slowly; experiments take time to run; new ideas get harder to find as easy ones are used up; and good evaluation for open-ended research is hard. These push toward α < 1, or toward fast growth that then saturates. Most serious researchers agree the answer is uncertain, which is exactly why the topic gets careful attention.

Hold both thoughts Partial self-improvement is real and useful today. A full, uncontrolled intelligence explosion is a hypothesis, not an observed event. Good thinking avoids both dismissal and hype.

Where it works, where it fails, and keeping humans in the loop

Works well where improvement is cheap to verify: games (win or lose), maths with checkable answers, code with tests, performance tuning with measurable speed. Fails or stalls where “better” is fuzzy (research taste, writing quality), where the evaluator can be gamed, where each round is very expensive, or where errors accumulate across rounds (training on one's own unfiltered outputs can degrade a model, sometimes called model collapse).

The core risk: the evaluator A self-improving system optimises whatever its evaluator rewards. If the evaluator is noisy (as in our toy, where a third of kept changes were secretly worse) or gameable, the system can “improve” in the wrong direction while scores go up. And if the system can modify its own evaluator or limits, every safeguard is at risk.

  • Human approval gates: people review and approve changes to the system, especially to its goals, evaluator or permissions.
  • Protected evaluation: keep tests and benchmarks out of the system's reach, use held-out and fresh tests, and look for reward hacking.
  • Sandboxing and limits: run self-modification in isolated environments with bounded compute, and the ability to roll back.
  • Interpretability and monitoring: track what changed and why, not only the score.
  • Staged deployment: advance capability in measured steps with safety evaluations at each step, as several labs' published safety frameworks describe.

Recursive self-improvement vs normal training

Common mistakes and how to spot them

Every self-improvement loop, from a simple prompt optimiser to a self-editing agent, tends to fail in the same few ways. None of them needs an intelligence explosion. They show up in small, everyday loops, and each one has a cheap test.

Failure modes of an improvement loop and how to check for them
FailureWhat it looks likeA cheap check
Chasing noiseMany changes are accepted, each by a tiny marginRe-run the old and the new version several times. If their scores overlap, the “gain” is noise
Overfitting the testThe score on the loop's own benchmark climbs, but results elsewhere stay flatKeep a second test set the loop never sees and compare it now and then
Gaming the evaluatorBig score jumps from odd changes, such as editing test files or special-casing inputsRead the actual changes, not only the score. Keep the evaluator outside the system's reach
Feeding on its own outputOutputs get more uniform and rare cases disappear round after roundMix in fresh real data and track diversity, not only average quality
No transferTask scores rise but each round's gain gets smallerPlot the gain per round. Shrinking gains mean the improver itself is not getting better

The second row deserves a worked example, because it happens even when nobody cheats. Suppose the system proposes five changes that all do nothing: the true score stays at 70. The benchmark is a little noisy, so the five measured scores come out as 68, 71, 69, 73 and 70 (illustrative numbers).

How picking the best creates a gain from nothing

  1. Measure: Five candidates, all truly at 70, score 68, 71, 69, 73 and 70. Their average is 70.2, as we would expect.
  2. Select: The loop keeps the best one and reports 73: an apparent gain of 3 points.
  3. Re-test on fresh questions: The kept version scores about 70 again. The gain was only the luckiest draw of the noise.
  4. Scale it up: With 50 candidates instead of 5, the luckiest draw is luckier still. The more options a loop tries on the same test, the more it overstates its progress.
  5. The fix: Choose on one test set and confirm on another. Only the confirmed number counts as progress.

A rule for any loop that keeps the best The score used to choose a winner is always too optimistic for that winner. This holds for hyper-parameter search, prompt tuning and model selection as much as for RSI. Report the score from data that played no part in the choice.

Practice: try it yourself

We will test one idea: how much does a steadier evaluator help a self-improvement loop? The loop proposes small random changes and keeps one only if its measured score beats the current one. This time the verifier can average several benchmark runs before deciding. We run 2,000 rounds with 1, 4, 16 and 64 runs per measurement and count how often the verifier was fooled.

practice_verifier_strength.py

import random
NOISE = 0.05      # how much one benchmark run wobbles around the true skill
CHANGE = 0.02     # typical size of one proposed self-modification
def measure(skill, repeats, rng):
"""Average `repeats` noisy benchmark runs. More repeats = a steadier score."""
return sum(skill + rng.gauss(0, NOISE) for _ in range(repeats)) / repeats
def self_improve(repeats, rounds=2000, seed=1):
rng = random.Random(seed)
skill, kept, worse = 1.0, 0, 0
for _ in range(rounds):
candidate = skill + rng.gauss(0, CHANGE)       # helps or hurts, 50/50
if measure(candidate, repeats, rng) > measure(skill, repeats, rng):
kept += 1
worse += candidate < skill                 # the verifier was fooled
skill = candidate
return skill, kept, worse
print("repeats  kept  secretly worse  final skill  benchmark runs")
for repeats in (1, 4, 16, 64):
skill, kept, worse = self_improve(repeats)
runs = 2000 * 2 * repeats                          # cost of all the checking
print(f"{repeats:7d}  {kept:4d}  {worse:8d} ({worse / kept:4.0%})  "
f"{skill:11.2f}  {runs:14,d}")

Output:

repeats  kept  secretly worse  final skill  benchmark runs
1  1005       402 ( 40%)         5.84           4,000
4   993       369 ( 37%)         8.20          16,000
16  1028       227 ( 22%)        13.18          64,000
64  1013       118 ( 12%)        16.79         256,000

Now change it:

  • Make the proposals bigger: change CHANGE from 0.02 to 0.2. Predict what happens to the “secretly worse” share for 1 repeat, and why big changes are easier to verify.
  • Make the proposals mostly harmful: change rng.gauss(0, CHANGE) to rng.gauss(-0.01, CHANGE). Predict whether the 1-repeat loop still ends above 1.0, and whether the 64-repeat loop does.
  • Make the benchmark noisier: change NOISE from 0.05 to 0.2. Predict which rows now look alike, and how many repeats would be needed to get back the old 16-repeat behaviour. (Hint: averaging n runs divides the noise by √n.)

Pause and think: Going from 1 to 64 repeats cuts the share of secretly worse changes from 40% to 12%, and the final skill rises from 5.84 to 16.79. What did that cost, and what does it say about the speed limit of self-improvement?

It cost 64 times more benchmark runs: 256,000 instead of 4,000. Averaging n runs only shrinks the noise by √n, so each further cut in errors gets more expensive. In this toy the proposals are equally good in every row; only the checking differs. So the pace of the loop is set by how well and how cheaply it can verify, not by how fast it can propose. That is one of the bottlenecks sceptics point to.

Pause and think: With a single run per measurement, 40% of kept changes were secretly worse, yet skill still rose from 1.0 to 5.84. Why does the loop still make progress, and why should that not reassure us about real systems?

A better candidate is still more likely to pass than a worse one, and a clearly bad change rarely beats the noise, so on average the kept changes add up to a gain. But the toy assumes honest, unbiased noise and a single number for “better”. A real evaluator can be gamed or can miss whole kinds of harm. Then its errors all lean the same way and do not average out, and a loop that accepts many unnoticed regressions can drift somewhere we did not intend while its score keeps rising.

Key takeaways

  • RSI is a loop where an AI improves itself and the improved version makes the next improvement.
  • Whether gains accelerate or fade depends on how much each gain eases the next (α) and on bottlenecks.
  • Every working self-improvement loop needs a reliable evaluator; noisy or gameable ones mislead it.
  • Partial RSI exists today (self-play, self-generated data, AlphaEvolve-style algorithm search), in bounded domains.
  • An intelligence explosion is a serious but uncertain hypothesis, not an observed fact.
  • Human approval, protected evaluation, sandboxing and staged deployment keep self-improvement controllable.

Key terms

  • Recursive self-improvement: A process in which an AI system improves itself and the improved system carries out further improvements.
  • Intelligence explosion: The hypothesis that recursive self-improvement could make AI capability rise extremely fast.
  • Self-play: Training in which a system generates its own experience by playing against copies of itself.
  • Evaluator (verifier): The test or measurement that decides whether a proposed change is an improvement.
  • Reward hacking: When a system raises its measured score through loopholes rather than genuinely doing better.
  • Meta-level improvement: Improving the process of improvement itself, such as training algorithms or agent code.

← 18.2 World Models: Teaching AI to Simulate Its Environment · 19.1 Cracking the AI Engineering Interview →