Modern AI Engineering

Lesson 7.3 · 26 min

Recursive Language Models: Self-Referential Generation

What if, instead of stuffing a 10-million-token document into a model's prompt, we let the model write code to explore it and call itself on the pieces?

In short: A Recursive Language Model (RLM) is an inference strategy, not a new network: the long input is stored as a variable in a programming environment (a Python REPL), and the model writes code to peek at it, search it, split it and call language models, including itself, on chosen pieces. Only small results ever enter the model's own context. This sidesteps context limits and "context rot", and it handles inputs far beyond any context window, at the cost of more calls, latency and reliance on the model's coding skill.

What is a Recursive Language Model (RLM)?

Recursive Language Models were proposed by Alex L. Zhang, Tim Kraska and Omar Khattab at MIT (a blog post in October 2025, then a paper on arXiv in December 2025). An RLM is a way of running an existing LLM. From the outside it looks like a normal model call: text in, answer out. Inside, the model never sees the whole input at once.

Instead, the input is loaded into a REPL (Read-Eval-Print Loop: an interactive programming session, like a Python console) as a variable, for example context. The model is told the variable exists and how big it is. It then writes code to look at parts of it, filter it, split it, and call a language model on selected pieces. Those sub-calls are where the recursion comes from: the model can invoke a model (even a copy of itself) on a sub-problem.

Think of it like a lawyer with a document room A lawyer handed 50 boxes of documents does not read every page into memory before answering. They check the index, search for key names, hand boxes to assistants with specific questions ("find every invoice from March"), and combine the assistants' short reports. The lawyer's own head only holds the summaries. An RLM is the lawyer; the REPL is the document room; sub-calls are the assistants.

Why do we need RLMs?

  • Hard context limits. Every model has a maximum context window. Inputs such as whole codebases, years of logs or thousands of documents can exceed even million-token windows.
  • Context rot. Even within the window, quality tends to drop as the context grows: models miss details, mix up facts, or ignore the middle. Benchmarks that need the model to use all of a long input (count, compare, aggregate) degrade especially fast.
  • Cost. Every token in the prompt is paid for on every call. Re-sending a huge context repeatedly is expensive.

Running example: a support team exports 50,000 lines of chat logs (about 1.4 million characters) and asks, "Which customer asked for a refund most often?" No single pass of an LLM over the raw text would reliably count hundreds of scattered mentions. We will solve it RLM-style.

How an RLM works

One RLM run, step by step

  1. Load: The 1.4M-character log goes into the REPL variable context. The root model's prompt stays small.
  2. Peek: The model prints the first two lines to learn the format: [00000] user=chen msg=hello.
  3. Plan: It decides to split the log into 5 chunks of 10,000 lines each.
  4. Delegate: For each chunk, it calls a sub-LM: "List the users who asked for a refund in this text." Each call sees only its chunk.
  5. Aggregate in code: It counts the returned names with ordinary Python, exactly and cheaply.
  6. Answer: It returns the final answer. In the authors' setup the model signals this with a special final-answer marker.

How the model writes and runs code

The root model behaves like a programmer at a console. Each turn it emits a code block; the environment runs it and returns the printed output, cut to a limited length so the model's context does not fill up. Variables persist between turns, so the model can build up intermediate results (lists of matches, partial summaries) outside its own context window.

The script below simulates this. The "sub-LM" is a stand-in function (a real RLM would call an LLM API there), but the structure, a variable holding the huge input, peeking, chunking, sub-calls and aggregation in code, is exactly the RLM pattern.

rlm_simulation.py

import random
random.seed(7)
# The huge input lives in a REPL variable, NOT in the model's prompt.
users = ["ana", "bo", "chen", "dev", "eli"]
context = "\n".join(
f"[{i:05d}] user={random.choice(users)} msg={'refund please' if random.random() < 0.03 else 'hello'}"
for i in range(50_000))
print("context size:", len(context), "chars,", context.count("\n") + 1, "lines")
calls = {"sub_llm": 0}
def llm_query(prompt, chunk):
# Stand-in for a sub-LM call on a small chunk; a real RLM would call an LLM here.
calls["sub_llm"] += 1
return [line.split()[1][5:] for line in chunk.splitlines() if "refund" in line]
# --- Code the root model might write, step by step ---
print("peek:", context[:200].splitlines()[:2])       # 1. look at the format
lines = context.splitlines()
chunks = ["\n".join(lines[i:i + 10_000]) for i in range(0, len(lines), 10_000)]  # 2. split
found = []
for ch in chunks:                                     # 3. recurse on each chunk
found += llm_query("Which users asked for a refund?", ch)
counts = {u: found.count(u) for u in sorted(set(found))}   # 4. aggregate in code
print("sub-LM calls:", calls["sub_llm"], "| refund requests:", len(found))
print("per user:", counts)
print("FINAL:", max(counts, key=counts.get), "asked most often")

Output:

context size: 1362328 chars, 50000 lines
peek: ['[00000] user=chen msg=hello', '[00001] user=dev msg=hello']
sub-LM calls: 5 | refund requests: 1527
per user: {'ana': 280, 'bo': 329, 'chen': 309, 'dev': 293, 'eli': 316}
FINAL: bo asked most often

Run model-written code in a sandbox An RLM executes code the model wrote. Always run it in an isolated sandbox with no secrets, limited network access, and limits on time, memory and the number of sub-calls. Never give the REPL access to production systems.

Why RLMs work better

  • Small, focused contexts. Each model call sees only what it needs, so it stays in the range where models are reliable, avoiding context rot.
  • Code does what code is good at. Exact searching, counting, sorting and joining are done by Python, not by a model guessing over a huge prompt.
  • The model chooses the strategy. Unlike a fixed pipeline, the model can inspect the data first and pick regex search, chunking or sampling as needed.
  • Inputs beyond the window. Since the full input never enters any prompt, its size is limited by the REPL's memory, not the context window.

The authors reported that on OOLONG, a long-context benchmark that requires aggregating information spread across the input, an RLM built on GPT-5-mini scored far higher than GPT-5 called directly (roughly double the correct answers on the 132K-token setting), at comparable cost per query. They also reported strong results when scaling to inputs of 10 million tokens and more (on BrowseComp-Plus with 1,000 documents), where direct calls are impossible. These are results from the authors' experiments; independent replications vary by task and model.

Recursion inside RLMs

Recursion means a procedure calling itself on a smaller version of the problem. In an RLM, a sub-call can be a plain LLM call, or another full RLM with its own REPL that can split its chunk further. The recursion depth is how many levels deep this goes.

In the original experiments the depth was 1: the root model called ordinary LMs on pieces, and those did not recurse further. Deeper recursion is possible in principle (a chunk that is still too big could be split again), but it multiplies calls and latency, so it needs limits.

Pause and think: A root model splits a 10M-token input into 100 chunks of 100K tokens, and each sub-call is itself an RLM that splits its chunk into 10 pieces. How many leaf LM calls happen, and what is the recursion depth?

100 × 10 = 1,000 leaf calls, plus the 100 intermediate RLM calls and the root. The depth is 2. This shows why depth and fan-out must be capped: calls grow multiplicatively.

How RLMs differ from simple chunking

Classic map-reduce chunking also splits a document and summarises each chunk, so what is new? In simple chunking a human designs a fixed pipeline in advance: split every N tokens, apply the same prompt to each chunk, merge summaries. The model has no say.

Advantages and limitations of RLMs

AdvantagesLimitations
Handles inputs far larger than any context windowMany sequential calls: answers can take seconds to minutes
Avoids context rot by keeping each call smallCost varies a lot by question; occasional runaway runs
Uses code for exact operations (count, sort, join)Only as good as the model's coding and planning skill
Works with existing models; no retraining neededSub-calls in the original design are blocking and do not reuse prefix caches
Model picks a strategy per questionRequires a secure sandbox for model-written code

Common mistake Using an RLM for a short input or a simple question. If the whole input fits comfortably in the context window and the question is easy, one direct call is faster, cheaper and just as accurate. RLMs pay off when inputs are huge or questions need exhaustive aggregation.

When to use RLMs

  • The input is larger than the context window, or so long that quality visibly drops.
  • The question needs information from all over the input: counting, comparing, listing every case, joining across documents.
  • The input has structure code can exploit: logs, CSVs, codebases, JSON, many separate documents.
  • Latency is acceptable (batch analysis, research, audits), not instant chat.

RLM vs RAG

Retrieval-Augmented Generation (RAG) embeds documents ahead of time, retrieves the top few chunks most similar to the question, and puts them into the prompt. It is fast and cheap, but it assumes the answer lives in a few retrievable chunks.

Pause and think: For "Which customer asked for a refund most often?" over 50,000 log lines, why would RAG struggle?

RAG returns only the top few chunks most similar to the question, but the answer requires counting every refund mention across the entire log. Most mentions would never be retrieved, so the count would be wrong. Exhaustive processing, as in an RLM, is needed.

A real use case

Auditing a large codebase A security team asks: "List every place where user input reaches a SQL query without parameterisation." The repository is millions of tokens. An RLM loads the file tree into the REPL, greps for database calls, groups hits by file, sends each group to a sub-LM with the question "Is the query built from unsanitised input? Answer yes/no with the line", and collects the yes-answers into a report in code. Every file is considered; each model call only sees a few hundred lines.

Other natural fits: analysing months of server logs during an incident review, comparing clauses across hundreds of contracts, and answering research questions over large document collections where the answer is spread across many sources.

Common mistakes and how to spot them

An RLM run can finish without any error and still return a wrong answer. Most such failures come from how the input was split or how the pieces were put back together. Here are the ones we meet most often and how to catch each.

RLM failure cases and their fixes
What we seeCauseFix
Counts come out slightly too lowA chunk boundary cut a record in half, so neither side matched itSplit on natural boundaries such as lines, files or records
Counts come out too highChunks overlap and the same record was counted twiceRemove duplicates by a record id before adding up
The root model answers from a tiny sampleThe printout was cut short and the model took it for the whole resultPrint lengths and counts first; keep full results in variables, not in printouts
The aggregation code crashes or skips itemsSub-calls answered in free text that code cannot parseAsk sub-calls for a strict format, such as one name per line
The run never seems to endChunks that are still too big keep being split againCap the depth and the total number of calls, and stop with a clear message

A cheap test catches the first two: run the RLM on a small input where we can compute the exact answer in ordinary code, and compare. Our practice script below does exactly that.

Estimating the number of calls before we run

  1. Start from the sizes: Say the input has 26,715 characters, one call may read 1,000, and each split makes 4 parts (toy numbers).
  2. Divide until it fits: 26,715 ÷ 4 ≈ 6,679, still too big. ÷ 4 again ≈ 1,670, too big. ÷ 4 again ≈ 417, which fits. That is 3 levels of splitting, so the depth is 3.
  3. Count the leaves: Each level multiplies the pieces by 4: 4 × 4 × 4 = 64 leaf calls.
  4. Count the inner calls: 1 at the root, 4 below it and 16 below those: 21 calls that only split and combine.
  5. Decide: 85 calls in total. If that is too many, raise the amount each call may read or search first and skip chunks with no hits.

Practice: try it yourself

The earlier script used recursion of depth 1. Here we will write a version that really calls itself: any piece that is still too big for one “model call” is split again. We count the calls at each level and check the answer against an exact count.

practice_recursive_split.py

WINDOW = 1_000          # most characters one "model call" may read (toy limit)
calls = {"leaf": 0, "inner": 0, "max_depth": 0}
def leaf_llm(chunk):
# stand-in for a sub-LM call: count the word ERROR in one small chunk
calls["leaf"] += 1
return chunk.count("ERROR")
def rlm(text, depth=0, fanout=4):
calls["max_depth"] = max(calls["max_depth"], depth)
if len(text) <= WINDOW:                  # small enough: answer directly
return leaf_llm(text)
calls["inner"] += 1                      # too big: split on lines and recurse
lines = text.splitlines(keepends=True)
size = -(-len(lines) // fanout)          # ceiling division
parts = ["".join(lines[i:i + size]) for i in range(0, len(lines), size)]
return sum(rlm(p, depth + 1, fanout) for p in parts)   # aggregate in code
log = "".join(f"line {i:04d} {'ERROR disk full' if i % 37 == 0 else 'ok'}\n"
for i in range(2000))
print("input chars:", len(log), "| window:", WINDOW)
print("ERROR count from RLM:", rlm(log), "| exact:", log.count("ERROR"))
print(calls)

Output:

input chars: 26715 | window: 1000
ERROR count from RLM: 55 | exact: 55
{'leaf': 64, 'inner': 21, 'max_depth': 3}

Now change it:

  • Set fanout=2 in the function definition. Predict the depth and the number of leaf calls before running.
  • Set WINDOW = 10_000. Predict how many leaf and inner calls are needed now.
  • Replace the line split with a character split: size = -(-len(text) // fanout) and parts = [text[i:i + size] for i in range(0, len(text), size)]. Predict whether the RLM count can still be trusted, then compare it with the exact count.

Pause and think: The run made 64 leaf calls. If each were a real model call taking 2 seconds, run one after another, how long would the leaves take, and what are two ways to cut that time?

64 × 2 = 128 seconds. We could run the sub-calls in parallel, since the chunks do not depend on each other. Or we could make fewer calls: let each call read more, or search for the word in code first and only send chunks that contain a hit.

Pause and think: Why does the function split on line boundaries instead of every N characters?

A cut at an arbitrary character can land inside a record, for example between “ERR” and “OR”. Then neither piece contains the full word and the record is silently missed. Splitting where records end keeps every record whole, so the parts can be counted independently and added up.

Key takeaways

  • An RLM keeps the long input in a REPL variable and lets the model explore it with code.
  • Sub-calls on chosen pieces (recursion) keep every model call small, avoiding context rot.
  • Exact operations such as counting and joining are done in code, not by the model guessing.
  • RLMs handle inputs far beyond the context window but are slower and less predictable in cost.
  • Use RAG for look-ups, RLMs for exhaustive analysis over huge inputs; sandbox all model-written code.

Key terms

  • Recursive Language Model (RLM): An inference strategy where a model explores a long input stored in a REPL and calls models on pieces of it.
  • REPL: Read-Eval-Print Loop: an interactive programming session that runs code and returns output.
  • Root model: The top-level model call that plans, writes code and returns the final answer.
  • Sub-call: A language-model call made from within the REPL on a selected piece of the input.
  • Recursion depth: How many levels of nested model calls are allowed below the root.
  • Context rot: The drop in model accuracy as the amount of text in its context grows.

← 7.2 Large Reasoning Models: Chain-of-Thought at Inference Time · 7.4 Diffusion Language Models: Text Generation Beyond Autoregression →