Lesson 11.15 · 25 min
Sakana Fugu: Lessons from an Open-Source Agent Study
What if, instead of training one giant model to be best at everything, we trained a model whose only skill is knowing which other model to ask, and how to make them work together?
In short: Sakana Fugu is an orchestration model from Sakana AI: a trained language model that sits in front of a pool of other LLMs and decides who should handle each request. Fugu routes each query to a single best worker using a small "selection head", trained first with supervised fine-tuning and then polished with evolutionary strategies. Fugu-Ultra writes whole multi-agent workflows in natural language, is trained with reinforcement learning (GRPO), and controls what each agent can see so agents do not simply copy each other.
What is Sakana Fugu?
Sakana Fugu is a family of orchestration models from the Tokyo-based lab Sakana AI, described in the "Sakana Fugu Technical Report" (arXiv, June 2026) and offered through an API that looks like an ordinary chat-model endpoint. You send a request to Fugu; behind the scenes Fugu decides which models in its worker pool (frontier LLMs from several providers, plus others) should handle it, calls them, and returns the answer.
An orchestrator here is not hand-written code with if/else rules. It is itself a trained model. Sakana calls the overall idea collective intelligence: many models with different strengths, coordinated well, can beat any one of them alone.
Think of it like a hospital's triage nurse and case manager A triage nurse is not the best surgeon or the best cardiologist. Their expertise is knowing, within seconds, which specialist a patient needs. For complex cases, a case manager goes further: they organise a team, decide who examines first, who reviews whose notes, and who makes the final call. Fugu is the triage nurse; Fugu-Ultra is the case manager.
This is a young system and many details (the backbone model, the exact pool, training data) are not public. In this lesson we explain the mechanisms the report describes, at the level we can state with confidence, and we mark what is uncertain.
Why Fugu was needed
- No model is best at everything. One frontier model may lead on coding, another on science questions, a third on long-context reading. Strengths shift with every release.
- Users and developers must choose. Picking a model per task by hand is tedious, and hand-written routing rules ("if the prompt mentions Python, use model A") are brittle.
- Combining models is hard. Debate, review and voting can help, but designing who does what, for each kind of question, is a skill in itself.
- New models arrive constantly. A system that composes models through their APIs can add a new worker without retraining it, or even seeing its weights.
Fugu's bet is that coordination can be learned: train a model on many tasks where we can check the answer, and let it discover which worker, or which team arrangement, tends to succeed.
The big picture: what Fugu does
Collective intelligence
Collective intelligence is the idea that a group can be smarter than its best member, if the members are different from each other and their contributions are combined well. For LLMs this works through two effects:
- Selection: for each question, pick the member most likely to be right. If models are strong in different areas, picking per question beats always using one.
- Combination: let several members contribute, then check, debate or merge. This can catch errors that any single model would make, if their errors are not all the same.
Pause and think: If every worker had exactly the same strengths and weaknesses, how much would per-question selection help?
Almost nothing. Selection only pays off when workers differ; if every row of the matrix were identical, choosing among them could not beat always using one. Diversity in the pool is what an orchestrator exploits.
How Fugu picks the right model: the lightweight selection head
A naive LLM router would generate text like "Use worker B" and parse it. That costs decoding time. Fugu instead adds a selection head: a small extra layer that sits next to the normal language-model output head of its backbone model.
How a selection decision is made
- Read the request: The backbone transformer processes the request, producing a hidden state (a vector) for each token position.
- Take one hidden state: Per the report, a hidden state from an early position is used, so no long text needs to be generated first.
- Score every worker: The selection head maps that vector to one logit (a raw score) per worker in the pool. With 4 workers we get 4 numbers.
- Pick and dispatch: The highest-scoring worker is called immediately with the request. The orchestration overhead is a single forward pass.
This is why Fugu can be fast: the expensive part (the chosen worker's answer) is unavoidable, and the routing part adds very little.
Teaching Fugu who is best: SFT, then evolutionary strategies
Supervised fine-tuning (SFT) means training on examples with known targets. For Fugu, the targets come from experiments: on tasks with checkable answers (code with tests, maths with known results), every worker is tried several times. Each worker's success rate becomes a reward, and a softmax turns the rewards into a soft target: a probability distribution that puts most weight on the best workers but still gives some to good alternatives. The head is trained to make its predicted distribution close to that target, by minimising the KL divergence (a measure of how different two probability distributions are).
SFT has a gap: it imitates per-step rewards measured offline, while what we really care about is the end result, including multi-turn and interactive tasks where one routing choice affects later ones. So the report describes a second stage using an evolutionary strategy (ES): a black-box optimiser that perturbs the parameters, measures the real end-to-end score of each perturbed version, and moves towards the better ones. The specific algorithm named is sep-CMA-ES, a variant of CMA-ES (Covariance Matrix Adaptation Evolution Strategy) that keeps only a diagonal covariance so it scales to many parameters. ES needs no gradients, so the reward can be any score, even one produced by running full agent sessions.
selection_head_toy.py
# A toy "selection head": learn which worker model to call for each query (numpy).
import numpy as np
rng = np.random.default_rng(0)
# 4 workers x 3 domains (code, math, science): true success rates (illustrative)
P = np.array([[0.80, 0.55, 0.60], # worker A: strong coder
[0.60, 0.85, 0.65], # worker B: strong at math
[0.62, 0.60, 0.82], # worker C: strong at science
[0.70, 0.70, 0.70]]) # worker D: good all-rounder
def queries(n): # hidden state stand-in: domain one-hot + noise
d = rng.integers(0, 3, n)
return np.eye(3)[d] + 0.3 * rng.normal(size=(n, 3)), d
X, d = queries(600)
# SFT targets: try every worker 8 times per query, softmax the average reward
wins = rng.binomial(8, P[:, d].T) / 8 # (600 queries, 4 workers)
T = np.exp(wins / 0.1); T /= T.sum(1, keepdims=True)
W = np.zeros((3, 4)) # the selection head: one logit per worker
for _ in range(300): # minimise KL(T || softmax(XW))
Z = X @ W; Q = np.exp(Z - Z.max(1, keepdims=True)); Q /= Q.sum(1, keepdims=True)
W -= 0.5 * X.T @ (Q - T) / len(X) # gradient of cross-entropy w.r.t. W
def success(W, n=3000): # expected success if we follow the head
Xt, dt = queries(n)
return P[(Xt @ W).argmax(1), dt].mean()
print("always one model (worker D): ", round(P[3].mean(), 3))
print("oracle (best worker per query):", round(P.max(0).mean(), 3))
print("after SFT selection head: ", round(success(W), 3))
for _ in range(30): # evolution strategy: perturb, keep the best
cands = [W] + [W + 0.2 * rng.normal(size=W.shape) for _ in range(8)]
W = max(cands, key=lambda c: success(c, 1000))
print("after ES polishing: ", round(success(W), 3))Output:
always one model (worker D): 0.7 oracle (best worker per query): 0.823 after SFT selection head: 0.819 after ES polishing: 0.82
In this toy, SFT already gets the head to 81.9%, very close to the 82.3% oracle and far above the 70% of always using one model. ES adds almost nothing here, because the SFT targets already match our goal exactly. ES matters in the real system where offline per-step rewards do not fully reflect end-to-end success.
How Fugu-Ultra conducts an orchestra: the Conductor, GRPO and isolation
Fugu-Ultra builds on Sakana's earlier research system called the Conductor. Instead of outputting a choice, the Conductor writes a workflow in natural language. Each step in the workflow contains three things:
- A subtask: what this step should accomplish, in plain language (for example, "write a failing test that reproduces the bug").
- An assigned worker: which model in the pool performs it (possibly Fugu itself, recursively).
- An access list: which earlier steps' outputs this worker is allowed to see.
Because the workflow is free-form text, any coordination shape that can be described in words is possible: a pipeline, parallel attempts with a judge, a debate, a tree. The report adds adaptive agent memory for long, tool-heavy tasks: persistent shared memory across turns of a conversation, but isolation inside a workflow (next section).
How do we teach a model to write good workflows? There is no dataset of "correct workflows". So Fugu-Ultra is trained with reinforcement learning using GRPO (Group Relative Policy Optimization). For one training question, the model samples a group of different workflows; each is executed and scored (for example, a reward for a well-formed workflow and a reward for a correct final answer). Each workflow's advantage is how much better it did than the group average. Workflows that beat their siblings become more likely; worse ones, less likely. No separate value model is needed.
Pause and think: For one question, GRPO samples four workflows with rewards [1, 1, 1, 1]. What do the advantages look like and what does the model learn from this group?
All rewards equal the mean, so every advantage is 0 (implementations guard against dividing by a zero std). The group gives no learning signal: when every workflow succeeds (or every one fails), there is nothing to prefer. Useful questions are those where some workflows succeed and others fail.
Stopping the agents from copying each other. A team of models helps only if members bring independent ideas. If the second agent can read everything the first agent did, it tends to follow the same path, including the same mistakes. The team collapses into one opinion, and the extra calls are wasted.
Fugu-Ultra handles this with intra-workflow isolation. Inside a workflow, an agent sees only its own past actions plus the outputs that the access list explicitly grants. So the Conductor can ask three workers to solve a problem independently, then give a fourth worker access to all three answers to compare and decide. Isolation is the default; sharing is a deliberate choice written into the plan.
A lesson that applies to any multi-agent system Independent first, then combine. Whether you are building your own debate, review or voting system, do not let every agent see every other agent's draft before it has formed its own answer.
How well does Fugu perform?
Sakana reports state-of-the-art or near state-of-the-art results across coding, science and reasoning benchmarks, often above the individual frontier models in its pool. Some headline numbers it published for Fugu-Ultra:
Sakana also reported strong results for the cheaper Fugu on some tasks (for example around 60% on SciCode). Treat all of these carefully:
- Self-reported: published by the vendor, not yet widely reproduced.
- The pool is not fully disclosed per request, so it is hard to tell how much gain comes from orchestration versus simply having access to strong workers.
- Cost and latency: early hands-on reviews noted that Fugu-Ultra can spend many tokens and much time on orchestration for simple questions. The faster Fugu is meant for those.
- Moving target: as workers are updated, scores change.
The clever strategies Fugu discovered on its own
Because the Conductor can describe any workflow in words, and GRPO rewards whatever works, Fugu-Ultra was not told which strategies to use. The report describes patterns that emerged from training, such as:
- Domain-aware routing: sending questions to different workers depending on subject area, even within one benchmark.
- Debate and aggregation: several workers answer independently, possibly over multiple rounds, and a final worker weighs the answers. Useful for knowledge-heavy questions.
- Tree-shaped teams: splitting a question into branches explored by different workers, then merging.
- Build-and-debug: for coding, one worker writes, another tests or reviews, and the work is revised.
These are familiar human-designed patterns from the orchestration lesson. The interesting part is that a trained orchestrator chose when to use each one, per question.
Do not over-read "it discovered strategies" Emergent here means the strategies were not hard-coded, not that the model invented new science. RL rediscovers what works on its training tasks, and it can also overfit to benchmarks. Real-world gains depend on whether your tasks look like those it was trained on.
Worked example, step by step
The training section said that measured success rates are turned into a soft target with a softmax, and that the head is trained to reduce the KL divergence to that target. Let us do both by hand for one query. These are toy numbers of our own, chosen to be easy to follow; they are not from the report.
Suppose we tried three workers on one training query and measured success rates of 0.9, 0.7 and 0.5.
From success rates to a soft target
- Divide by the temperature: With temperature 0.1, the rates become 9, 7 and 5. A small temperature stretches the gaps between workers.
- Exponentiate: Only the gaps to the best worker matter: e⁰ = 1, e⁻² ≈ 0.135 and e⁻⁴ ≈ 0.018.
- Normalise: The sum is about 1.154. Dividing gives the soft target [0.867, 0.117, 0.016]. The best worker gets most of the weight, the second still gets some, the third almost none.
- Try a higher temperature: With temperature 0.5 the same rates give [0.472, 0.316, 0.212]: much flatter. The ranking is the same, but the target now says “all three are fairly close”.
Now the loss. KL divergence compares the target T with the head's prediction Q as ∑ Tᵢ · ln(Tᵢ / Qᵢ). It is 0 when they match and grows as they differ. Using the temperature 0.1 target:
| Head's prediction Q | What it means | KL(T ‖ Q) |
|---|---|---|
| [0.34, 0.33, 0.33] | An untrained head: no idea who is best | 0.642 |
| [0.60, 0.30, 0.10] | Right ranking, not confident enough | 0.180 |
| [0.85, 0.12, 0.03] | Close to the target | 0.004 |
Training pushes the head down this table. Notice that the middle row already picks the right worker: taking the top score gives worker 1. The remaining loss is about confidence, not about the choice. That is one reason a loss on per-query targets and the end-to-end success we care about are not the same thing, which is the gap the second training stage is there to close.
Practice: try it yourself
The earlier code trained a toy selection head, the Fugu side of the story. Here we explore the Fugu-Ultra side: a workflow made of steps, where each step has a subtask, an assigned worker and an access list. We write a small executor that enforces the access lists, and scripted workers that copy any earlier answer they are allowed to see. This is our own sketch for learning, not Sakana's implementation.
practice_access_lists.py
# Toy workflow executor with access lists (our own sketch, not Sakana's code).
from collections import Counter
ALONE = {"A": "41", "B": "42", "C": "42"} # each worker's own answer; 42 is correct
def solver(name, visible):
"""Scripted worker: if it can see an earlier answer, it follows the first one."""
return visible[0] if visible else ALONE[name]
def judge(name, visible):
"""Scripted judge: picks the most common answer among those it may see."""
return Counter(visible).most_common(1)[0][0]
def run(title, workflow):
outputs = []
for number, step in enumerate(workflow, start=1):
assert all(n < number for n in step["access"]) # only earlier steps
visible = [outputs[n - 1] for n in step["access"]] # isolation is enforced here
role = judge if step["task"] == "decide" else solver
outputs.append(role(step["worker"], visible))
print(f" step {number} {step['worker']} sees {visible} -> {outputs[-1]}")
print(f"{title}: final answer {outputs[-1]}, distinct opinions {len(set(outputs[:3]))}")
# Each step: a subtask, an assigned worker, and an access list of earlier steps.
shared = [{"task": "solve", "worker": "A", "access": []},
{"task": "solve", "worker": "B", "access": [1]},
{"task": "solve", "worker": "C", "access": [1, 2]},
{"task": "decide", "worker": "D", "access": [1, 2, 3]}]
isolated = [{"task": "solve", "worker": "A", "access": []},
{"task": "solve", "worker": "B", "access": []},
{"task": "solve", "worker": "C", "access": []},
{"task": "decide", "worker": "D", "access": [1, 2, 3]}]
run("everyone sees everything", shared)
run("independent, then combine", isolated)Output:
step 1 A sees [] -> 41 step 2 B sees ['41'] -> 41 step 3 C sees ['41', '41'] -> 41 step 4 D sees ['41', '41', '41'] -> 41 everyone sees everything: final answer 41, distinct opinions 1 step 1 A sees [] -> 41 step 2 B sees [] -> 42 step 3 C sees [] -> 42 step 4 D sees ['41', '42', '42'] -> 42 independent, then combine: final answer 42, distinct opinions 2
Now change it:
- In
shared, reorder the solvers so that B goes first and A second (keep the access lists as they are). Predict the final answer. Is the shared workflow now “good”, or just lucky? - In
isolated, change the judge's access list to[1]. Predict the final answer. What does this say about the combining step? - Make the pool less diverse: set
ALONE = {"A": "41", "B": "41", "C": "42"}. Predict the result of the isolated workflow. Does isolation help when most workers share the same mistake?
Pause and think: In the shared workflow, the judge saw three answers that all agreed. Why is that agreement worth less than it looks?
The three answers are not three opinions. B and C copied A, so the judge is really looking at one opinion repeated three times; the output even reports 1 distinct opinion. Agreement only counts as evidence when the answers were formed independently. In the isolated workflow the judge sees a real 2-to-1 split and can outvote the one wrong worker.
Pause and think: Using the GRPO idea from the lesson: suppose these two workflows were sampled in one group for this question, with reward 1 for a correct final answer and 0 otherwise. Which one gets the positive advantage, and what would the orchestrator slowly learn?
The rewards are [0, 1] for [shared, isolated]. The mean is 0.5, so the isolated workflow is above average and gets the positive advantage, and the shared one gets the negative. Over many such questions, workflows that keep solvers apart and combine afterwards would become more likely. Nobody has to hard-code “isolate first”; it is favoured because it scores better than its siblings. (This holds in our toy, where workers copy what they see.)
Quick summary
Sakana Fugu is a trained orchestration model that coordinates a pool of other LLMs. Fugu picks one worker per request using a light selection head on the backbone's hidden state, trained with SFT on soft targets from measured worker success and polished with sep-CMA-ES on end-to-end results. Fugu-Ultra builds on the Conductor: it writes natural-language workflows of subtasks, workers and access lists, and is trained with GRPO, which rewards workflows that beat their siblings. Intra-workflow isolation keeps agents independent so the team does not collapse into one opinion. Reported benchmark results are strong but self-reported; the pool and training details are only partly public.
Key takeaways
- Fugu is a trained orchestration model: its skill is choosing and coordinating other LLMs.
- Fugu routes each request to one worker via a light selection head on the backbone's hidden state.
- Training: SFT towards soft targets from measured worker success, then sep-CMA-ES on end-to-end results.
- Fugu-Ultra writes natural-language workflows (subtask, worker, access list) and is trained with GRPO.
- Isolation inside a workflow keeps agents independent, so teamwork adds diversity instead of copies.
- Reported scores are strong but vendor-reported; details of the pool and training are only partly public.
Key terms
- Orchestration model: A model trained to decide which other models to call and how to combine them.
- Worker pool: The set of LLMs an orchestrator can delegate to.
- Selection head: A small layer that maps the backbone's hidden state to one score per worker.
- Soft target: A probability distribution used as a training label instead of a single correct class.
- sep-CMA-ES: A scalable evolution strategy that searches parameters using only end-to-end scores, no gradients.
- GRPO: Group Relative Policy Optimization: RL that scores each sample relative to the average of its group.
- Intra-workflow isolation: Letting each agent see only its own work plus outputs explicitly granted by the access list.
← 11.14 AI Orchestration: Coordinating Agents, Tools, and Flows · 11.16 Computer-Use Agents: Controlling Interfaces with AI →