What a Model Orchestrator Can and Cannot Add

A reading of the four papers behind Sakana's Fugu (AB-MCTS, Trinity, the Conductor, the Fugu technical report). The orchestrators are small models trained against ground truth to choose which frontier model answers, whether to rewrite the question, and when to stop; at inference nothing checks the answer. On the papers' own tables they beat the best single model in their pool by one to nine points. Trinity and Fugu are routers and cannot answer what no model in the pool answers; Trinity, the only one that measures that ceiling, stays under it on every task. The Conductor and Fugu-Ultra build the answer from several models' work and on held-out tasks gain more than any router could, measured with every worker cut to short answers; neither reports its own ceiling. The June 2026 claim of parity with Fable 5 rests on provider-reported scores and a previous-generation pool; Fugu-Ultra trails Fable 5 by up to six points on SWE-Bench Pro, SciCode and Humanity's Last Exam. Reading note, nothing measured. Live version: https://alatip.github.io/orchestrator-can-and-cannot-add/

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-09
DOI
https://doi.org/10.5281/zenodo.23258162
Primary Topic
Topic Modeling
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

What a Model Orchestrator Can and Cannot Add

Aziz Latipov
Zenodo (CERN European Organization for Nuclear Research)
Topic Modeling
preprint

What a Model Orchestrator Can and Cannot Add

Aziz Latipov
preprint en

Abstract

A reading of the four papers behind Sakana's Fugu (AB-MCTS, Trinity, the Conductor, the Fugu technical report). The orchestrators are small models trained against ground truth to choose which frontier model answers, whether to rewrite the question, and when to stop; at inference nothing checks the answer. On the papers' own tables they beat the best single model in their pool by one to nine points. Trinity and Fugu are routers and cannot answer what no model in the pool answers; Trinity, the only one that measures that ceiling, stays under it on every task. The Conductor and Fugu-Ultra build the answer from several models' work and on held-out tasks gain more than any router could, measured with every worker cut to short answers; neither reports its own ceiling. The June 2026 claim of parity with Fable 5 rests on provider-reported scores and a previous-generation pool; Fugu-Ultra trails Fable 5 by up to six points on SWE-Bench Pro, SciCode and Humanity's Last Exam. Reading note, nothing measured. Live version: https://alatip.github.io/orchestrator-can-and-cannot-add/

Zenodo (CERN European Organization for Nuclear Research)
Topic Modeling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

What a Model Orchestrator Can and Cannot Add — Aziz Latipov · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS