What a Model Orchestrator Can and Cannot Add
A reading of the four papers behind Sakana's Fugu (AB-MCTS, Trinity, the Conductor, the Fugu technical report). The orchestrators are small models trained against ground truth to choose which frontier model answers, whether to rewrite the question, and when to stop; at inference nothing checks the answer. On the papers' own tables they beat the best single model in their pool by one to nine points. Trinity and Fugu are routers and cannot answer what no model in the pool answers; Trinity, the only one that measures that ceiling, stays under it on every task. The Conductor and Fugu-Ultra build the answer from several models' work and on held-out tasks gain more than any router could, measured with every worker cut to short answers; neither reports its own ceiling. The June 2026 claim of parity with Fable 5 rests on provider-reported scores and a previous-generation pool; Fugu-Ultra trails Fable 5 by up to six points on SWE-Bench Pro, SciCode and Humanity's Last Exam. Reading note, nothing measured. Live version: https://alatip.github.io/orchestrator-can-and-cannot-add/
Authors
- Aziz Latipov (ORCID: https://orcid.org/0009-0003-8938-8223)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-09
- DOI
- https://doi.org/10.5281/zenodo.23258162
- Primary Topic
- Topic Modeling
- Type
- preprint