The Gate Does Not Choose: An Expert-Initialised Mixture-of-Experts Outperforms Training-Free Consensus under the Official BraTS-2023 Metrics
Given several trained segmentation models, a practitioner has two ways to combine them: a deterministic consensus operator that costs no training at all, or a learned gate that costs a full one. This paper measures both against each other on BraTS-2023 adult glioma, under the official lesion-wise Dice and HD95 and nothing else — and the learned gate wins, on one condition: it must be initialised from the experts it hosts. Twenty-nine arms are scored by one code path on the same 239 held-out patients of each of five disjoint folds: four dense experts that differ only by an auxiliary loss term (none, signed-distance regression, boundary loss, blob loss), twenty-four deterministic consensus operators built from them (ordered two-expert vetoes, one-against-three vetoes, symmetric votes), and two intra-model patch-wise mixtures-of-experts whose gates are trained inside the network. Four findings. (i) The best-scoring arm of the study is a learned gate: the V3 mixture-of-experts — four experts initialised from the last decoder block of the four already-trained specialists — gains +0.0057 lesion-wise Dice over its own baseline averaged across the five folds, is positive on five of five, and clears the pre-declared smallest effect of interest (0.003) on four. (ii) No free operator reaches that effect of interest: the best of the twenty-four consensus arms gains +0.0013, and paired directly against the V3 on the same patients all of them have a negative mean margin — 21 of them behind on five folds out of five, the three exceptions (each involving K) on four of five. The gate's margin over the best free operator is +0.0044, 1.5 times the effect of interest. (iii) The gate does not choose. Across the seven training runs the routing stays close to uniform — largest expert share 0.507 to 0.583 where 0.25 would be no preference at all on 4 experts, normalised patch entropy 0.892 to 0.995 where 1.0 is maximum, and no expert dead in any run. The gain therefore comes from *hosting* four loss specialists in one network and initialising them from their own checkpoints, not from learning to route between them; the title states the negative result that makes the positive one interpretable. (iv) The cost is stated rather than hidden: 300 epochs on top of four already-trained experts, against zero GPU-hours for a veto whose passes are independent and parallelise across GPUs — and the blob-loss expert the V3 hosts is the *worst* of the four alone (−0.0038, no fold positive) yet its veto on the baseline is the *best* two-expert veto (+0.0011). What does not change is the resolution of any single fold — every per-fold delta remains below its own minimum detectable effect — so the claim rests on a constant sign across five disjoint populations, not on a significant per-fold test. Twenty-nine arms are scored under the official BraTS-2023 lesion-wise metrics (clean regime, 1000/250/500 component filtering) on the same held-out patients of five disjoint folds: four dense MedNeXt/nnU-Net experts that differ only by an auxiliary loss term (baseline, signed-distance head, Kervadec boundary loss, blob loss), twenty-four training-free consensus operators built from them (vetoes and votes), and two learned patch-wise gates. The archive ships the EN and FR manuscripts (markdown + rendered PDF), the LaTeX sources are kept in the project repository, every published table, all ten figures, the measurement JSONs behind every number, and the reproduction code (official-metric evaluation, statistics, tables, figures). Measured verdicts shipped in data/v3_vs_consensus.json: gate V3 mean cross-validated delta to its own baseline +0.00566 Dice against a pre-declared smallest effect of interest of 0.003, best training-free operator (cc_B_and_DKG) +0.00127; the routing stays close to uniform (largest expert share 0.507-0.583 where 0.25 is no preference on four experts), so the gain comes from hosting four expert-initialised specialists, not from learned routing.
Authors
- Stanislas Larnier
- Guillaume Cassez (ORCID: https://orcid.org/0009-0007-0987-3931)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-15
- DOI
- https://doi.org/10.5281/zenodo.22776410
- Primary Topic
- Glioma Diagnosis and Treatment
- Type
- preprint