The Gate Does Not Choose: An Expert-Initialised Mixture-of-Experts Outperforms Training-Free Consensus under the Official BraTS-2023 Metrics

Given several trained segmentation models, a practitioner has two ways to combine them: a deterministic consensus operator that costs no training at all, or a learned gate that costs a full one. This paper measures both against each other on BraTS-2023 adult glioma, under the official lesion-wise Dice and HD95 and nothing else — and the learned gate wins, on one condition: it must be initialised from the experts it hosts. Twenty-nine arms are scored by one code path on the same 239 held-out patients of each of five disjoint folds: four dense experts that differ only by an auxiliary loss term (none, signed-distance regression, boundary loss, blob loss), twenty-four deterministic consensus operators built from them (ordered two-expert vetoes, one-against-three vetoes, symmetric votes), and two intra-model patch-wise mixtures-of-experts whose gates are trained inside the network. Four findings. (i) The best-scoring arm of the study is a learned gate: the V3 mixture-of-experts — four experts initialised from the last decoder block of the four already-trained specialists — gains +0.0057 lesion-wise Dice over its own baseline averaged across the five folds, is positive on five of five, and clears the pre-declared smallest effect of interest (0.003) on four. (ii) No free operator reaches that effect of interest: the best of the twenty-four consensus arms gains +0.0013, and paired directly against the V3 on the same patients all of them have a negative mean margin — 21 of them behind on five folds out of five, the three exceptions (each involving K) on four of five. The gate's margin over the best free operator is +0.0044, 1.5 times the effect of interest. (iii) The gate does not choose. Across the seven training runs the routing stays close to uniform — largest expert share 0.507 to 0.583 where 0.25 would be no preference at all on 4 experts, normalised patch entropy 0.892 to 0.995 where 1.0 is maximum, and no expert dead in any run. The gain therefore comes from *hosting* four loss specialists in one network and initialising them from their own checkpoints, not from learning to route between them; the title states the negative result that makes the positive one interpretable. (iv) The cost is stated rather than hidden: 300 epochs on top of four already-trained experts, against zero GPU-hours for a veto whose passes are independent and parallelise across GPUs — and the blob-loss expert the V3 hosts is the *worst* of the four alone (−0.0038, no fold positive) yet its veto on the baseline is the *best* two-expert veto (+0.0011). What does not change is the resolution of any single fold — every per-fold delta remains below its own minimum detectable effect — so the claim rests on a constant sign across five disjoint populations, not on a significant per-fold test. Twenty-nine arms are scored under the official BraTS-2023 lesion-wise metrics (clean regime, 1000/250/500 component filtering) on the same held-out patients of five disjoint folds: four dense MedNeXt/nnU-Net experts that differ only by an auxiliary loss term (baseline, signed-distance head, Kervadec boundary loss, blob loss), twenty-four training-free consensus operators built from them (vetoes and votes), and two learned patch-wise gates. The archive ships the EN and FR manuscripts (markdown + rendered PDF), the LaTeX sources are kept in the project repository, every published table, all ten figures, the measurement JSONs behind every number, and the reproduction code (official-metric evaluation, statistics, tables, figures). Measured verdicts shipped in data/v3_vs_consensus.json: gate V3 mean cross-validated delta to its own baseline +0.00566 Dice against a pre-declared smallest effect of interest of 0.003, best training-free operator (cc_B_and_DKG) +0.00127; the routing stays close to uniform (largest expert share 0.507-0.583 where 0.25 is no preference on four experts), so the gain comes from hosting four expert-initialised specialists, not from learned routing.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-15
DOI
https://doi.org/10.5281/zenodo.22776412
Primary Topic
Glioma Diagnosis and Treatment
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

The Gate Does Not Choose: An Expert-Initialised Mixture-of-Experts Outperforms Training-Free Consensus under the Official BraTS-2023 Metrics

Stanislas Larnier, Guillaume Cassez
Zenodo (CERN European Organization for Nuclear Research)
Glioma Diagnosis and Treatment
preprint

The Gate Does Not Choose: An Expert-Initialised Mixture-of-Experts Outperforms Training-Free Consensus under the Official BraTS-2023 Metrics

Stanislas Larnier, Guillaume Cassez
preprint en

Abstract

Given several trained segmentation models, a practitioner has two ways to combine them: a deterministic consensus operator that costs no training at all, or a learned gate that costs a full one. This paper measures both against each other on BraTS-2023 adult glioma, under the official lesion-wise Dice and HD95 and nothing else — and the learned gate wins, on one condition: it must be initialised from the experts it hosts. Twenty-nine arms are scored by one code path on the same 239 held-out patients of each of five disjoint folds: four dense experts that differ only by an auxiliary loss term (none, signed-distance regression, boundary loss, blob loss), twenty-four deterministic consensus operators built from them (ordered two-expert vetoes, one-against-three vetoes, symmetric votes), and two intra-model patch-wise mixtures-of-experts whose gates are trained inside the network. Four findings. (i) The best-scoring arm of the study is a learned gate: the V3 mixture-of-experts — four experts initialised from the last decoder block of the four already-trained specialists — gains +0.0057 lesion-wise Dice over its own baseline averaged across the five folds, is positive on five of five, and clears the pre-declared smallest effect of interest (0.003) on four. (ii) No free operator reaches that effect of interest: the best of the twenty-four consensus arms gains +0.0013, and paired directly against the V3 on the same patients all of them have a negative mean margin — 21 of them behind on five folds out of five, the three exceptions (each involving K) on four of five. The gate's margin over the best free operator is +0.0044, 1.5 times the effect of interest. (iii) The gate does not choose. Across the seven training runs the routing stays close to uniform — largest expert share 0.507 to 0.583 where 0.25 would be no preference at all on 4 experts, normalised patch entropy 0.892 to 0.995 where 1.0 is maximum, and no expert dead in any run. The gain therefore comes from *hosting* four loss specialists in one network and initialising them from their own checkpoints, not from learning to route between them; the title states the negative result that makes the positive one interpretable. (iv) The cost is stated rather than hidden: 300 epochs on top of four already-trained experts, against zero GPU-hours for a veto whose passes are independent and parallelise across GPUs — and the blob-loss expert the V3 hosts is the *worst* of the four alone (−0.0038, no fold positive) yet its veto on the baseline is the *best* two-expert veto (+0.0011). What does not change is the resolution of any single fold — every per-fold delta remains below its own minimum detectable effect — so the claim rests on a constant sign across five disjoint populations, not on a significant per-fold test. Twenty-nine arms are scored under the official BraTS-2023 lesion-wise metrics (clean regime, 1000/250/500 component filtering) on the same held-out patients of five disjoint folds: four dense MedNeXt/nnU-Net experts that differ only by an auxiliary loss term (baseline, signed-distance head, Kervadec boundary loss, blob loss), twenty-four training-free consensus operators built from them (vetoes and votes), and two learned patch-wise gates. The archive ships the EN and FR manuscripts (markdown + rendered PDF), the LaTeX sources are kept in the project repository, every published table, all ten figures, the measurement JSONs behind every number, and the reproduction code (official-metric evaluation, statistics, tables, figures). Measured verdicts shipped in data/v3_vs_consensus.json: gate V3 mean cross-validated delta to its own baseline +0.00566 Dice against a pre-declared smallest effect of interest of 0.003, best training-free operator (cc_B_and_DKG) +0.00127; the routing stays close to uniform (largest expert share 0.507-0.583 where 0.25 is no preference on four experts), so the gain comes from hosting four expert-initialised specialists, not from learned routing.

Zenodo (CERN European Organization for Nuclear Research)
Peace, Justice and strong institutions
Glioma Diagnosis and Treatment
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.