The Gate Still Does Not Choose: An Expert-Initialised Mixture-of-Experts Beats Its Matched Control and Is the Only All-Rounder Arm under Full-Resolution Cityscapes Metrics
We report a pre-registered, controlled evaluation of **MoE-V3-CS**, a patch-wise **mixture-of-experts** inserted into a full-resolution Cityscapes segmenter (ConvNeXt-V2-Base + UPerNet, 1024×2048, 19 classes, BF16). Its four experts are **initialised bit-for-bit from the `head.fpn_convs.0` block of four independently trained loss specialists of the same seed** — B (CE+Dice), D (CE+Kervadec EDT), Dp (CE+SDT distance map), G (CE+Dice+0.5·blob) — and a top-2 gate over 3×3 patches is then trained for 80 epochs from a residual scale γ initialised to zero, so that epoch 0 is bit-exact to the control. The **pre-registered primary endpoint** is the dataset-level official mIoU (cityscapesScripts, 19 classes) on the shared 500-image val holdout, MoE-V3-CS against its **recipe-matched control** (same B initialisation, same recipe, 80 epochs, no MoE layer), paired image-bootstrap B = 10 000, three seeds averaged within each replicate. The primary endpoint is **positive at the raw threshold and inside the pre-registered single-pair family**: Δ = **+0.449 pt** (MoE 81.618 vs control 81.168), 95 % CI [+0.109 ; +0.820], two-sided p = **0.0066**, Holm = 0.0066 over the 1-pair family. It **does not survive the exploratory multiplicity families**: Holm = 0.0726 over 12 pairs and Holm = 0.0924 over 15 pairs — and **no arm of the plateau survives either** (best 0.0750). All three families are reported together, none is chosen because it flatters. The central result is a **routing diagnostic**: at the final epoch the gate distribution is flat (minimal normalised entropy 0.99771 over the three seeds, `part_max` 0.5030 / 0.5019 / 0.5019 against 0.5000 for exact equidistribution, every expert between 49.79 % and 50.30 % of the patches, **zero dead experts**), yet γ grows from 0.00024 to ≈ 0.020 (×77 to ×84) — the mixture **helps**, but as an **average of experts**, not as a selection. The gain does not come from routing; it comes from the expert initialisation. This **replicates on a second dataset, a second architecture and a second attach point** the BRATS result *The Gate Does Not Choose* [DOI 10.5281/zenodo.22903668], of which this paper is a replication, not a republication: no BRATS number is reused. On the program's pre-registered versatility criterion (36 endpoints × 13 arms), MoE-V3-CS is the **only all-rounder**: maximal damage **−0.53 pt** (crowd-group instances [email protected]) against −1.90 (consensus C⊘B) and −22.20 (D, pedestrian pixel precision 70.0 → 47.8), a margin of 1.37 pt over the next arm, and the cheapest plateau of the field at **−1.2 pt of damage per mIoU point** against −43.1 for D. It is nonetheless **first on none of the 36 endpoints** (D wins 16) and its mean percentile is 60.4 (5th of 11): *good everywhere is not best everywhere*. Two business metrics survive the 15-pair Holm — `fragments` (−31.2 connected components per image) and `instances_taille_T3_rappel` (+0.26 pt, largest-instance recall) — while no per-class IoU rise survives the 19-class Holm (the only survivor is a fall, bicycle 0.0038). **Contributions.** (1) A pre-registered **positive primary** for an expert-initialised mixture against its recipe-matched control, reported with its **three** multiplicity families together and the explicit statement that it does not survive the exploratory one. (2) A **routing diagnostic that separates a tautology from a proof**: the equality `entropy_token == entropy_token_clean` at the final epoch holds *by construction* (the Shazeer noise is annealed to zero) and proves nothing; the probative comparison runs over the 40 active-noise epochs, where the clean distribution is already flat (≥ 0.99821) and `part_max` stays within [0.50011 ; 0.50217] of the 0.5000 equidistribution. (3) The correct naming of two routinely conflated routing quantities — `sum(frac) == top_k` (= 2), not 1, and `top_expert_share` (mean of per-batch maxima, 0.667 at epoch 0) ≠ `part_max` (maximum of per-batch means, 0.501). (4) A **versatility verdict recomputed from raw deltas and ranks**, with every rank carrying the denominator of its own endpoint (the MoE's worst rank is 11/12 on `instances_foule_rappel`, because arm A is in partial coverage and the denominator is therefore not constant). (5) A measured **architectural difference with BRATS**: the residual guard γ is indispensable here (without it the raw recipe costs −28.1 mIoU points at epoch 0) and absent there. (6) Public release of code, configs, tables, figures and regeneration scripts. ---
Authors
- Stanislas Larnier
- Guillaume Cassez (ORCID: https://orcid.org/0009-0007-0987-3931)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-04
- DOI
- https://doi.org/10.5281/zenodo.23146586
- Primary Topic
- Advanced Neural Network Applications
- Type
- preprint