The Gate Still Does Not Choose: An Expert-Initialised Mixture-of-Experts Beats Its Matched Control and Is the Only All-Rounder Arm under Full-Resolution Cityscapes Metrics
We report a pre-registered, controlled evaluation of MoE-V3-CS, a patch-wise mixture-of-experts inserted into a full-resolution Cityscapes segmenter (ConvNeXt-V2-Base + UPerNet, 1024×2048, 19 classes, BF16). Its four experts are initialised bit-for-bit from the head.fpn_convs.0 block of four independently trained loss specialists of the same seed — B (CE+Dice), D (CE+Kervadec EDT), Dp (CE+SDT distance map), G (CE+Dice+0.5·blob) — and a top-2 gate over 3×3 patches is then trained for 80 epochs from a residual scale γ initialised to zero, so that epoch 0 is bit-exact to the control. The pre-registered primary endpoint is the dataset-level official mIoU (cityscapesScripts, 19 classes) on the shared 500-image val holdout, MoE-V3-CS against its recipe-matched control (same B initialisation, same recipe, 80 epochs, no MoE layer), paired image-bootstrap B = 10 000, three seeds averaged within each replicate. The primary endpoint is positive at the raw threshold and inside the pre-registered single-pair family: Δ = +0.449 pt (MoE 81.618 vs control 81.168), 95 % CI [+0.109 ; +0.820], two-sided p = 0.0066, Holm = 0.0066 over the 1-pair family. It does not survive the exploratory multiplicity families: Holm = 0.0726 over 12 pairs and Holm = 0.0924 over 15 pairs — and no arm of the plateau survives either (best 0.0750). All three families are reported together, none is chosen because it flatters. The central result is a routing diagnostic: at the final epoch the gate distribution is flat (minimal normalised entropy 0.99771 over the three seeds, part_max 0.5030 / 0.5019 / 0.5019 against 0.5000 for exact equidistribution, every expert between 49.79 % and 50.30 % of the patches, zero dead experts), yet γ grows from 0.00024 to ≈ 0.020 (×77 to ×84) — the mixture helps, but as an average of experts, not as a selection. The gain does not come from routing; it comes from the expert initialisation. This replicates on a second dataset, a second architecture and a second attach point the BRATS result The Gate Does Not Choose [DOI 10.5281/zenodo.22903668], of which this paper is a replication, not a republication: no BRATS number is reused. On the program's pre-registered versatility criterion (36 endpoints × 13 arms), MoE-V3-CS is the only all-rounder: maximal damage −0.53 pt (crowd-group instances [email protected]) against −1.90 (consensus C⊘B) and −22.20 (D, pedestrian pixel precision 70.0 → 47.8), a margin of 1.37 pt over the next arm, and the cheapest plateau of the field at −1.2 pt of damage per mIoU point against −43.1 for D. It is nonetheless first on none of the 36 endpoints (D wins 16) and its mean percentile is 60.4 (5th of 11): good everywhere is not best everywhere. Two business metrics survive the 15-pair Holm — fragments (−31.2 connected components per image) and instances_taille_T3_rappel (+0.26 pt, largest-instance recall) — while no per-class IoU rise survives the 19-class Holm (the only survivor is a fall, bicycle 0.0038). Contributions. (1) A pre-registered positive primary for an expert-initialised mixture against its recipe-matched control, reported with its three multiplicity families together and the explicit statement that it does not survive the exploratory one. (2) A routing diagnostic that separates a tautology from a proof: the equality entropy_token == entropy_token_clean at the final epoch holds by construction (the Shazeer noise is annealed to zero) and proves nothing; the probative comparison runs over the 40 active-noise epochs, where the clean distribution is already flat (≥ 0.99821) and part_max stays within [0.50011 ; 0.50217] of the 0.5000 equidistribution. (3) The correct naming of two routinely conflated routing quantities — sum(frac) == top_k (= 2), not 1, and top_expert_share (mean of per-batch maxima, 0.667 at epoch 0) ≠ part_max (maximum of per-batch means, 0.501). (4) A versatility verdict recomputed from raw deltas and ranks, with every rank carrying the denominator of its own endpoint (the MoE's worst rank is 11/12 on instances_foule_rappel, because arm A is in partial coverage and the denominator is therefore not constant). (5) A measured architectural difference with BRATS: the residual guard γ is indispensable here (without it the raw recipe costs −28.1 mIoU points at epoch 0) and absent there. (6) Public release of code, configs, tables, figures and regeneration scripts.
Authors
- Stanislas Larnier
- Guillaume Cassez (ORCID: https://orcid.org/0009-0007-0987-3931)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-05
- DOI
- https://doi.org/10.5281/zenodo.23090081
- Primary Topic
- Advanced Neural Network Applications
- Type
- preprint