The Gate Still Does Not Choose: An Expert-Initialised Mixture-of-Experts Beats Its Matched Control and Is the Only All-Rounder Arm under Full-Resolution Cityscapes Metrics

We report a pre-registered, controlled evaluation of MoE-V3-CS, a patch-wise mixture-of-experts inserted into a full-resolution Cityscapes segmenter (ConvNeXt-V2-Base + UPerNet, 1024×2048, 19 classes, BF16). Its four experts are initialised bit-for-bit from the head.fpn_convs.0 block of four independently trained loss specialists of the same seed — B (CE+Dice), D (CE+Kervadec EDT), Dp (CE+SDT distance map), G (CE+Dice+0.5·blob) — and a top-2 gate over 3×3 patches is then trained for 80 epochs from a residual scale γ initialised to zero, so that epoch 0 is bit-exact to the control. The pre-registered primary endpoint is the dataset-level official mIoU (cityscapesScripts, 19 classes) on the shared 500-image val holdout, MoE-V3-CS against its recipe-matched control (same B initialisation, same recipe, 80 epochs, no MoE layer), paired image-bootstrap B = 10 000, three seeds averaged within each replicate. The primary endpoint is positive at the raw threshold and inside the pre-registered single-pair family: Δ = +0.449 pt (MoE 81.618 vs control 81.168), 95 % CI [+0.109 ; +0.820], two-sided p = 0.0066, Holm = 0.0066 over the 1-pair family. It does not survive the exploratory multiplicity families: Holm = 0.0726 over 12 pairs and Holm = 0.0924 over 15 pairs — and no arm of the plateau survives either (best 0.0750). All three families are reported together, none is chosen because it flatters. The central result is a routing diagnostic: at the final epoch the gate distribution is flat (minimal normalised entropy 0.99771 over the three seeds, part_max 0.5030 / 0.5019 / 0.5019 against 0.5000 for exact equidistribution, every expert between 49.79 % and 50.30 % of the patches, zero dead experts), yet γ grows from 0.00024 to ≈ 0.020 (×77 to ×84) — the mixture helps, but as an average of experts, not as a selection. The gain does not come from routing; it comes from the expert initialisation. This replicates on a second dataset, a second architecture and a second attach point the BRATS result The Gate Does Not Choose [DOI 10.5281/zenodo.22903668], of which this paper is a replication, not a republication: no BRATS number is reused. On the program's pre-registered versatility criterion (36 endpoints × 13 arms), MoE-V3-CS is the only all-rounder: maximal damage −0.53 pt (crowd-group instances [email protected]) against −1.90 (consensus C⊘B) and −22.20 (D, pedestrian pixel precision 70.0 → 47.8), a margin of 1.37 pt over the next arm, and the cheapest plateau of the field at −1.2 pt of damage per mIoU point against −43.1 for D. It is nonetheless first on none of the 36 endpoints (D wins 16) and its mean percentile is 60.4 (5th of 11): good everywhere is not best everywhere. Two business metrics survive the 15-pair Holm — fragments (−31.2 connected components per image) and instances_taille_T3_rappel (+0.26 pt, largest-instance recall) — while no per-class IoU rise survives the 19-class Holm (the only survivor is a fall, bicycle 0.0038). Contributions. (1) A pre-registered positive primary for an expert-initialised mixture against its recipe-matched control, reported with its three multiplicity families together and the explicit statement that it does not survive the exploratory one. (2) A routing diagnostic that separates a tautology from a proof: the equality entropy_token == entropy_token_clean at the final epoch holds by construction (the Shazeer noise is annealed to zero) and proves nothing; the probative comparison runs over the 40 active-noise epochs, where the clean distribution is already flat (≥ 0.99821) and part_max stays within [0.50011 ; 0.50217] of the 0.5000 equidistribution. (3) The correct naming of two routinely conflated routing quantities — sum(frac) == top_k (= 2), not 1, and top_expert_share (mean of per-batch maxima, 0.667 at epoch 0) ≠ part_max (maximum of per-batch means, 0.501). (4) A versatility verdict recomputed from raw deltas and ranks, with every rank carrying the denominator of its own endpoint (the MoE's worst rank is 11/12 on instances_foule_rappel, because arm A is in partial coverage and the denominator is therefore not constant). (5) A measured architectural difference with BRATS: the residual guard γ is indispensable here (without it the raw recipe costs −28.1 mIoU points at epoch 0) and absent there. (6) Public release of code, configs, tables, figures and regeneration scripts.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-05
DOI
https://doi.org/10.5281/zenodo.23147252
Primary Topic
Advanced Neural Network Applications
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

The Gate Still Does Not Choose: An Expert-Initialised Mixture-of-Experts Beats Its Matched Control and Is the Only All-Rounder Arm under Full-Resolution Cityscapes Metrics

Stanislas Larnier, Guillaume Cassez
Zenodo (CERN European Organization for Nuclear Research)
Advanced Neural Network Applications
preprint

The Gate Still Does Not Choose: An Expert-Initialised Mixture-of-Experts Beats Its Matched Control and Is the Only All-Rounder Arm under Full-Resolution Cityscapes Metrics

Stanislas Larnier, Guillaume Cassez
preprint en

Abstract

We report a pre-registered, controlled evaluation of MoE-V3-CS, a patch-wise mixture-of-experts inserted into a full-resolution Cityscapes segmenter (ConvNeXt-V2-Base + UPerNet, 1024×2048, 19 classes, BF16). Its four experts are initialised bit-for-bit from the head.fpn_convs.0 block of four independently trained loss specialists of the same seed — B (CE+Dice), D (CE+Kervadec EDT), Dp (CE+SDT distance map), G (CE+Dice+0.5·blob) — and a top-2 gate over 3×3 patches is then trained for 80 epochs from a residual scale γ initialised to zero, so that epoch 0 is bit-exact to the control. The pre-registered primary endpoint is the dataset-level official mIoU (cityscapesScripts, 19 classes) on the shared 500-image val holdout, MoE-V3-CS against its recipe-matched control (same B initialisation, same recipe, 80 epochs, no MoE layer), paired image-bootstrap B = 10 000, three seeds averaged within each replicate. The primary endpoint is positive at the raw threshold and inside the pre-registered single-pair family: Δ = +0.449 pt (MoE 81.618 vs control 81.168), 95 % CI [+0.109 ; +0.820], two-sided p = 0.0066, Holm = 0.0066 over the 1-pair family. It does not survive the exploratory multiplicity families: Holm = 0.0726 over 12 pairs and Holm = 0.0924 over 15 pairs — and no arm of the plateau survives either (best 0.0750). All three families are reported together, none is chosen because it flatters. The central result is a routing diagnostic: at the final epoch the gate distribution is flat (minimal normalised entropy 0.99771 over the three seeds, part_max 0.5030 / 0.5019 / 0.5019 against 0.5000 for exact equidistribution, every expert between 49.79 % and 50.30 % of the patches, zero dead experts), yet γ grows from 0.00024 to ≈ 0.020 (×77 to ×84) — the mixture helps, but as an average of experts, not as a selection. The gain does not come from routing; it comes from the expert initialisation. This replicates on a second dataset, a second architecture and a second attach point the BRATS result The Gate Does Not Choose [DOI 10.5281/zenodo.22903668], of which this paper is a replication, not a republication: no BRATS number is reused. On the program's pre-registered versatility criterion (36 endpoints × 13 arms), MoE-V3-CS is the only all-rounder: maximal damage −0.53 pt (crowd-group instances [email protected]) against −1.90 (consensus C⊘B) and −22.20 (D, pedestrian pixel precision 70.0 → 47.8), a margin of 1.37 pt over the next arm, and the cheapest plateau of the field at −1.2 pt of damage per mIoU point against −43.1 for D. It is nonetheless first on none of the 36 endpoints (D wins 16) and its mean percentile is 60.4 (5th of 11): good everywhere is not best everywhere. Two business metrics survive the 15-pair Holm — fragments (−31.2 connected components per image) and instances_taille_T3_rappel (+0.26 pt, largest-instance recall) — while no per-class IoU rise survives the 19-class Holm (the only survivor is a fall, bicycle 0.0038). Contributions. (1) A pre-registered positive primary for an expert-initialised mixture against its recipe-matched control, reported with its three multiplicity families together and the explicit statement that it does not survive the exploratory one. (2) A routing diagnostic that separates a tautology from a proof: the equality entropy_token == entropy_token_clean at the final epoch holds by construction (the Shazeer noise is annealed to zero) and proves nothing; the probative comparison runs over the 40 active-noise epochs, where the clean distribution is already flat (≥ 0.99821) and part_max stays within [0.50011 ; 0.50217] of the 0.5000 equidistribution. (3) The correct naming of two routinely conflated routing quantities — sum(frac) == top_k (= 2), not 1, and top_expert_share (mean of per-batch maxima, 0.667 at epoch 0) ≠ part_max (maximum of per-batch means, 0.501). (4) A versatility verdict recomputed from raw deltas and ranks, with every rank carrying the denominator of its own endpoint (the MoE's worst rank is 11/12 on instances_foule_rappel, because arm A is in partial coverage and the denominator is therefore not constant). (5) A measured architectural difference with BRATS: the residual guard γ is indispensable here (without it the raw recipe costs −28.1 mIoU points at epoch 0) and absent there. (6) Public release of code, configs, tables, figures and regeneration scripts.

Zenodo (CERN European Organization for Nuclear Research)
Sustainable cities and communities
Advanced Neural Network Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.