Equal Weight per Instance Does Not Pay Alone: A Blob-Loss Auxiliary on Full-Resolution Cityscapes Buys Thin-Object IoU with Pedestrian Recall, and Only Pays as an Expert of a Mixture-of-Experts

We report a pre-registered, controlled evaluation of the blob loss of Kofler et al. [IPMI 2023] — an auxiliary term that gives every ground-truth instance the same weight regardless of its pixel count — ported to exclusive multi-class semantic segmentation at the native Cityscapes resolution (1024×2048, 19 classes, softmax regime). Arm G (CE + Dice + 0.5·blob) is compared against its exact paired reference B (CE + Dice): same ConvNeXt-V2-Base + UPerNet architecture, same 160-epoch recipe, same augmentation distribution, three shared seeds (42, 123, 456) — only the loss changes. The pre-registered primary endpoint is null: dataset-level official mIoU on the shared 500-image holdout, Δ(G−B) = +0.170 pt, 95 % CI [−0.233 ; +0.551], two-sided paired image-bootstrap p = 0.4032 (B = 10 000 replicates). Equal weight per instance does not improve global mIoU in this regime. What the term actually does is a measured trade-off. It buys thin-object pixel coverage — traffic light +0.98, pole +0.58, bicycle +0.57 IoU vs B (Holm-significant within the 19-class family), truck +3.78 (CI excluding zero), Boundary F1 (3 px) +0.63 — and it pays in pedestrian instance integrity: strict pedestrian recall −4.34, small-instance stratum (T1) −5.65, crowd-group instances −5.74, pedestrian pixel precision −2.70 (all vs paired B, Holm = 0), and a general fragmentation of the masks: connected components rise from 615.1 to 933.9 per image (×1.52), up in 16 of 19 classes against the paired arm B (14 Holm-significant rises; the only material decrease is terrain, −1.8) and in 19 of 19 against the control arm (person masks ×2.2) — a new per-class decomposition that revises the working mechanistic hypothesis of the program (the term does not prune small components; it creates holes). On the program's pre-registered versatility criterion (36 endpoints × 13 arms), G is a dominated specialist: mean percentile 41.9, worst rank 13/13, maximal damage −6.76 pt, 13 significant losses against 4 significant gains. The same verdict holds on a second dataset and a second probability regime: on BRATS 2023 (region-based sigmoid, MedNeXt, nnU-Net recipe, 5-fold CV, n = 1 196), the blob arm alone scored −0.00376 Dice vs baseline, yet as expert 3 of the winning MoE-V3 it contributed +0.00566 Dice, first of 29 arms [DOI 10.5281/zenodo.22903668]. On the Cityscapes program plateau, the expert-mixture arm initialised from these specialists (MoE-V3-CS, master table P3.14) points the same way, as context: ΔmIoU +0.45 pt [+0.11 ; +0.82], p = 0.0066 vs its recipe-paired control, maximal damage −0.53 pt, small-instance recall T1 +0.80, traffic light not degraded — the profile of an arm that puts G's thin-object speciality to use without carrying its damage. The dedicated analysis of that mixture arm (routing diagnostics, multiplicity families) is not the subject of this paper. Contributions. (1) A pre-registered null primary for equal-weight-per-instance auxiliary loss in full-resolution exclusive-softmax segmentation, reported as such. (2) A complete trade-off diagnostic: per-class IoU, seven business-metric families on pedestrian instances, and a new per-class connected-component decomposition (933.9 vs 615.1 components/image) that contradicts the initially hypothesised compaction mechanism — the paper reports what the tables show. (3) An exact port of Kofler's algebra (eq. 1) to the exclusive-softmax regime with precomputed instance packs: +1–2 % epoch time against +870 s/epoch for the naive path, naive↔︎precomputed parity locked by unit tests, and the horizontal-flip alignment pitfall documented. (4) Evidence, on two datasets and two probability regimes, that the term pays as an initialised expert of a mixture while null-or-harmful alone. (5) Public release of code, configs, tables, figures and regeneration scripts.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-05
DOI
https://doi.org/10.5281/zenodo.23083559
Primary Topic
Advanced Neural Network Applications
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Equal Weight per Instance Does Not Pay Alone: A Blob-Loss Auxiliary on Full-Resolution Cityscapes Buys Thin-Object IoU with Pedestrian Recall, and Only Pays as an Expert of a Mixture-of-Experts

Stanislas Larnier, Guillaume Cassez
Zenodo (CERN European Organization for Nuclear Research)
Advanced Neural Network Applications
preprint

Equal Weight per Instance Does Not Pay Alone: A Blob-Loss Auxiliary on Full-Resolution Cityscapes Buys Thin-Object IoU with Pedestrian Recall, and Only Pays as an Expert of a Mixture-of-Experts

Stanislas Larnier, Guillaume Cassez
preprint en

Abstract

We report a pre-registered, controlled evaluation of the blob loss of Kofler et al. [IPMI 2023] — an auxiliary term that gives every ground-truth instance the same weight regardless of its pixel count — ported to exclusive multi-class semantic segmentation at the native Cityscapes resolution (1024×2048, 19 classes, softmax regime). Arm G (CE + Dice + 0.5·blob) is compared against its exact paired reference B (CE + Dice): same ConvNeXt-V2-Base + UPerNet architecture, same 160-epoch recipe, same augmentation distribution, three shared seeds (42, 123, 456) — only the loss changes. The pre-registered primary endpoint is null: dataset-level official mIoU on the shared 500-image holdout, Δ(G−B) = +0.170 pt, 95 % CI [−0.233 ; +0.551], two-sided paired image-bootstrap p = 0.4032 (B = 10 000 replicates). Equal weight per instance does not improve global mIoU in this regime. What the term actually does is a measured trade-off. It buys thin-object pixel coverage — traffic light +0.98, pole +0.58, bicycle +0.57 IoU vs B (Holm-significant within the 19-class family), truck +3.78 (CI excluding zero), Boundary F1 (3 px) +0.63 — and it pays in pedestrian instance integrity: strict pedestrian recall −4.34, small-instance stratum (T1) −5.65, crowd-group instances −5.74, pedestrian pixel precision −2.70 (all vs paired B, Holm = 0), and a general fragmentation of the masks: connected components rise from 615.1 to 933.9 per image (×1.52), up in 16 of 19 classes against the paired arm B (14 Holm-significant rises; the only material decrease is terrain, −1.8) and in 19 of 19 against the control arm (person masks ×2.2) — a new per-class decomposition that revises the working mechanistic hypothesis of the program (the term does not prune small components; it creates holes). On the program's pre-registered versatility criterion (36 endpoints × 13 arms), G is a dominated specialist: mean percentile 41.9, worst rank 13/13, maximal damage −6.76 pt, 13 significant losses against 4 significant gains. The same verdict holds on a second dataset and a second probability regime: on BRATS 2023 (region-based sigmoid, MedNeXt, nnU-Net recipe, 5-fold CV, n = 1 196), the blob arm alone scored −0.00376 Dice vs baseline, yet as expert 3 of the winning MoE-V3 it contributed +0.00566 Dice, first of 29 arms [DOI 10.5281/zenodo.22903668]. On the Cityscapes program plateau, the expert-mixture arm initialised from these specialists (MoE-V3-CS, master table P3.14) points the same way, as context: ΔmIoU +0.45 pt [+0.11 ; +0.82], p = 0.0066 vs its recipe-paired control, maximal damage −0.53 pt, small-instance recall T1 +0.80, traffic light not degraded — the profile of an arm that puts G's thin-object speciality to use without carrying its damage. The dedicated analysis of that mixture arm (routing diagnostics, multiplicity families) is not the subject of this paper. Contributions. (1) A pre-registered null primary for equal-weight-per-instance auxiliary loss in full-resolution exclusive-softmax segmentation, reported as such. (2) A complete trade-off diagnostic: per-class IoU, seven business-metric families on pedestrian instances, and a new per-class connected-component decomposition (933.9 vs 615.1 components/image) that contradicts the initially hypothesised compaction mechanism — the paper reports what the tables show. (3) An exact port of Kofler's algebra (eq. 1) to the exclusive-softmax regime with precomputed instance packs: +1–2 % epoch time against +870 s/epoch for the naive path, naive↔︎precomputed parity locked by unit tests, and the horizontal-flip alignment pitfall documented. (4) Evidence, on two datasets and two probability regimes, that the term pays as an initialised expert of a mixture while null-or-harmful alone. (5) Public release of code, configs, tables, figures and regeneration scripts.

Zenodo (CERN European Organization for Nuclear Research)
Advanced Neural Network Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.