Equal Weight per Instance Does Not Pay Alone: A Blob-Loss Auxiliary on Full-Resolution Cityscapes Buys Thin-Object IoU with Pedestrian Recall, and Only Pays as an Expert of a Mixture-of-Experts

We report a pre-registered, controlled evaluation of the **blob loss** of Kofler *et al.* [IPMI 2023] — an auxiliary term that gives every ground-truth instance the same weight regardless of its pixel count — ported to **exclusive multi-class semantic segmentation at the native Cityscapes resolution** (1024×2048, 19 classes, softmax regime). Arm **G** (CE + Dice + 0.5·blob) is compared against its exact paired reference **B** (CE + Dice): same ConvNeXt-V2-Base + UPerNet architecture, same 160-epoch recipe, same augmentation distribution, three shared seeds (42, 123, 456) — only the loss changes. The **pre-registered primary endpoint is null**: dataset-level official mIoU on the shared 500-image holdout, Δ(G−B) = **+0.170 pt**, 95 % CI [−0.233 ; +0.551], two-sided paired image-bootstrap p = **0.4032** (B = 10 000 replicates). Equal weight per instance does **not** improve global mIoU in this regime. What the term actually does is a **measured trade-off**. It *buys* thin-object pixel coverage — traffic light +0.98, pole +0.58, bicycle +0.57 IoU vs B (Holm-significant within the 19-class family), truck +3.78 (CI excluding zero), Boundary F1 (3 px) +0.63 — and it *pays* in pedestrian instance integrity: strict pedestrian recall **−4.34**, small-instance stratum (T1) **−5.65**, crowd-group instances **−5.74**, pedestrian pixel precision **−2.70** (all vs paired B, Holm = 0), and a general **fragmentation** of the masks: connected components rise from 615.1 to 933.9 per image (**×1.52**), up in **16 of 19 classes** against the paired arm B (14 Holm-significant rises; the only material decrease is terrain, −1.8) and in **19 of 19** against the control arm (person masks ×2.2) — a new per-class decomposition that *revises* the working mechanistic hypothesis of the program (the term does not prune small components; it creates holes). On the program's pre-registered versatility criterion (36 endpoints × 13 arms), G is a **dominated specialist**: mean percentile 41.9, worst rank 13/13, maximal damage −6.76 pt, 13 significant losses against 4 significant gains. The same verdict holds on a second dataset and a second probability regime: on BRATS 2023 (region-based sigmoid, MedNeXt, nnU-Net recipe, 5-fold CV, n = 1 196), the blob arm alone scored **−0.00376 Dice** vs baseline, yet as **expert 3 of the winning MoE-V3** it contributed **+0.00566 Dice, first of 29 arms** [DOI 10.5281/zenodo.22903668]. The Cityscapes companion mixture **MoE-V3-CS** (four experts initialised from B, D, Dp, G; top-2 patch-wise gate) tells the same story: it absorbs G's thin-object speciality — small-instance recall T1 +0.80 (significant), traffic light not degraded — without inheriting its damage (maximal damage −0.53 pt; ΔmIoU +0.45 pt, p = 0.0066 vs its paired control). **Contributions.** (1) A pre-registered **null primary** for equal-weight-per-instance auxiliary loss in full-resolution exclusive-softmax segmentation, reported as such. (2) A complete **trade-off diagnostic**: per-class IoU, seven business-metric families on pedestrian instances, and a **new per-class connected-component decomposition** (933.9 vs 615.1 components/image) that contradicts the initially hypothesised compaction mechanism — the paper reports what the tables show. (3) An **exact port** of Kofler's algebra (eq. 1) to the exclusive-softmax regime with precomputed instance packs: +1–2 % epoch time against +870 s/epoch for the naive path, naive↔precomputed parity locked by unit tests, and the horizontal-flip alignment pitfall documented. (4) Evidence, on two datasets and two probability regimes, that the term **pays as an initialised expert of a mixture** while null-or-harmful alone. (5) Public release of code, configs, tables, figures and regeneration scripts. ---

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-01
DOI
https://doi.org/10.5281/zenodo.23083560
Primary Topic
Advanced Neural Network Applications
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Equal Weight per Instance Does Not Pay Alone: A Blob-Loss Auxiliary on Full-Resolution Cityscapes Buys Thin-Object IoU with Pedestrian Recall, and Only Pays as an Expert of a Mixture-of-Experts

Guillaume Cassez
Zenodo (CERN European Organization for Nuclear Research)
Advanced Neural Network Applications
preprint

Equal Weight per Instance Does Not Pay Alone: A Blob-Loss Auxiliary on Full-Resolution Cityscapes Buys Thin-Object IoU with Pedestrian Recall, and Only Pays as an Expert of a Mixture-of-Experts

Guillaume Cassez
preprint en

Abstract

We report a pre-registered, controlled evaluation of the **blob loss** of Kofler *et al.* [IPMI 2023] — an auxiliary term that gives every ground-truth instance the same weight regardless of its pixel count — ported to **exclusive multi-class semantic segmentation at the native Cityscapes resolution** (1024×2048, 19 classes, softmax regime). Arm **G** (CE + Dice + 0.5·blob) is compared against its exact paired reference **B** (CE + Dice): same ConvNeXt-V2-Base + UPerNet architecture, same 160-epoch recipe, same augmentation distribution, three shared seeds (42, 123, 456) — only the loss changes. The **pre-registered primary endpoint is null**: dataset-level official mIoU on the shared 500-image holdout, Δ(G−B) = **+0.170 pt**, 95 % CI [−0.233 ; +0.551], two-sided paired image-bootstrap p = **0.4032** (B = 10 000 replicates). Equal weight per instance does **not** improve global mIoU in this regime. What the term actually does is a **measured trade-off**. It *buys* thin-object pixel coverage — traffic light +0.98, pole +0.58, bicycle +0.57 IoU vs B (Holm-significant within the 19-class family), truck +3.78 (CI excluding zero), Boundary F1 (3 px) +0.63 — and it *pays* in pedestrian instance integrity: strict pedestrian recall **−4.34**, small-instance stratum (T1) **−5.65**, crowd-group instances **−5.74**, pedestrian pixel precision **−2.70** (all vs paired B, Holm = 0), and a general **fragmentation** of the masks: connected components rise from 615.1 to 933.9 per image (**×1.52**), up in **16 of 19 classes** against the paired arm B (14 Holm-significant rises; the only material decrease is terrain, −1.8) and in **19 of 19** against the control arm (person masks ×2.2) — a new per-class decomposition that *revises* the working mechanistic hypothesis of the program (the term does not prune small components; it creates holes). On the program's pre-registered versatility criterion (36 endpoints × 13 arms), G is a **dominated specialist**: mean percentile 41.9, worst rank 13/13, maximal damage −6.76 pt, 13 significant losses against 4 significant gains. The same verdict holds on a second dataset and a second probability regime: on BRATS 2023 (region-based sigmoid, MedNeXt, nnU-Net recipe, 5-fold CV, n = 1 196), the blob arm alone scored **−0.00376 Dice** vs baseline, yet as **expert 3 of the winning MoE-V3** it contributed **+0.00566 Dice, first of 29 arms** [DOI 10.5281/zenodo.22903668]. The Cityscapes companion mixture **MoE-V3-CS** (four experts initialised from B, D, Dp, G; top-2 patch-wise gate) tells the same story: it absorbs G's thin-object speciality — small-instance recall T1 +0.80 (significant), traffic light not degraded — without inheriting its damage (maximal damage −0.53 pt; ΔmIoU +0.45 pt, p = 0.0066 vs its paired control). **Contributions.** (1) A pre-registered **null primary** for equal-weight-per-instance auxiliary loss in full-resolution exclusive-softmax segmentation, reported as such. (2) A complete **trade-off diagnostic**: per-class IoU, seven business-metric families on pedestrian instances, and a **new per-class connected-component decomposition** (933.9 vs 615.1 components/image) that contradicts the initially hypothesised compaction mechanism — the paper reports what the tables show. (3) An **exact port** of Kofler's algebra (eq. 1) to the exclusive-softmax regime with precomputed instance packs: +1–2 % epoch time against +870 s/epoch for the naive path, naive↔precomputed parity locked by unit tests, and the horizontal-flip alignment pitfall documented. (4) Evidence, on two datasets and two probability regimes, that the term **pays as an initialised expert of a mixture** while null-or-harmful alone. (5) Public release of code, configs, tables, figures and regeneration scripts. ---

Zenodo (CERN European Organization for Nuclear Research)
Sustainable cities and communities
Advanced Neural Network Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.