It Detects More, It Scores Less: A Pre-Registered Blob Loss Refutes Its Own Precision Prediction and Pays Only as a Consensus Veto under the Official BraTS-2023 Metrics

Instance-wise losses are advertised as the remedy for small-lesion under-segmentation: they score each lesion separately, so a missed small instance costs as much as a missed large one. We pre-registered a test of one of them — blob loss (Kofler et al. 2023) — on BraTS-2023 adult glioma before reading any lesion-wise number, predicting from the source paper's multiple-sclerosis cohort that it would move a MedNeXt-B / nnU-Net baseline (Roy et al. 2023; Isensee et al. 2021) toward precision: no drop in ground-truth lesions at Dice = 0, and fewer lesion false positives. Both clauses are refuted, in opposite directions. On 1195 held-out cases pooled over five disjoint folds and scored by the official BraTS-2023 capsule, the blob arm misses 129 ground-truth lesions against 190 for the baseline (significant in 3/3 regions after Holm) and produces 2029 lesion false positives against 1418 (3/3 regions): it buys recall and pays in precision, the exact inverse of the prediction. The official post-processing then absorbs 93.2 % of the raw Dice gap and 96.2 % of the raw HD95 gap, leaving -0.00376 lesion-wise Dice (95% CI [-0.00648, -0.00103], Holm p = 8.0e-14) — 1.25 times the pre-specified smallest effect of interest (0.003) and negative on 5 folds out of 5. Of the three auxiliary terms we have measured on this backbone, blob is the only one whose confidence interval excludes zero, and it does so behind: signed-distance regression gives -0.00060 (Holm p = 1) and boundary loss -0.00116 (Holm p = 0.9547), both indistinguishable from the baseline. The arm's single positive delta is as a corroborator: the ordered veto cc_B_G ranks 1 of 12 at +0.001137, below the effect of interest and not significant after Holm (p = 0.156), and reversing the order drops it to rank 11 (-0.002306). The finding is not that the loss fails; it is that a directional prediction made in blind was contradicted by measurement, and that the sign of an instance-wise loss is a property of the dataset's instance-size distribution, not of the loss. Paper 5 of the BRATS auxiliary-loss programme: the blob loss (Kofler et al., IPMI 2023) trained ALONE, at its published weight, against the MedNeXt-B / nnU-Net baseline it was supposed to improve. The test was pre-registered in blind on 2026-08-06, before any lesion-wise number was read, and the pre-registration document ships in this archive (docs/blob_loss_preregistrement.md). Both of its clauses are refuted, in opposite directions: over 1195 held-out cases pooled on 5 disjoint folds and scored by the official BraTS-2023 capsule, the blob arm misses 129 ground-truth lesions against 190 for the baseline (significant in 3/3 regions after Holm) and produces 2029 lesion false positives against 1418 (3/3 regions). The official component filter then absorbs 93.2 % of the raw Dice gap and 96.2 % of the raw HD95 gap, leaving -0.00376 lesion-wise Dice (95% CI [-0.00648, -0.00103], Holm p = 8.0e-14), that is 1.25 times the pre-specified smallest effect of interest (0.003) and negative on 5 folds out of 5. Of the three auxiliary terms measured on this backbone, blob is the only one whose interval excludes zero, and it does so behind: signed-distance regression gives -0.00060 (Holm p = 1) and boundary loss -0.00116 (Holm p = 0.9547). Its single positive delta is as a CORROBORATOR: the ordered veto cc_B_G ranks 1 of 12 at +0.001137, under the effect of interest and not significant after Holm (p = 0.156); reversing the order (cc_G_B) drops it to rank 11 (-0.002306). Training: 7 runs at 300 epochs (5 folds at seed 42 plus 2 extra seeds on fold 0), beta = 0.5, signed-distance weight = 0.0, each value re-read from the run's own debug.json. The archive ships the EN and FR manuscripts (markdown + rendered PDF), every published table with the JSON path of each of its values, the five figures with their manifest, the prose templates the manuscripts are assembled from, the consolidated corpus (analysis/blobloss_stats.json, 36/36 internal controls), the 16 source artefacts it reads (per-patient official-metric CSVs of the five folds, the fold-0 evaluation, the campaign log and one debug.json per run) under data/ at their original paths with size and md5 verified against the consolidation, the pre-registration, and the reproduction code (official-metric evaluation, statistics, tables, figures, manuscript assembly, coherence and PDF-layout measurement). No stage of the pipeline requires a GPU.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-04
DOI
https://doi.org/10.5281/zenodo.23142692
Primary Topic
Medical Image Segmentation Techniques
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

It Detects More, It Scores Less: A Pre-Registered Blob Loss Refutes Its Own Precision Prediction and Pays Only as a Consensus Veto under the Official BraTS-2023 Metrics

Stanislas Larnier, Guillaume Cassez
Zenodo (CERN European Organization for Nuclear Research)
Medical Image Segmentation Techniques
preprint

It Detects More, It Scores Less: A Pre-Registered Blob Loss Refutes Its Own Precision Prediction and Pays Only as a Consensus Veto under the Official BraTS-2023 Metrics

Stanislas Larnier, Guillaume Cassez
preprint en

Abstract

Instance-wise losses are advertised as the remedy for small-lesion under-segmentation: they score each lesion separately, so a missed small instance costs as much as a missed large one. We pre-registered a test of one of them — blob loss (Kofler et al. 2023) — on BraTS-2023 adult glioma before reading any lesion-wise number, predicting from the source paper's multiple-sclerosis cohort that it would move a MedNeXt-B / nnU-Net baseline (Roy et al. 2023; Isensee et al. 2021) toward precision: no drop in ground-truth lesions at Dice = 0, and fewer lesion false positives. Both clauses are refuted, in opposite directions. On 1195 held-out cases pooled over five disjoint folds and scored by the official BraTS-2023 capsule, the blob arm misses 129 ground-truth lesions against 190 for the baseline (significant in 3/3 regions after Holm) and produces 2029 lesion false positives against 1418 (3/3 regions): it buys recall and pays in precision, the exact inverse of the prediction. The official post-processing then absorbs 93.2 % of the raw Dice gap and 96.2 % of the raw HD95 gap, leaving -0.00376 lesion-wise Dice (95% CI [-0.00648, -0.00103], Holm p = 8.0e-14) — 1.25 times the pre-specified smallest effect of interest (0.003) and negative on 5 folds out of 5. Of the three auxiliary terms we have measured on this backbone, blob is the only one whose confidence interval excludes zero, and it does so behind: signed-distance regression gives -0.00060 (Holm p = 1) and boundary loss -0.00116 (Holm p = 0.9547), both indistinguishable from the baseline. The arm's single positive delta is as a corroborator: the ordered veto cc_B_G ranks 1 of 12 at +0.001137, below the effect of interest and not significant after Holm (p = 0.156), and reversing the order drops it to rank 11 (-0.002306). The finding is not that the loss fails; it is that a directional prediction made in blind was contradicted by measurement, and that the sign of an instance-wise loss is a property of the dataset's instance-size distribution, not of the loss. Paper 5 of the BRATS auxiliary-loss programme: the blob loss (Kofler et al., IPMI 2023) trained ALONE, at its published weight, against the MedNeXt-B / nnU-Net baseline it was supposed to improve. The test was pre-registered in blind on 2026-08-06, before any lesion-wise number was read, and the pre-registration document ships in this archive (docs/blob_loss_preregistrement.md). Both of its clauses are refuted, in opposite directions: over 1195 held-out cases pooled on 5 disjoint folds and scored by the official BraTS-2023 capsule, the blob arm misses 129 ground-truth lesions against 190 for the baseline (significant in 3/3 regions after Holm) and produces 2029 lesion false positives against 1418 (3/3 regions). The official component filter then absorbs 93.2 % of the raw Dice gap and 96.2 % of the raw HD95 gap, leaving -0.00376 lesion-wise Dice (95% CI [-0.00648, -0.00103], Holm p = 8.0e-14), that is 1.25 times the pre-specified smallest effect of interest (0.003) and negative on 5 folds out of 5. Of the three auxiliary terms measured on this backbone, blob is the only one whose interval excludes zero, and it does so behind: signed-distance regression gives -0.00060 (Holm p = 1) and boundary loss -0.00116 (Holm p = 0.9547). Its single positive delta is as a CORROBORATOR: the ordered veto cc_B_G ranks 1 of 12 at +0.001137, under the effect of interest and not significant after Holm (p = 0.156); reversing the order (cc_G_B) drops it to rank 11 (-0.002306). Training: 7 runs at 300 epochs (5 folds at seed 42 plus 2 extra seeds on fold 0), beta = 0.5, signed-distance weight = 0.0, each value re-read from the run's own debug.json. The archive ships the EN and FR manuscripts (markdown + rendered PDF), every published table with the JSON path of each of its values, the five figures with their manifest, the prose templates the manuscripts are assembled from, the consolidated corpus (analysis/blobloss_stats.json, 36/36 internal controls), the 16 source artefacts it reads (per-patient official-metric CSVs of the five folds, the fold-0 evaluation, the campaign log and one debug.json per run) under data/ at their original paths with size and md5 verified against the consolidation, the pre-registration, and the reproduction code (official-metric evaluation, statistics, tables, figures, manuscript assembly, coherence and PDF-layout measurement). No stage of the pipeline requires a GPU.

Zenodo (CERN European Organization for Nuclear Research)
Good health and well-being
Medical Image Segmentation Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.