How Much Do Earth-Observation Foundation Models Help Flood Mapping When Labels Are Scarce? Frozen, LoRA and Full Fine-Tuning of Prithvi-EO 2.0 on Sen1Floods11
Earth-observation foundation models are expected to reduce the number of labels needed for tasks such as flood mapping. We test this expectation in a controlled setting on the Sen1Floods11 benchmark. Two U-Nets (trained from scratch and ImageNet-initialized) are compared with Prithvi-EO 2.0, a 300M-parameter masked-autoencoder foundation model, under three fine-tuning regimes: frozen encoder, low-rank adaptation (LoRA) and full fine-tuning. All models see the same six Sentinel-2 bands and the same budget of optimizer steps, with 1% to 100% of the training labels (3 to 252 chips), three seeds per setting and 84 runs in total. The U-Nets lead at every label fraction: 0.829 versus 0.767 water IoU for Prithvi with LoRA at full labels, and 0.809 versus 0.704 with only three labeled chips. The gap therefore does not close as labels become scarce; it widens. Among the fine-tuning regimes, LoRA matches full fine-tuning (0.767 versus 0.770) with about 140 times fewer trainable parameters, and a frozen encoder trails both by about seven points. Pretraining also does not reduce the accuracy drop on the held-out Bolivia region. A pixel-level error analysis locates the foundation model's shortfall at water edges and on water bodies smaller than 0.1 km^2, consistent with its 16-pixel patch resolution. For Sentinel-2 flood mapping on this benchmark, a small well-tuned CNN remains the better choice, even with a handful of labeled scenes.
Authors
- Mert Semih Sarıyerli
Institutions
- Gazi University (TR)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-08
- DOI
- https://doi.org/10.5281/zenodo.23105699
- Primary Topic
- Flood Risk Assessment and Management
- Type
- preprint