Diagnosing landslide segmentation across geographic clusters with limited target annotations and no access to source imagery

Abstract Landslide segmentation models that perform well within their training distribution can degrade sharply when transferred to a different geographic region, limiting their reliability for rapid post-event mapping. We study this problem when the original training data are unavailable and only a small target-annotation budget is permitted. Using six geographic clusters from Sen12Landslides, we separately probe decision-threshold calibration, trainable parameter scope, annotation and optimization budgets, and target-unlabelled adaptation. Three independently trained source chains obtain mean positive-class pixel $$F_1$$ scores of 0.181, 0.199, and 0.180 on matched support-excluded query pools. Oracle scalar thresholds selected using query labels improve these scores to 0.221, 0.256, and 0.212, whereas BN-clean decoder-plus-head fine-tuning reaches 0.199, 0.220, and 0.203 and full-network fine-tuning reaches 0.260, 0.289, and 0.270. The paired gains obtained by additionally opening the encoder-like path are 0.061, 0.069, and 0.067, with positive cluster-level differences in five, six, and five of six clusters. Across the tested prevalence-preserving sampling regimes, larger support sets are associated with higher but heterogeneous mean scores; because support identities and support-excluded query pools differ across K , this comparison does not isolate the causal effect of tile count alone. Under the historical independent-overlap component convention, mean patch-component recall increases from 0.116 to 0.205, with a corresponding false-positive cost. Class-balanced pseudo-label self-training also raises the mean for all three source chains, whereas entropy minimization collapses positive-class prediction. In a single-seed CAS analysis, one of two full-network variants ranks highest within the finite candidate set for each of five held-out events. Together, these results identify encoder-inclusive full-network updating as the strongest supervised repair path on average among the evaluated interventions and establish a reproducible diagnostic and annotation-budget analysis for source-free cross-cluster landslide segmentation.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-09-10
DOI
https://doi.org/10.1038/s41598-026-70729-6
Primary Topic
Landslides and related hazards
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Diagnosing landslide segmentation across geographic clusters with limited target annotations and no access to source imagery

Xinfang Chen, Cun Weng, Zhuangzhi Wang
Scientific Reports
Landslides and related hazards
article

Diagnosing landslide segmentation across geographic clusters with limited target annotations and no access to source imagery

Xinfang Chen, Cun Weng, Zhuangzhi Wang
article en

Abstract

Abstract Landslide segmentation models that perform well within their training distribution can degrade sharply when transferred to a different geographic region, limiting their reliability for rapid post-event mapping. We study this problem when the original training data are unavailable and only a small target-annotation budget is permitted. Using six geographic clusters from Sen12Landslides, we separately probe decision-threshold calibration, trainable parameter scope, annotation and optimization budgets, and target-unlabelled adaptation. Three independently trained source chains obtain mean positive-class pixel $$F_1$$ scores of 0.181, 0.199, and 0.180 on matched support-excluded query pools. Oracle scalar thresholds selected using query labels improve these scores to 0.221, 0.256, and 0.212, whereas BN-clean decoder-plus-head fine-tuning reaches 0.199, 0.220, and 0.203 and full-network fine-tuning reaches 0.260, 0.289, and 0.270. The paired gains obtained by additionally opening the encoder-like path are 0.061, 0.069, and 0.067, with positive cluster-level differences in five, six, and five of six clusters. Across the tested prevalence-preserving sampling regimes, larger support sets are associated with higher but heterogeneous mean scores; because support identities and support-excluded query pools differ across K , this comparison does not isolate the causal effect of tile count alone. Under the historical independent-overlap component convention, mean patch-component recall increases from 0.116 to 0.205, with a corresponding false-positive cost. Class-balanced pseudo-label self-training also raises the mean for all three source chains, whereas entropy minimization collapses positive-class prediction. In a single-seed CAS analysis, one of two full-network variants ranks highest within the finite candidate set for each of five held-out events. Together, these results identify encoder-inclusive full-network updating as the strongest supervised repair path on average among the evaluated interventions and establish a reproducible diagnostic and annotation-budget analysis for source-free cross-cluster landslide segmentation.

Scientific Reports
ENN (China) (CN)
Openalex Percentile: Top 5%
Landslides and related hazards
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.