A Benchmark for Cultural-Heritage Deterioration Segmentation: Semi-Supervised Baselines, Auto-Discovery, and Calibration-Aware Evaluation

Cultural-heritage preservation increasingly relies on computing technologies for documentation, restoration, analysis, and risk-informed conservation. Pixel-level deterioration segmentation remains challenging because heritage archives are heterogeneous and weakly curated, expert annotations are scarce, foreground regions are highly imbalanced, and authentic damage differs substantially from standard computer-vision benchmarks. We address a practical archival setting in which many image fragments are available but only a small subset has reliable pixel-level masks. Our benchmark integrates archive auto-discovery, Mean Teacher semi-supervised learning, and calibration-aware evaluation within a reproducible protocol. On the curated MuralDH subset ( \(N=961\) ), we compare LR-ASPP, DeepLabV3+, U-Net++, and SegFormer under identical conditions. SegFormer (MiT-B2) provides the strongest discrimination–calibration trade-off on the held-out test set ( \(N=201\) ), achieving 94.9% pixel accuracy, 0.783 F1, 0.643 IoU, and 0.019 ECE. We also apply the protocol to DeepCrack ( \(N=537\) ) and CrackSeg9k ( \(N=9{,}159\) ) as auxiliary cross-domain stress tests. SegFormer remains the strongest model, while the lower CrackSeg9k performance exposes sensitivity to thin, elongated, low-contrast structures. Qualitative comparisons, error maps, MC-Dropout uncertainty maps, and failure analysis show how decorative contours, pigment transitions, and textured regions can be confused with deterioration. The benchmark therefore provides both segmentation performance and diagnostic evidence for interpreting model behavior in conservation-oriented workflows.

Authors

Institutions

Publication Details

Journal
Journal on Computing and Cultural Heritage
Published
2026-10-06
DOI
https://doi.org/10.1145/3850157
Primary Topic
Infrastructure Maintenance and Monitoring
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

A Benchmark for Cultural-Heritage Deterioration Segmentation: Semi-Supervised Baselines, Auto-Discovery, and Calibration-Aware Evaluation

José Mendoza-Valdés, Christian Quesada López
Journal on Computing and Cultural Heritage
Infrastructure Maintenance and Monitoring
article

A Benchmark for Cultural-Heritage Deterioration Segmentation: Semi-Supervised Baselines, Auto-Discovery, and Calibration-Aware Evaluation

José Mendoza-Valdés, Christian Quesada López
article en

Abstract

Cultural-heritage preservation increasingly relies on computing technologies for documentation, restoration, analysis, and risk-informed conservation. Pixel-level deterioration segmentation remains challenging because heritage archives are heterogeneous and weakly curated, expert annotations are scarce, foreground regions are highly imbalanced, and authentic damage differs substantially from standard computer-vision benchmarks. We address a practical archival setting in which many image fragments are available but only a small subset has reliable pixel-level masks. Our benchmark integrates archive auto-discovery, Mean Teacher semi-supervised learning, and calibration-aware evaluation within a reproducible protocol. On the curated MuralDH subset ( \(N=961\) ), we compare LR-ASPP, DeepLabV3+, U-Net++, and SegFormer under identical conditions. SegFormer (MiT-B2) provides the strongest discrimination–calibration trade-off on the held-out test set ( \(N=201\) ), achieving 94.9% pixel accuracy, 0.783 F1, 0.643 IoU, and 0.019 ECE. We also apply the protocol to DeepCrack ( \(N=537\) ) and CrackSeg9k ( \(N=9{,}159\) ) as auxiliary cross-domain stress tests. SegFormer remains the strongest model, while the lower CrackSeg9k performance exposes sensitivity to thin, elongated, low-contrast structures. Qualitative comparisons, error maps, MC-Dropout uncertainty maps, and failure analysis show how decorative contours, pigment transitions, and textured regions can be confused with deterioration. The benchmark therefore provides both segmentation performance and diagnostic evidence for interpreting model behavior in conservation-oriented workflows.

Journal on Computing and Cultural Heritage
Universidad de Costa Rica (CR), Universidad Tecnológica de Panamá (PA)
Openalex Percentile: Top 17%
Infrastructure Maintenance and Monitoring
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

A Benchmark for Cultural-Heritage Deterioration Segmentation: Semi-Supervised Baselines, Auto-Discovery, and Calibration-Aware Evaluation — José Mendoza-Valdés, Christian Quesada López · Journal on Computing and Cultural Heritage (2026) | TGRS Research Map | TGRS