CAA-CLIP: Decoupled Pair and Cultural-Term Adaptation for Chinese Ancient Architecture Image–Text Retrieval

Image–text retrieval for Chinese ancient architecture requires fine-grained discrimination among visually similar sites and alignment with specialist terms that are sparse in general vision–language corpora. We propose CAA-CLIP, a parameter-efficient framework for exact image–annotation retrieval and annotation-derived cultural-term retrieval. CAA-CLIP adapts five frozen CLIP-family encoders through separate pair and term branches and fuses row-normalized scores with validation-selected weights. Evaluation uses 581 image–text pairs from 58 sites, five site-grouped outer folds, and three optimization seeds per fold. After seed averaging, CAA-CLIP obtains 29.06±3.05% Avg R@1, 63.19±6.34% Avg R@5, 78.65±6.64% Avg R@10, and 43.82±6.50% term mAP. Against matched validation-selected plain fusion, the pair-retrieval differences are 0.43–0.98 points, and the term-mAP difference is 9.05 points. No primary comparison reaches significance after Holm correction over the five folds, although the term difference is positive in every fold. Equal CAA fusion attains 44.39% term mAP, indicating that validation-selected term weights do not improve this endpoint. These findings distinguish expert fusion for exact-pair retrieval from direct term supervision for annotation-defined term retrieval in small, site-structured heritage collections.

Authors

Institutions

Publication Details

Journal
Mathematics
Published
2026-09-10
DOI
https://doi.org/10.3390/math14183289
Primary Topic
Handwritten Text Recognition Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

CAA-CLIP: Decoupled Pair and Cultural-Term Adaptation for Chinese Ancient Architecture Image–Text Retrieval

Wangyu Wu, Qingqing Hu, Xiaoyu Zhang, Hao Wang
Mathematics
Handwritten Text Recognition Techniques
article

CAA-CLIP: Decoupled Pair and Cultural-Term Adaptation for Chinese Ancient Architecture Image–Text Retrieval

Wangyu Wu, Qingqing Hu, Xiaoyu Zhang, Hao Wang
article en

Abstract

Image–text retrieval for Chinese ancient architecture requires fine-grained discrimination among visually similar sites and alignment with specialist terms that are sparse in general vision–language corpora. We propose CAA-CLIP, a parameter-efficient framework for exact image–annotation retrieval and annotation-derived cultural-term retrieval. CAA-CLIP adapts five frozen CLIP-family encoders through separate pair and term branches and fuses row-normalized scores with validation-selected weights. Evaluation uses 581 image–text pairs from 58 sites, five site-grouped outer folds, and three optimization seeds per fold. After seed averaging, CAA-CLIP obtains 29.06±3.05% Avg R@1, 63.19±6.34% Avg R@5, 78.65±6.64% Avg R@10, and 43.82±6.50% term mAP. Against matched validation-selected plain fusion, the pair-retrieval differences are 0.43–0.98 points, and the term-mAP difference is 9.05 points. No primary comparison reaches significance after Holm correction over the five folds, although the term difference is positive in every fold. Equal CAA fusion attains 44.39% term mAP, indicating that validation-selected term weights do not improve this endpoint. These findings distinguish expert fusion for exact-pair retrieval from direct term supervision for annotation-defined term retrieval in small, site-structured heritage collections.

MathematicsVol. 14(18)
Macau University of Science and Technology (MO), University of Liverpool (GB), Zhejiang University (CN)
Reduced inequalities
Openalex Percentile: Top 13%
Handwritten Text Recognition Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.