A Fine-Grained Image Dataset of Gold and Silver Artifacts from the Tomb of Prince Liangzhuang and Preliminary LoRA-Based Generation Experiments

The digital preservation of cultural heritage is evolving from static archiving towards intelligent creation. As representative examples of early Ming Dynasty gold and silver craftsmanship, the gold and silver artifacts from the Tomb of Prince Liangzhuang, with their complex craftsmanship and rich cultural significance embodied in their decorative motifs, pose substantial challenges for digital preservation and innovative design. Generative artificial intelligence currently faces difficulties in reproducing fine-grained forms and decorative features, while the lack of high-quality, multi-dimensionally annotated image data specifically designed for such artifacts further constrains their digital preservation and application. To address these issues, this study constructed a fine-grained, multi-dimensional annotated image dataset of gold and silver artifacts from the Tomb of Prince Liangzhuang. The dataset covers five major categories, namely headwear, headgear and sashes, accessories, daily necessities, and ritual objects, comprising 82 artifacts and 4449 images, including 4127 panoramic images and 322 detail images. All images were annotated according to four visual feature dimensions: form and structure, decorative patterns, material appearance, and color composition. Based on this dataset, LoRA fine-tuning was performed on three mainstream diffusion models, namely SD1.5, SDXL, and FLUX. In the experimental evaluation, a total of 1968 images were generated for the 82 artifacts using the three models and eight random seeds (82 artifacts × 3 models × 8 random seeds). The generation results were systematically evaluated using quantitative metrics, including Structural Similarity Index Measure (SSIM), Learned Perceptual Image Patch Similarity (LPIPS), CLIP Score, and Kernel Inception Distance (KID), together with five types of subjective evaluation questionnaires, yielding 101 valid questionnaire responses in total. The results showed that all three LoRA-fine-tuned models were capable of generating images exhibiting typical characteristics of the gold and silver artifacts from the Tomb of Prince Liangzhuang. SD1.5_LoRA achieved the highest SSIM score (0.627) and the lowest seed-to-seed standard deviation (0.050), with statistically significant differences observed in most pairwise comparisons across categories after Bonferroni correction. SDXL_LoRA achieved the highest composite score in the subjective evaluation. CLIP Score showed category-dependent differences across the five artifact categories, and no consistent overall ranking of the three models was observed, suggesting comparable semantic alignment performance across the models. The resolution comparison experiment showed that, for the headwear category using SDXL_LoRA, no significant differences were observed across the evaluated metrics between the 768 × 768 and 1024 × 1024 resolutions. The results demonstrate that the proposed dataset can effectively support LoRA fine-tuning for the generation of gold and silver artifacts from the Tomb of Prince Liangzhuang. The three models showed different strengths in structural fidelity, semantic alignment, and material appearance, indicating that model selection can be tailored to specific application requirements. This study provides a high-quality data resource and a reusable technical approach for the digital preservation of Ming Dynasty gold and silver artifacts.

Authors

Institutions

Publication Details

Journal
Sensors
Published
2026-10-07
DOI
https://doi.org/10.3390/s26196322
Primary Topic
Generative Adversarial Networks and Image Synthesis
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

A Fine-Grained Image Dataset of Gold and Silver Artifacts from the Tomb of Prince Liangzhuang and Preliminary LoRA-Based Generation Experiments

Lei Shu, Yelei Chen, Yifan Xu, Ru Han
Sensors
Generative Adversarial Networks and Image Synthesis
article

A Fine-Grained Image Dataset of Gold and Silver Artifacts from the Tomb of Prince Liangzhuang and Preliminary LoRA-Based Generation Experiments

Lei Shu, Yelei Chen, Yifan Xu, Ru Han
article en

Abstract

The digital preservation of cultural heritage is evolving from static archiving towards intelligent creation. As representative examples of early Ming Dynasty gold and silver craftsmanship, the gold and silver artifacts from the Tomb of Prince Liangzhuang, with their complex craftsmanship and rich cultural significance embodied in their decorative motifs, pose substantial challenges for digital preservation and innovative design. Generative artificial intelligence currently faces difficulties in reproducing fine-grained forms and decorative features, while the lack of high-quality, multi-dimensionally annotated image data specifically designed for such artifacts further constrains their digital preservation and application. To address these issues, this study constructed a fine-grained, multi-dimensional annotated image dataset of gold and silver artifacts from the Tomb of Prince Liangzhuang. The dataset covers five major categories, namely headwear, headgear and sashes, accessories, daily necessities, and ritual objects, comprising 82 artifacts and 4449 images, including 4127 panoramic images and 322 detail images. All images were annotated according to four visual feature dimensions: form and structure, decorative patterns, material appearance, and color composition. Based on this dataset, LoRA fine-tuning was performed on three mainstream diffusion models, namely SD1.5, SDXL, and FLUX. In the experimental evaluation, a total of 1968 images were generated for the 82 artifacts using the three models and eight random seeds (82 artifacts × 3 models × 8 random seeds). The generation results were systematically evaluated using quantitative metrics, including Structural Similarity Index Measure (SSIM), Learned Perceptual Image Patch Similarity (LPIPS), CLIP Score, and Kernel Inception Distance (KID), together with five types of subjective evaluation questionnaires, yielding 101 valid questionnaire responses in total. The results showed that all three LoRA-fine-tuned models were capable of generating images exhibiting typical characteristics of the gold and silver artifacts from the Tomb of Prince Liangzhuang. SD1.5_LoRA achieved the highest SSIM score (0.627) and the lowest seed-to-seed standard deviation (0.050), with statistically significant differences observed in most pairwise comparisons across categories after Bonferroni correction. SDXL_LoRA achieved the highest composite score in the subjective evaluation. CLIP Score showed category-dependent differences across the five artifact categories, and no consistent overall ranking of the three models was observed, suggesting comparable semantic alignment performance across the models. The resolution comparison experiment showed that, for the headwear category using SDXL_LoRA, no significant differences were observed across the evaluated metrics between the 768 × 768 and 1024 × 1024 resolutions. The results demonstrate that the proposed dataset can effectively support LoRA fine-tuning for the generation of gold and silver artifacts from the Tomb of Prince Liangzhuang. The three models showed different strengths in structural fidelity, semantic alignment, and material appearance, indicating that model selection can be tailored to specific application requirements. This study provides a high-quality data resource and a reusable technical approach for the digital preservation of Ming Dynasty gold and silver artifacts.

SensorsVol. 26(19)
Nanjing Agricultural University (CN), University of Lincoln (GB), Hubei University of Technology (CN)
Openalex Percentile: Top 15%
Generative Adversarial Networks and Image Synthesis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.