The dual-net multi-scale fusion framework for the digital preservation and creative transformation of intangible cultural heritage
Abstract This study proposes a deep neural network model based on the Convolutional Neural Network (CNN) algorithm combined with multi-scale feature learning—Dual-Net Multi-Scale Fusion (DN-MSF). It aims to address the problems of difficulty in high-precision restoration and low efficiency of innovative design in the current digitalization process of Intangible Cultural Heritage (ICH). Thus, the model can better achieve the dual guarantee of structural integrity and semantic coherence for multi-category ICH. The model adopts a dual-branch structure. One branch captures local multi-scale details and enhances textures, while the other branch models global semantics. It dynamically allocates cross-scale weights and maintains semantic consistency through a feature fusion module. Simultaneously, the study introduces variable receptive field convolution and hierarchical attention to optimize performance in static and dynamic scenarios. The research objects cover video materials of 8 categories of the 6th batch of national-level ICH (including folk literature, traditional music, traditional dance, etc.) from the China ICH Network ( https://www.ihchina.cn/ ). Moreover, pre-training and verification experiments are conducted using these videos. The results show that: (1) In training and testing with 80% of the dataset, DN-MSF maintains accuracy and recall above 83% across eight major categories of ICH, and its overall F1 score reaches 0.899. This finding shows that DN-MSF achieves high-fidelity restoration and accurate semantic recognition for different categories. (2) In terms of digital protection, DN-MSF demonstrates outstanding precision in image reconstruction of static ICH. It also achieves remarkable motion modeling results for dynamic ICH. (3) In innovative design tasks, categories with prominent static features such as traditional crafts and fine arts obtain a style retention score above 0.97 for generated static ICH samples. The distribution similarity index is lower than 30. This result proves that DN-MSF can achieve high-quality reconstruction of cultural styles with limited samples. The semantic consistency across all categories exhibits a correlation of Δ < 0.02 between the similarity of text-image semantic matching and text precision, verifying the semantic reliability of cross-modal outputs. The purpose of this study is to construct a deep learning framework that can support digital preservation and innovative design tasks. Through visual feature modeling, ICH reconstruction, and semantic consistency control, it provides computational support for digital expression and innovative design of ICH data under the constructed dataset and evaluation indicator system.
Authors
- Zhihong Chen
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-09-16
- DOI
- https://doi.org/10.1038/s41598-026-71995-0
- Primary Topic
- Cultural Heritage Management and Preservation
- Type
- article
- Field-Weighted Citation Impact
- 0.00