DPBERT: A domain-specific pre-trained model for Chinese data policy texts and its interpretability

Chinese data policy texts are characterised by dense terminology, standardised expressions and complex domain-specific semantics, which pose challenges for general-purpose language models. This study develops Data-Policy BERT (DPBERT), a domain-specific pre-trained model tailored to Chinese data policy texts. Using Chinese-bidirectional encoder representations from transformer-whole-word masking as the backbone, we conduct continued pre-training on 65,498 Chinese data-related policy documents with two masking strategies: masked language modelling and whole-word masking. The resulting models are evaluated on policy text classification, named entity recognition and an interpretability analysis based on gradient-based saliency. Experimental results show that DPBERT-whole-word masking outperforms the baseline models on both downstream tasks, indicating that whole-word masking is better suited to capturing Chinese policy terms and composite concepts. We further construct an interpretability evaluation framework using gradient-based saliency and a feature salient value metric to examine how models attend to core policy tokens. DPBERT-whole-word masking exhibits more concentrated attention on high-saliency policy features. DPBERT and Large Language Model Meta AI are compared under the same Chinese data, fine-tuning procedure and evaluation metrics; the results indicate that DPBERT achieves better performance on policy text classification and named entity recognition within this study’s task scope. This work provides a reference for semantic modelling, entity recognition and interpretable analysis of Chinese data policy texts.

Authors

Institutions

Publication Details

Journal
Journal of Information Science
Published
2026-09-30
DOI
https://doi.org/10.1177/01655515261489505
Primary Topic
Topic Modeling
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

DPBERT: A domain-specific pre-trained model for Chinese data policy texts and its interpretability

Ma Haiqun, Tao Zhang, Zhang Ce, Wang Hangong et al.
Journal of Information Science
Topic Modeling
article

DPBERT: A domain-specific pre-trained model for Chinese data policy texts and its interpretability

Ma Haiqun, Tao Zhang, Zhang Ce, Wang Hangong, Jiang Lei
article en

Abstract

Chinese data policy texts are characterised by dense terminology, standardised expressions and complex domain-specific semantics, which pose challenges for general-purpose language models. This study develops Data-Policy BERT (DPBERT), a domain-specific pre-trained model tailored to Chinese data policy texts. Using Chinese-bidirectional encoder representations from transformer-whole-word masking as the backbone, we conduct continued pre-training on 65,498 Chinese data-related policy documents with two masking strategies: masked language modelling and whole-word masking. The resulting models are evaluated on policy text classification, named entity recognition and an interpretability analysis based on gradient-based saliency. Experimental results show that DPBERT-whole-word masking outperforms the baseline models on both downstream tasks, indicating that whole-word masking is better suited to capturing Chinese policy terms and composite concepts. We further construct an interpretability evaluation framework using gradient-based saliency and a feature salient value metric to examine how models attend to core policy tokens. DPBERT-whole-word masking exhibits more concentrated attention on high-saliency policy features. DPBERT and Large Language Model Meta AI are compared under the same Chinese data, fine-tuning procedure and evaluation metrics; the results indicate that DPBERT achieves better performance on policy text classification and named entity recognition within this study’s task scope. This work provides a reference for semantic modelling, entity recognition and interpretable analysis of Chinese data policy texts.

Journal of Information Science
Sichuan University (CN), Heilongjiang University (CN)
Peace, Justice and strong institutions
Openalex Percentile: Top 9%
Topic Modeling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

DPBERT: A domain-specific pre-trained model for Chinese data policy texts and its interpretability — Ma Haiqun, Tao Zhang, et al. · Journal of Information Science (2026) | TGRS Research Map | TGRS