Tao Zhang, Zhang Ce, Wang Hangong, Ma Haiqun, Jiang Lei · Journal of Information Science 2026 · 2026
DOI: 10.1177/01655515261489505
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Chinese data policy texts are characterised by dense terminology, standardised expressions and complex domain-specific semantics, which pose challenges for general-purpose language models. This study develops Data-Policy BERT (DPBERT), a domain-specific pre-trained model tailored to Chinese data policy texts. Using Chinese-bidirectional encoder representations from transformer-whole-word masking as the backbone, we conduct continued pre-training on 65,498 Chinese data-related policy documents with two masking strategies: masked language modelling and whole-word masking. The resulting models are evaluated on policy text classification, named entity recognition and an interpretability analysis based on gradient-based saliency. Experimental results show that DPBERT-whole-word masking outperforms the baseline models on both downstream tasks, indicating that whole-word masking is better suited to capturing Chinese policy terms and composite concepts. We further construct an interpretability evaluation framework using gradient-based saliency and a feature salient value metric to examine how models attend to core policy tokens. DPBERT-whole-word masking exhibits more concentrated attention on high-saliency policy features. DPBERT and Large Language Model Meta AI are compared under the same Chinese data, fine-tuning procedure and evaluation metrics; the results indicate that DPBERT achieves better performance on policy text classification and named entity recognition within this study’s task scope. This work provides a reference for semantic modelling, entity recognition and interpretable analysis of Chinese data policy texts.
No comments yet — start the discussion below.