Automated assessment and cross-system benchmarking of Chinese reading text difficulty via multi-model analysis
Abstract Text difficulty assessment is a significant task in natural language processing (NLP) and language education, providing an important basis for reading material selection and curriculum articulation. Using computational models as analytical tools, this study develops and evaluates a multidimensional automated framework for assessing the difficulty of Chinese reading texts in the International Baccalaureate Diploma Programme (IBDP), with the aim of investigating the construct of reading text difficulty and its cross-system correspondences. We extracted 173 features across five linguistic dimensions (character, lexical, syntactic, discourse, and affective features) and compared machine learning (ML), deep learning (DL), and large language model (LLM) approaches in terms of predictive performance, interpretability, and applicability. The findings indicate that Chinese reading text difficulty within the IBDP context examined in this study is multidimensional and continuous, with partly overlapping boundaries between adjacent levels. Lexical and character features made the strongest contributions to level differentiation, while syntactic, discourse, and affective features provided complementary evidence. The three computational approaches offered complementary methodological evidence: raw-text BERT achieved the highest accuracy (91.18%), feature-based XGBoost combined relatively high accuracy (85.00%) with interpretability, and the LLM showed potential for preliminary screening in zero-shot settings. The selected model was then applied to HSK reading texts, and its predictions were triangulated with external expert judgments and linguistic-feature analysis. HSK Level 4 (L4) texts were generally below the IBDP Ab Initio range; HSK L5 formed a transitional range between Ab Initio and SL; and HSK L6 was closest to HL while retaining some overlap with SL. Together, these findings show that cross-system benchmarking can reveal not only approximate correspondences but also differences in difficulty ranges and proficiency-progression patterns across assessment systems. The framework thereby provides a transparent text-level basis for reading-material selection, curriculum articulation, and future cross-system benchmarking.
Authors
- Shuangyun Yao (ORCID: https://orcid.org/0000-0002-2171-6521)
- Lu Xue (ORCID: https://orcid.org/0000-0001-6284-0492)
- Ke Shen (ORCID: https://orcid.org/0009-0005-8948-4199)
Institutions
- Central China Normal University (CN)
Publication Details
- Journal
- Humanities and Social Sciences Communications
- Published
- 2026-09-30
- DOI
- https://doi.org/10.1057/s41599-026-09234-0
- Primary Topic
- Text Readability and Simplification
- Type
- article
- Field-Weighted Citation Impact
- 0.00