Automated assessment and cross-system benchmarking of Chinese reading text difficulty via multi-model analysis

Abstract Text difficulty assessment is a significant task in natural language processing (NLP) and language education, providing an important basis for reading material selection and curriculum articulation. Using computational models as analytical tools, this study develops and evaluates a multidimensional automated framework for assessing the difficulty of Chinese reading texts in the International Baccalaureate Diploma Programme (IBDP), with the aim of investigating the construct of reading text difficulty and its cross-system correspondences. We extracted 173 features across five linguistic dimensions (character, lexical, syntactic, discourse, and affective features) and compared machine learning (ML), deep learning (DL), and large language model (LLM) approaches in terms of predictive performance, interpretability, and applicability. The findings indicate that Chinese reading text difficulty within the IBDP context examined in this study is multidimensional and continuous, with partly overlapping boundaries between adjacent levels. Lexical and character features made the strongest contributions to level differentiation, while syntactic, discourse, and affective features provided complementary evidence. The three computational approaches offered complementary methodological evidence: raw-text BERT achieved the highest accuracy (91.18%), feature-based XGBoost combined relatively high accuracy (85.00%) with interpretability, and the LLM showed potential for preliminary screening in zero-shot settings. The selected model was then applied to HSK reading texts, and its predictions were triangulated with external expert judgments and linguistic-feature analysis. HSK Level 4 (L4) texts were generally below the IBDP Ab Initio range; HSK L5 formed a transitional range between Ab Initio and SL; and HSK L6 was closest to HL while retaining some overlap with SL. Together, these findings show that cross-system benchmarking can reveal not only approximate correspondences but also differences in difficulty ranges and proficiency-progression patterns across assessment systems. The framework thereby provides a transparent text-level basis for reading-material selection, curriculum articulation, and future cross-system benchmarking.

Authors

Institutions

Publication Details

Journal
Humanities and Social Sciences Communications
Published
2026-09-30
DOI
https://doi.org/10.1057/s41599-026-09234-0
Primary Topic
Text Readability and Simplification
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Automated assessment and cross-system benchmarking of Chinese reading text difficulty via multi-model analysis

Shuangyun Yao, Lu Xue, Ke Shen
Humanities and Social Sciences Communications
Text Readability and Simplification
article

Automated assessment and cross-system benchmarking of Chinese reading text difficulty via multi-model analysis

Shuangyun Yao, Lu Xue, Ke Shen
article en

Abstract

Abstract Text difficulty assessment is a significant task in natural language processing (NLP) and language education, providing an important basis for reading material selection and curriculum articulation. Using computational models as analytical tools, this study develops and evaluates a multidimensional automated framework for assessing the difficulty of Chinese reading texts in the International Baccalaureate Diploma Programme (IBDP), with the aim of investigating the construct of reading text difficulty and its cross-system correspondences. We extracted 173 features across five linguistic dimensions (character, lexical, syntactic, discourse, and affective features) and compared machine learning (ML), deep learning (DL), and large language model (LLM) approaches in terms of predictive performance, interpretability, and applicability. The findings indicate that Chinese reading text difficulty within the IBDP context examined in this study is multidimensional and continuous, with partly overlapping boundaries between adjacent levels. Lexical and character features made the strongest contributions to level differentiation, while syntactic, discourse, and affective features provided complementary evidence. The three computational approaches offered complementary methodological evidence: raw-text BERT achieved the highest accuracy (91.18%), feature-based XGBoost combined relatively high accuracy (85.00%) with interpretability, and the LLM showed potential for preliminary screening in zero-shot settings. The selected model was then applied to HSK reading texts, and its predictions were triangulated with external expert judgments and linguistic-feature analysis. HSK Level 4 (L4) texts were generally below the IBDP Ab Initio range; HSK L5 formed a transitional range between Ab Initio and SL; and HSK L6 was closest to HL while retaining some overlap with SL. Together, these findings show that cross-system benchmarking can reveal not only approximate correspondences but also differences in difficulty ranges and proficiency-progression patterns across assessment systems. The framework thereby provides a transparent text-level basis for reading-material selection, curriculum articulation, and future cross-system benchmarking.

Humanities and Social Sciences Communications
Central China Normal University (CN)
Quality Education
Openalex Percentile: Top 9%
Text Readability and Simplification
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.