Evaluating the consistency of text readability levels across corpora for English education

Text readability assessment is important for natural language processing, educational text adaptation, and learner-oriented content selection. However, readability labels defined in different corpora are often not directly comparable because they are based on corpus-specific grading schemes, annotation criteria, and target learner populations. This lack of comparability limits cross-corpus evaluation, weakens model transferability, and reduces the reliable reuse of existing readability resources. To address this issue, this study formulates cross-corpus readability alignment as a consistency evaluation problem and proposes a model-agnostic framework for analyzing the compatibility of readability levels across heterogeneous corpora. The framework combines cross-corpus readability prediction with three complementary consistency metrics, namely Reverse Jensen–Shannon Divergence (RJSD), Reverse Rank Normalized Sum of Squares (RRNSS), and Normalized Discounted Cumulative Gain (NDCG). To instantiate the framework, linguistic features, GloVe-based distributed representations, fused feature settings, and transformer-based contextualized baselines are examined under conventional machine learning, recurrent neural, and transformer-based modeling settings. Experiments on six English educational readability corpora suggest that the proposed framework reveals interpretable cross-corpus consistency patterns across the evaluated corpora and model families. In general, richer feature or representation settings tend to yield more favorable consistency scores than simpler settings, although the magnitude of the advantage varies across corpus pairs and modeling choices. Additional statistical analyses, including metric correlation analysis, repeated-run variability analysis, and selected significance testing, provide further support for the descriptive findings. These results suggest that the framework offers a useful basis for evaluating readability level alignment across corpora and may support exploratory corpus reuse, educational text selection, and cross-dataset readability research.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-10-09
DOI
https://doi.org/10.1038/s41598-026-74479-3
Primary Topic
Text Readability and Simplification
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Evaluating the consistency of text readability levels across corpora for English education

Yang Liu
Scientific Reports
Text Readability and Simplification
article

Evaluating the consistency of text readability levels across corpora for English education

Yang Liu
article en

Abstract

Text readability assessment is important for natural language processing, educational text adaptation, and learner-oriented content selection. However, readability labels defined in different corpora are often not directly comparable because they are based on corpus-specific grading schemes, annotation criteria, and target learner populations. This lack of comparability limits cross-corpus evaluation, weakens model transferability, and reduces the reliable reuse of existing readability resources. To address this issue, this study formulates cross-corpus readability alignment as a consistency evaluation problem and proposes a model-agnostic framework for analyzing the compatibility of readability levels across heterogeneous corpora. The framework combines cross-corpus readability prediction with three complementary consistency metrics, namely Reverse Jensen–Shannon Divergence (RJSD), Reverse Rank Normalized Sum of Squares (RRNSS), and Normalized Discounted Cumulative Gain (NDCG). To instantiate the framework, linguistic features, GloVe-based distributed representations, fused feature settings, and transformer-based contextualized baselines are examined under conventional machine learning, recurrent neural, and transformer-based modeling settings. Experiments on six English educational readability corpora suggest that the proposed framework reveals interpretable cross-corpus consistency patterns across the evaluated corpora and model families. In general, richer feature or representation settings tend to yield more favorable consistency scores than simpler settings, although the magnitude of the advantage varies across corpus pairs and modeling choices. Additional statistical analyses, including metric correlation analysis, repeated-run variability analysis, and selected significance testing, provide further support for the descriptive findings. These results suggest that the framework offers a useful basis for evaluating readability level alignment across corpora and may support exploratory corpus reuse, educational text selection, and cross-dataset readability research.

Scientific Reports
Chifeng University (CN)
Openalex Percentile: Top 12%
Text Readability and Simplification
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.