Are complex sentences reliable proxies for text-level linguistic complexity? A dependency-based corpus analysis of English

Abstract This study investigates whether English complex sentences can serve as reliable proxies for the overall linguistic complexity of the texts in which they appear. Focusing on both lexical and syntactic complexity, we analyze a dependency-annotated version of the Brown Corpus (1 million words, 15 genres), comparing full-text measures against those derived exclusively from complex sentences. Lexical complexity is operationalized via Standardized Type-Token Ratio (STTR) and Lexical Density (LD); syntactic complexity is measured using Mean Dependency Distance (MDD) for linear complexity and Mean Hierarchical Distance (MHD) for hierarchical complexity. Results show significant positive correlations between complex sentences and their host texts across all four indices and across all genres. Regression analyses further indicate that complex-sentence-based measures strongly predict full-text LD and MHD (R 2 > 0.8, large effect sizes), while their predictiveness is moderate for STTR and MDD (R 2 ≈ 0.42, medium-to-large effect sizes). These findings demonstrate that complex sentences capture key dimensions of textual complexity, with LD and MHD as robust proxies and STTR and MDD as more moderate ones. This offers a linguistically motivated and computationally efficient entry point for large-scale text analysis, provided that index-specific representativeness is taken into account. The study contributes to corpus linguistics by empirically validating a reductionist yet theoretically grounded approach to complexity measurement, and opens new avenues for research on the representativeness of syntactic structures in text-level analysis.

Authors

Institutions

Publication Details

Journal
Folia Linguistica
Published
2026-10-01
DOI
https://doi.org/10.1515/flin-2026-0052
Primary Topic
Text Readability and Simplification
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Are complex sentences reliable proxies for text-level linguistic complexity? A dependency-based corpus analysis of English

Jinlu Liu, 柴改英, Jinyi Zhang
Folia Linguistica
Text Readability and Simplification
article

Are complex sentences reliable proxies for text-level linguistic complexity? A dependency-based corpus analysis of English

Jinlu Liu, 柴改英, Jinyi Zhang
article en

Abstract

Abstract This study investigates whether English complex sentences can serve as reliable proxies for the overall linguistic complexity of the texts in which they appear. Focusing on both lexical and syntactic complexity, we analyze a dependency-annotated version of the Brown Corpus (1 million words, 15 genres), comparing full-text measures against those derived exclusively from complex sentences. Lexical complexity is operationalized via Standardized Type-Token Ratio (STTR) and Lexical Density (LD); syntactic complexity is measured using Mean Dependency Distance (MDD) for linear complexity and Mean Hierarchical Distance (MHD) for hierarchical complexity. Results show significant positive correlations between complex sentences and their host texts across all four indices and across all genres. Regression analyses further indicate that complex-sentence-based measures strongly predict full-text LD and MHD (R 2 > 0.8, large effect sizes), while their predictiveness is moderate for STTR and MDD (R 2 ≈ 0.42, medium-to-large effect sizes). These findings demonstrate that complex sentences capture key dimensions of textual complexity, with LD and MHD as robust proxies and STTR and MDD as more moderate ones. This offers a linguistically motivated and computationally efficient entry point for large-scale text analysis, provided that index-specific representativeness is taken into account. The study contributes to corpus linguistics by empirically validating a reductionist yet theoretically grounded approach to complexity measurement, and opens new avenues for research on the representativeness of syntactic structures in text-level analysis.

Folia Linguistica
Zhejiang International Studies University (CN), Zhejiang University of Finance and Economics (CN)
Quality Education
Openalex Percentile: Top 9%
Text Readability and Simplification
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Are complex sentences reliable proxies for text-level linguistic complexity? A dependency-based corpus analysis of English — Jinlu Liu, 柴改英, et al. · Folia Linguistica (2026) | TGRS Research Map | TGRS