Bayesian Hierarchical Multinomial Modeling of Multifunctional Word Usage in Learner-Corpus Data: Evidence from the English Word That
This study applies a Bayesian hierarchical multinomial model to examine the functional distribution of the English word ‘that’ in learner-corpus data. The corpus consisted of 154,694 running words of learner-generated online comments and peer replies produced by 139 first-year Chinese university EFL learners during an 18-week collaborative critical-reading course. Occurrences of that were classified into three target categories: subordinating conjunction (S), demonstrative use (D), and relative marker (R). For each learner, the category-count vector was modeled as a multinomial outcome, with standardized log text contribution included as a learner-level predictor and learner-level random effects used to account for between-learner heterogeneity. The inferential model included 135 learners with non-zero target counts and 2162 tokens of that. Subordinating-conjunction that was dominant, accounting for 69.01% of the observed tokens, followed by relative-marker that at 16.19% and demonstrative that at 14.80%. Posterior predictive checks indicated that the hierarchical model reproduced the observed aggregate category totals reasonably well. In the hierarchical model, greater text contribution was associated with lower relative odds of demonstrative that versus subordinating-conjunction that (β = −0.223, 95% CrI [−0.408, −0.036]; OR = 0.800, 95% CrI [0.665, 0.964]), whereas the corresponding association for relative-marker that was positive but uncertain (β = 0.102, 95% CrI [−0.115, 0.318]; OR = 1.108, 95% CrI [0.891, 1.374]). Learner-level variation was greater for the relative-marker contrast than for the demonstrative contrast. A pooled model without learner-level random effects produced a similar D-versus-S estimate but a larger and more precise positive R-versus-S estimate whose credible interval excluded zero (β = 0.182, 95% CrI [0.023, 0.340]). A complementary negative-binomial analysis suggested a modestly higher rate of that per 1000 running words with greater text contribution, although the evidence was uncertain (IRR = 1.097, 95% CrI [1.000, 1.204]). These findings indicate that the functional distribution of that in this dataset reflects both text-contribution patterns and meaningful learner-level heterogeneity. Because the multinomial analysis concerns the relative composition of that functions conditional on the number of that tokens produced by each learner, the observed associations should not be interpreted as evidence that longer texts necessarily contain more that tokens per unit of text or as direct evidence of general syntactic development. Rather, the findings describe patterns observed within this particular collaborative critical-reading context. The study demonstrates how Bayesian categorical-data modeling can support uncertainty-aware analysis of multifunctional word usage in learner corpora.
Authors
- Chenggang Wu (ORCID: https://orcid.org/0000-0003-3837-3841)
- Ke Song (ORCID: https://orcid.org/0000-0002-4501-8680)
- Haoxin Xu (ORCID: https://orcid.org/0000-0002-5007-509X)
- Xiyang Li (ORCID: https://orcid.org/0009-0005-0450-6708)
- Guowei Chen
Institutions
- Shanghai International Studies University (CN)
- Shanghai Normal University (CN)
Publication Details
- Journal
- Mathematics
- Published
- 2026-09-21
- DOI
- https://doi.org/10.3390/math14183433
- Primary Topic
- Second Language Acquisition and Learning
- Type
- article
- Field-Weighted Citation Impact
- 0.00