Benchmark of linguistic variation in LLM‑generated texts
Abstract This study investigates the register variation in texts written by humans and comparable texts produced by large language models (LLMs). Multidimensional analysis (MDA) is applied to a sample of human-written texts and AI-generated counterpart texts to find the dimensions of variation in which LLMs differ most significantly and most systematically from humans. We introduce the LLM-generated corpus AI-Brown, which is comparable to BE-21: a Brown family corpus representing contemporary British English. Since all languages except English are underrepresented in the training data of frontier LLMs, similar analysis is replicated on Czech language data using the AI-Koditex corpus and a Czech multidimensional model Cvrček et al. (2021) . Sixteen frontier models were examined in various settings and prompts, with emphasis placed on the difference between base models and instruction-tuned models. We offer a benchmark through which models can be compared with each other and ranked in interpretable dimensions.
Authors
- Jiří Milička (ORCID: https://orcid.org/0000-0001-8605-1199)
- Václav Cvrček (ORCID: https://orcid.org/0000-0003-3977-2393)
- Anna Marklová (ORCID: https://orcid.org/0000-0003-3392-1028)
Institutions
- Charles University (CZ)
Publication Details
- Journal
- International Journal of Corpus Linguistics
- Published
- 2026-09-15
- DOI
- https://doi.org/10.1075/ijcl.25151.mil
- Primary Topic
- Authorship Attribution and Profiling
- Type
- article
- Field-Weighted Citation Impact
- 0.00