Benchmark of linguistic variation in LLM‑generated texts

Abstract This study investigates the register variation in texts written by humans and comparable texts produced by large language models (LLMs). Multidimensional analysis (MDA) is applied to a sample of human-written texts and AI-generated counterpart texts to find the dimensions of variation in which LLMs differ most significantly and most systematically from humans. We introduce the LLM-generated corpus AI-Brown, which is comparable to BE-21: a Brown family corpus representing contemporary British English. Since all languages except English are underrepresented in the training data of frontier LLMs, similar analysis is replicated on Czech language data using the AI-Koditex corpus and a Czech multidimensional model Cvrček et al. (2021) . Sixteen frontier models were examined in various settings and prompts, with emphasis placed on the difference between base models and instruction-tuned models. We offer a benchmark through which models can be compared with each other and ranked in interpretable dimensions.

Authors

Institutions

Publication Details

Journal
International Journal of Corpus Linguistics
Published
2026-09-15
DOI
https://doi.org/10.1075/ijcl.25151.mil
Primary Topic
Authorship Attribution and Profiling
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Benchmark of linguistic variation in LLM‑generated texts

Jiří Milička, Václav Cvrček, Anna Marklová
International Journal of Corpus Linguistics
Authorship Attribution and Profiling
article

Benchmark of linguistic variation in LLM‑generated texts

Jiří Milička, Václav Cvrček, Anna Marklová
article en

Abstract

Abstract This study investigates the register variation in texts written by humans and comparable texts produced by large language models (LLMs). Multidimensional analysis (MDA) is applied to a sample of human-written texts and AI-generated counterpart texts to find the dimensions of variation in which LLMs differ most significantly and most systematically from humans. We introduce the LLM-generated corpus AI-Brown, which is comparable to BE-21: a Brown family corpus representing contemporary British English. Since all languages except English are underrepresented in the training data of frontier LLMs, similar analysis is replicated on Czech language data using the AI-Koditex corpus and a Czech multidimensional model Cvrček et al. (2021) . Sixteen frontier models were examined in various settings and prompts, with emphasis placed on the difference between base models and instruction-tuned models. We offer a benchmark through which models can be compared with each other and ranked in interpretable dimensions.

International Journal of Corpus Linguistics
Charles University (CZ)
Quality Education
Openalex Percentile: Top 9%
Authorship Attribution and Profiling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Benchmark of linguistic variation in LLM‑generated texts — Jiří Milička, Václav Cvrček, et al. · International Journal of Corpus Linguistics (2026) | TGRS Research Map | TGRS