What the Score Conceals: A Qualitative Study of LLM-Generated Metalinguistic Accounts of Variation, Norms and Acceptability in German
Research on large language models as linguistic judges has focused primarily on numerical outcomes, while their accompanying explanations remain understudied. This article qualitatively analyzes 750 explanations generated by five publicly accessible systems rating 150 contextually embedded German stimuli. The findings show that similar scores frequently conceal substantial differences in the accompanying metalinguistic accounts. Overt morphosyntactic violations generally elicit accurate diagnoses, whereas gray-zone constructions are often normalized or reduced to categorical standard/non-standard oppositions. Diatopic forms are usually accepted, but their regional classification is frequently broad or inconsistent. Register-sensitive items produce the clearest contextual reasoning, although often in formulaic terms. The explanations also reveal recurrent practices of norm invocation, repair, target substitution and sociolinguistic labeling. Although they do not transparently represent model-internal computation, they constitute behavioral data revealing the normative categories and sociolinguistic assumptions that models make available in interaction. Evaluations of LLMs as linguistic judges should therefore assess scores and explanations jointly.
Authors
- Nicholas Catasso (ORCID: https://orcid.org/0000-0002-3315-4140)
Institutions
- University of Wuppertal (DE)
Publication Details
- Journal
- AI
- Published
- 2026-09-24
- DOI
- https://doi.org/10.3390/ai7100389
- Primary Topic
- Linguistic Variation and Morphology
- Type
- article
- Field-Weighted Citation Impact
- 0.00