What the Score Conceals: A Qualitative Study of LLM-Generated Metalinguistic Accounts of Variation, Norms and Acceptability in German

Research on large language models as linguistic judges has focused primarily on numerical outcomes, while their accompanying explanations remain understudied. This article qualitatively analyzes 750 explanations generated by five publicly accessible systems rating 150 contextually embedded German stimuli. The findings show that similar scores frequently conceal substantial differences in the accompanying metalinguistic accounts. Overt morphosyntactic violations generally elicit accurate diagnoses, whereas gray-zone constructions are often normalized or reduced to categorical standard/non-standard oppositions. Diatopic forms are usually accepted, but their regional classification is frequently broad or inconsistent. Register-sensitive items produce the clearest contextual reasoning, although often in formulaic terms. The explanations also reveal recurrent practices of norm invocation, repair, target substitution and sociolinguistic labeling. Although they do not transparently represent model-internal computation, they constitute behavioral data revealing the normative categories and sociolinguistic assumptions that models make available in interaction. Evaluations of LLMs as linguistic judges should therefore assess scores and explanations jointly.

Authors

Institutions

Publication Details

Journal
AI
Published
2026-09-24
DOI
https://doi.org/10.3390/ai7100389
Primary Topic
Linguistic Variation and Morphology
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

What the Score Conceals: A Qualitative Study of LLM-Generated Metalinguistic Accounts of Variation, Norms and Acceptability in German

Nicholas Catasso
AI
Linguistic Variation and Morphology
article

What the Score Conceals: A Qualitative Study of LLM-Generated Metalinguistic Accounts of Variation, Norms and Acceptability in German

Nicholas Catasso
article en

Abstract

Research on large language models as linguistic judges has focused primarily on numerical outcomes, while their accompanying explanations remain understudied. This article qualitatively analyzes 750 explanations generated by five publicly accessible systems rating 150 contextually embedded German stimuli. The findings show that similar scores frequently conceal substantial differences in the accompanying metalinguistic accounts. Overt morphosyntactic violations generally elicit accurate diagnoses, whereas gray-zone constructions are often normalized or reduced to categorical standard/non-standard oppositions. Diatopic forms are usually accepted, but their regional classification is frequently broad or inconsistent. Register-sensitive items produce the clearest contextual reasoning, although often in formulaic terms. The explanations also reveal recurrent practices of norm invocation, repair, target substitution and sociolinguistic labeling. Although they do not transparently represent model-internal computation, they constitute behavioral data revealing the normative categories and sociolinguistic assumptions that models make available in interaction. Evaluations of LLMs as linguistic judges should therefore assess scores and explanations jointly.

AIVol. 7(10)
University of Wuppertal (DE)
Quality Education
Openalex Percentile: Top 4%
Linguistic Variation and Morphology
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

What the Score Conceals: A Qualitative Study of LLM-Generated Metalinguistic Accounts of Variation, Norms and Acceptability in German — Nicholas Catasso · AI (2026) | TGRS Research Map | TGRS