Classifying 25 misinterpretations of statistical tests: a comparison of six large language models
Misinterpretations of statistical tests remain widespread and can be amplified by tools increasingly used to support scientific reasoning, including large language models (LLMs). This study evaluates whether LLMs endorse or reproduce well documented interpretive errors when asked to assess benchmark statements about frequentist inference. We used a fixed benchmark of 25 statements incorrect by construction. Six LLM configurations accessed via their user interfaces were tested under two prompting conditions (a discursive prompt requesting correctness judgments and corrections; a minimal prompt requesting only correct/incorrect labels). To assess within-model variability and prompt dependence, we repeated prompts across separate chat sessions under two administration formats: batch submission of all 25 statements and single-item submission. Summary outcomes captured whether the model rejected or endorsed each incorrect statement. Expanded outcomes were assessed via structured textual analysis of explanations using an a priori glossary of fallacy categories. Across configurations, most incorrect benchmark statements were correctly rejected at the label level; however, both sporadic errors and systematic misclassifications were frequently observed. Under batch administration, erroneous endorsements concentrated in a subset of recurrent failure items and varied across models and prompting conditions. Under single-item administration, erroneous endorsements were markedly reduced, with some configurations producing no classification errors. However, textual analysis revealed that outputs correctly rejecting the benchmark claim often introduced additional fallacies. The most recurrent patterns included subordination of assumptions, overconfident interpretations of interval estimates, mixing of inferential logics, ritualistic reliance on thresholds, null privileging, oversimplification, and unqualified use of dichotomizing language. In this study, conducted between December 2025 and January 2026, LLM responses to statistical-testing misinterpretations depended on discourse constraints and prompting. Single-item administration yielded higher label accuracy, but explanatory text often introduced misleading inferential rhetoric. In light of these findings and causal knowledge about LLM generation, caution is warranted when using LLM explanations for methodological guidance. Therefore, LLM-generated explanations should not be relied upon for methodological guidance without critical evaluation by qualified human experts. This evaluation should explicitly account for the fact that LLMs are likely to reproduce not only machine-generated errors, but historical human distortions already embedded and normalized in statistical language and scientific communication. The transportability of these conclusions to other model versions requires separate evaluation.
Authors
- Mohammad Alì Mansournia (ORCID: https://orcid.org/0000-0003-3343-2718)
- Alessandro Rovetta (ORCID: https://orcid.org/0000-0002-4634-279X)
- Lucia Castaldo
Institutions
- Istituto Nazionale di Fisica Nucleare, Galileo Galilei Institute for Theoretical Physics (IT)
- Redwood Scientific (United States) (US)
- Statistical Service (CY)
- Tehran University of Medical Sciences (IR)
Publication Details
- Journal
- BMC Medical Research Methodology
- Published
- 2026-09-14
- DOI
- https://doi.org/10.1186/s12874-026-03005-w
- Primary Topic
- Explainable Artificial Intelligence (XAI)
- Type
- article
- Field-Weighted Citation Impact
- 0.00