Classifying 25 misinterpretations of statistical tests: a comparison of six large language models

Misinterpretations of statistical tests remain widespread and can be amplified by tools increasingly used to support scientific reasoning, including large language models (LLMs). This study evaluates whether LLMs endorse or reproduce well documented interpretive errors when asked to assess benchmark statements about frequentist inference. We used a fixed benchmark of 25 statements incorrect by construction. Six LLM configurations accessed via their user interfaces were tested under two prompting conditions (a discursive prompt requesting correctness judgments and corrections; a minimal prompt requesting only correct/incorrect labels). To assess within-model variability and prompt dependence, we repeated prompts across separate chat sessions under two administration formats: batch submission of all 25 statements and single-item submission. Summary outcomes captured whether the model rejected or endorsed each incorrect statement. Expanded outcomes were assessed via structured textual analysis of explanations using an a priori glossary of fallacy categories. Across configurations, most incorrect benchmark statements were correctly rejected at the label level; however, both sporadic errors and systematic misclassifications were frequently observed. Under batch administration, erroneous endorsements concentrated in a subset of recurrent failure items and varied across models and prompting conditions. Under single-item administration, erroneous endorsements were markedly reduced, with some configurations producing no classification errors. However, textual analysis revealed that outputs correctly rejecting the benchmark claim often introduced additional fallacies. The most recurrent patterns included subordination of assumptions, overconfident interpretations of interval estimates, mixing of inferential logics, ritualistic reliance on thresholds, null privileging, oversimplification, and unqualified use of dichotomizing language. In this study, conducted between December 2025 and January 2026, LLM responses to statistical-testing misinterpretations depended on discourse constraints and prompting. Single-item administration yielded higher label accuracy, but explanatory text often introduced misleading inferential rhetoric. In light of these findings and causal knowledge about LLM generation, caution is warranted when using LLM explanations for methodological guidance. Therefore, LLM-generated explanations should not be relied upon for methodological guidance without critical evaluation by qualified human experts. This evaluation should explicitly account for the fact that LLMs are likely to reproduce not only machine-generated errors, but historical human distortions already embedded and normalized in statistical language and scientific communication. The transportability of these conclusions to other model versions requires separate evaluation.

Authors

Institutions

Publication Details

Journal
BMC Medical Research Methodology
Published
2026-09-14
DOI
https://doi.org/10.1186/s12874-026-03005-w
Primary Topic
Explainable Artificial Intelligence (XAI)
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Classifying 25 misinterpretations of statistical tests: a comparison of six large language models

Mohammad Alì Mansournia, Alessandro Rovetta, Lucia Castaldo
BMC Medical Research Methodology
Explainable Artificial Intelligence (XAI)
article

Classifying 25 misinterpretations of statistical tests: a comparison of six large language models

Mohammad Alì Mansournia, Alessandro Rovetta, Lucia Castaldo
article en

Abstract

Misinterpretations of statistical tests remain widespread and can be amplified by tools increasingly used to support scientific reasoning, including large language models (LLMs). This study evaluates whether LLMs endorse or reproduce well documented interpretive errors when asked to assess benchmark statements about frequentist inference. We used a fixed benchmark of 25 statements incorrect by construction. Six LLM configurations accessed via their user interfaces were tested under two prompting conditions (a discursive prompt requesting correctness judgments and corrections; a minimal prompt requesting only correct/incorrect labels). To assess within-model variability and prompt dependence, we repeated prompts across separate chat sessions under two administration formats: batch submission of all 25 statements and single-item submission. Summary outcomes captured whether the model rejected or endorsed each incorrect statement. Expanded outcomes were assessed via structured textual analysis of explanations using an a priori glossary of fallacy categories. Across configurations, most incorrect benchmark statements were correctly rejected at the label level; however, both sporadic errors and systematic misclassifications were frequently observed. Under batch administration, erroneous endorsements concentrated in a subset of recurrent failure items and varied across models and prompting conditions. Under single-item administration, erroneous endorsements were markedly reduced, with some configurations producing no classification errors. However, textual analysis revealed that outputs correctly rejecting the benchmark claim often introduced additional fallacies. The most recurrent patterns included subordination of assumptions, overconfident interpretations of interval estimates, mixing of inferential logics, ritualistic reliance on thresholds, null privileging, oversimplification, and unqualified use of dichotomizing language. In this study, conducted between December 2025 and January 2026, LLM responses to statistical-testing misinterpretations depended on discourse constraints and prompting. Single-item administration yielded higher label accuracy, but explanatory text often introduced misleading inferential rhetoric. In light of these findings and causal knowledge about LLM generation, caution is warranted when using LLM explanations for methodological guidance. Therefore, LLM-generated explanations should not be relied upon for methodological guidance without critical evaluation by qualified human experts. This evaluation should explicitly account for the fact that LLMs are likely to reproduce not only machine-generated errors, but historical human distortions already embedded and normalized in statistical language and scientific communication. The transportability of these conclusions to other model versions requires separate evaluation.

BMC Medical Research Methodology
Istituto Nazionale di Fisica Nucleare, Galileo Galilei Institute for Theoretical Physics (IT), Redwood Scientific (United States) (US), Statistical Service (CY), Tehran University of Medical Sciences (IR)
Peace, Justice and strong institutions
Openalex Percentile: Top 8%
Explainable Artificial Intelligence (XAI)
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.