Diagnostic reasoning with and without AI: automation bias in pre-clerkship medical students

BACKGROUND: As large language models become increasingly integrated into clinical workflows, medical students need structured opportunities to learn how to engage critically with artificial intelligence (AI) during diagnostic reasoning. Empirical evaluation of automation bias and other risks of AI use in pre-clerkship training remains limited. METHODS: We piloted a two-component exercise for second-year pre-clerkship medical students: an introductory lecture on AI capabilities and limitations followed by a custom-built web application integrating an AI chatbot into a diagnostic reasoning case in which students ranked their differential diagnoses before and after AI access and rated the perceived influence of AI on their reasoning. Diagnostic accuracy was scored against predefined criteria. Descriptive and inferential statistics were calculated. RESULTS: In a sample of 185 students, AI use was associated with increased diagnostic accuracy (Wilcoxon signed-rank Z = -4.21, P < .001), with the greatest benefit concentrated among students with lower performance before AI use. Both students whose accuracy improved (n = 65) and those whose accuracy worsened (n = 26) after AI use rated AI as more influential than students whose accuracy did not change (Z = -2.74, P = .006 and Z = -3.23, P = .001, respectively), suggesting that perceived influence was related to whether AI changed students' rankings, regardless of whether the change was beneficial or detrimental. CONCLUSIONS: Students benefited most from AI when their baseline diagnostic accuracy was weakest but did not appear to reliably distinguish helpful from harmful AI influence, consistent with automation bias. Foundational coursework on AI and its limitations, paired with AI-integrated case practice and faculty-led debriefing, offers one training approach to address this challenge.

Authors

Institutions

Publication Details

Journal
BMC Medical Education
Published
2026-07-21
DOI
https://doi.org/10.1186/s12909-026-09979-x
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Diagnostic reasoning with and without AI: automation bias in pre-clerkship medical students

Saradha Ramesh, Michael R. Blanco, Robert Trowbridge, Robert Hayden et al.
BMC Medical Education
Artificial Intelligence in Healthcare and Education
article

Diagnostic reasoning with and without AI: automation bias in pre-clerkship medical students

Saradha Ramesh, Michael R. Blanco, Robert Trowbridge, Robert Hayden, Scott Epstein
article en

Abstract

BACKGROUND: As large language models become increasingly integrated into clinical workflows, medical students need structured opportunities to learn how to engage critically with artificial intelligence (AI) during diagnostic reasoning. Empirical evaluation of automation bias and other risks of AI use in pre-clerkship training remains limited. METHODS: We piloted a two-component exercise for second-year pre-clerkship medical students: an introductory lecture on AI capabilities and limitations followed by a custom-built web application integrating an AI chatbot into a diagnostic reasoning case in which students ranked their differential diagnoses before and after AI access and rated the perceived influence of AI on their reasoning. Diagnostic accuracy was scored against predefined criteria. Descriptive and inferential statistics were calculated. RESULTS: In a sample of 185 students, AI use was associated with increased diagnostic accuracy (Wilcoxon signed-rank Z = -4.21, P < .001), with the greatest benefit concentrated among students with lower performance before AI use. Both students whose accuracy improved (n = 65) and those whose accuracy worsened (n = 26) after AI use rated AI as more influential than students whose accuracy did not change (Z = -2.74, P = .006 and Z = -3.23, P = .001, respectively), suggesting that perceived influence was related to whether AI changed students' rankings, regardless of whether the change was beneficial or detrimental. CONCLUSIONS: Students benefited most from AI when their baseline diagnostic accuracy was weakest but did not appear to reliably distinguish helpful from harmful AI influence, consistent with automation bias. Foundational coursework on AI and its limitations, paired with AI-integrated case practice and faculty-led debriefing, offers one training approach to address this challenge.

BMC Medical Education
Tufts University (US), Maine Medical Center (US), MaineHealth (US)
Quality Education
Openalex Percentile: Top 12%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.