Impact of automation bias on artificial intelligence-assisted bone age assessment: a randomized crossover study

Integrating artificial intelligence (AI) into clinical practice introduces the risk of automation bias, whereby clinicians may over-rely on AI outputs even when they are incorrect. This study examined the impact of automation bias on AI-assisted bone age assessment among radiologists with different levels of seniority. This exploratory randomized crossover study included six radiologists representing three levels of seniority. Each radiologist assessed bone age in 200 radiographs with either accurate AI assistance (true-AI) or deliberately altered AI outputs (sham-AI). The sham-AI outputs were generated by randomizing the true-AI predictions. Radiologists also indicated whether they trusted or distrusted the provided AI predictions. The intra-class correlation coefficient (ICC) for repeated measures was calculated to evaluate reliability. The mean absolute difference (MAD) between each radiologist’s AI-assisted assessment and the ground-truth bone age was calculated to evaluate accuracy. Paired ttests were used to compare the MAD between the true-AI and sham-AI conditions. The radiologist (M1) with the highest ICC and lowest MAD in the unassisted condition was not affected by discrepancies between true and sham AI outputs ( P = 0.297). All other radiologists were influenced by AI accuracy, exhibiting significant differences in MAD between the true-AI and sham-AI assistance ( P < 0.001). The radiologist (J2) with the lowest ICC and highest MAD showed consistently high trust in sham-AI predictions. Senior radiologists (S1 and S2) displayed higher rates of disagreement with AI outputs than other participants, even when the AI predictions were accurate. When the discrepancy between true and sham AI predictions exceeded six months, the MAD of bone age assessments increased significantly under the sham-AI assistance, particularly among junior radiologists. Radiologists with higher baseline reliability and diagnostic accuracy appeared to be less susceptible to automation bias. The greater susceptibility observed in the junior radiologist was an exploratory finding that requires validation in larger studies.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-09-28
DOI
https://doi.org/10.1038/s41598-026-73218-y
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Impact of automation bias on artificial intelligence-assisted bone age assessment: a randomized crossover study

Yen-Huai Lin, Tien-Yu Chang, Yeong-Seng Yuh
Scientific Reports
Artificial Intelligence in Healthcare and Education
article

Impact of automation bias on artificial intelligence-assisted bone age assessment: a randomized crossover study

Yen-Huai Lin, Tien-Yu Chang, Yeong-Seng Yuh
article en

Abstract

Integrating artificial intelligence (AI) into clinical practice introduces the risk of automation bias, whereby clinicians may over-rely on AI outputs even when they are incorrect. This study examined the impact of automation bias on AI-assisted bone age assessment among radiologists with different levels of seniority. This exploratory randomized crossover study included six radiologists representing three levels of seniority. Each radiologist assessed bone age in 200 radiographs with either accurate AI assistance (true-AI) or deliberately altered AI outputs (sham-AI). The sham-AI outputs were generated by randomizing the true-AI predictions. Radiologists also indicated whether they trusted or distrusted the provided AI predictions. The intra-class correlation coefficient (ICC) for repeated measures was calculated to evaluate reliability. The mean absolute difference (MAD) between each radiologist’s AI-assisted assessment and the ground-truth bone age was calculated to evaluate accuracy. Paired ttests were used to compare the MAD between the true-AI and sham-AI conditions. The radiologist (M1) with the highest ICC and lowest MAD in the unassisted condition was not affected by discrepancies between true and sham AI outputs ( P = 0.297). All other radiologists were influenced by AI accuracy, exhibiting significant differences in MAD between the true-AI and sham-AI assistance ( P < 0.001). The radiologist (J2) with the lowest ICC and highest MAD showed consistently high trust in sham-AI predictions. Senior radiologists (S1 and S2) displayed higher rates of disagreement with AI outputs than other participants, even when the AI predictions were accurate. When the discrepancy between true and sham AI predictions exceeded six months, the MAD of bone age assessments increased significantly under the sham-AI assistance, particularly among junior radiologists. Radiologists with higher baseline reliability and diagnostic accuracy appeared to be less susceptible to automation bias. The greater susceptibility observed in the junior radiologist was an exploratory finding that requires validation in larger studies.

Scientific Reports
National Yang Ming Chiao Tung University (TW), Cheng Hsin General Hospital (TW)
Quality Education
Openalex Percentile: Top 15%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.