Clinicians’ Trust in AI-Based Clinical Recommendations Across Controlled Primary Care–Style Clinical Vignettes: Nurse-Dominant Experimental Pilot Study

BACKGROUND Safe integration of AI-enabled clinical decision support requires understanding whether users’ trust and behavioral reliance are appropriately calibrated to recommendation quality. OBJECTIVE This pilot study examined a 3-item vignette-level trust construct and accept or reject behavior across sequential primary care–style clinical vignettes. METHODS A total of 68 health care professionals, including 59 (86.8%) registered nurses and 9 (13.2%) physicians, each evaluated 21 clinical vignettes from 1 of 2 series. Each series contained 10 correct and 11 intentionally incorrect AI recommendations. Participants accepted or rejected each recommendation and responded to 3 questionnaire items measuring trust, perceived transparency, and likelihood of acting on the recommendation before receiving correctness feedback and points. The unrounded item mean formed the trust composite. In the behavioral analysis, 50 unrecorded accept or reject responses were classified as rejections. A pooled cross-classified linear mixed-effects model was constructed to assess associations between the trust composite, recommendation correctness, vignette position, and participant and vignette characteristics. It included random intercepts for participant and vignette item, with case series included as a nuisance adjustment. RESULTS Participants accepted 374 of 748 (50%) incorrect recommendations and rejected 112 of 680 (16.47%) correct recommendations. Incorrect recommendations received lower prefeedback trust than correct recommendations (β=−0.772, 95% CI −1.036 to −0.507; standardized effect=−0.419). Trust declined modestly across vignette positions (β=−0.027, 95% CI −0.048 to −0.005; standardized effect=−0.088), with a decline in case series A but not case series B. Baseline intention to use AI was positively associated with trust (β=0.391, 95% CI 0.110-0.672; standardized effect=0.243), whereas higher perceived diagnostic difficulty was negatively associated with trust (β=−0.474, 95% CI −0.648 to −0.301; standardized effect=−0.257). In the exploratory lagged model, trust in the preceding recommendation was associated with trust in the current recommendation (β=0.266, 95% CI 0.210-0.322; standardized effect=0.264). CONCLUSIONS Self-reported trust differentiated correct from incorrect recommendations in aggregate, while acceptance and rejection were not fully aligned with recommendation correctness. These preliminary findings identify a potential evaluation problem: lower trust in incorrect advice does not necessarily imply that users will reject it. The findings do not establish real-world clinical effects.

Authors

Publication Details

Journal
JMIR Human Factors
Published
2026-10-08
DOI
https://doi.org/10.2196/97649
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Clinicians’ Trust in AI-Based Clinical Recommendations Across Controlled Primary Care–Style Clinical Vignettes: Nurse-Dominant Experimental Pilot Study

Avishek Choudhury, Yeganeh Shahsavar, Ayşe P. Gürses
JMIR Human Factors
Artificial Intelligence in Healthcare and Education
article

Clinicians’ Trust in AI-Based Clinical Recommendations Across Controlled Primary Care–Style Clinical Vignettes: Nurse-Dominant Experimental Pilot Study

Avishek Choudhury, Yeganeh Shahsavar, Ayşe P. Gürses
article en

Abstract

BACKGROUND Safe integration of AI-enabled clinical decision support requires understanding whether users’ trust and behavioral reliance are appropriately calibrated to recommendation quality. OBJECTIVE This pilot study examined a 3-item vignette-level trust construct and accept or reject behavior across sequential primary care–style clinical vignettes. METHODS A total of 68 health care professionals, including 59 (86.8%) registered nurses and 9 (13.2%) physicians, each evaluated 21 clinical vignettes from 1 of 2 series. Each series contained 10 correct and 11 intentionally incorrect AI recommendations. Participants accepted or rejected each recommendation and responded to 3 questionnaire items measuring trust, perceived transparency, and likelihood of acting on the recommendation before receiving correctness feedback and points. The unrounded item mean formed the trust composite. In the behavioral analysis, 50 unrecorded accept or reject responses were classified as rejections. A pooled cross-classified linear mixed-effects model was constructed to assess associations between the trust composite, recommendation correctness, vignette position, and participant and vignette characteristics. It included random intercepts for participant and vignette item, with case series included as a nuisance adjustment. RESULTS Participants accepted 374 of 748 (50%) incorrect recommendations and rejected 112 of 680 (16.47%) correct recommendations. Incorrect recommendations received lower prefeedback trust than correct recommendations (β=−0.772, 95% CI −1.036 to −0.507; standardized effect=−0.419). Trust declined modestly across vignette positions (β=−0.027, 95% CI −0.048 to −0.005; standardized effect=−0.088), with a decline in case series A but not case series B. Baseline intention to use AI was positively associated with trust (β=0.391, 95% CI 0.110-0.672; standardized effect=0.243), whereas higher perceived diagnostic difficulty was negatively associated with trust (β=−0.474, 95% CI −0.648 to −0.301; standardized effect=−0.257). In the exploratory lagged model, trust in the preceding recommendation was associated with trust in the current recommendation (β=0.266, 95% CI 0.210-0.322; standardized effect=0.264). CONCLUSIONS Self-reported trust differentiated correct from incorrect recommendations in aggregate, while acceptance and rejection were not fully aligned with recommendation correctness. These preliminary findings identify a potential evaluation problem: lower trust in incorrect advice does not necessarily imply that users will reject it. The findings do not establish real-world clinical effects.

JMIR Human FactorsVol. 13
Openalex Percentile: Top 19%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.