Clinicians’ Trust in AI-Based Clinical Recommendations Across Controlled Primary Care–Style Clinical Vignettes: Nurse-Dominant Experimental Pilot Study
BACKGROUND Safe integration of AI-enabled clinical decision support requires understanding whether users’ trust and behavioral reliance are appropriately calibrated to recommendation quality. OBJECTIVE This pilot study examined a 3-item vignette-level trust construct and accept or reject behavior across sequential primary care–style clinical vignettes. METHODS A total of 68 health care professionals, including 59 (86.8%) registered nurses and 9 (13.2%) physicians, each evaluated 21 clinical vignettes from 1 of 2 series. Each series contained 10 correct and 11 intentionally incorrect AI recommendations. Participants accepted or rejected each recommendation and responded to 3 questionnaire items measuring trust, perceived transparency, and likelihood of acting on the recommendation before receiving correctness feedback and points. The unrounded item mean formed the trust composite. In the behavioral analysis, 50 unrecorded accept or reject responses were classified as rejections. A pooled cross-classified linear mixed-effects model was constructed to assess associations between the trust composite, recommendation correctness, vignette position, and participant and vignette characteristics. It included random intercepts for participant and vignette item, with case series included as a nuisance adjustment. RESULTS Participants accepted 374 of 748 (50%) incorrect recommendations and rejected 112 of 680 (16.47%) correct recommendations. Incorrect recommendations received lower prefeedback trust than correct recommendations (β=−0.772, 95% CI −1.036 to −0.507; standardized effect=−0.419). Trust declined modestly across vignette positions (β=−0.027, 95% CI −0.048 to −0.005; standardized effect=−0.088), with a decline in case series A but not case series B. Baseline intention to use AI was positively associated with trust (β=0.391, 95% CI 0.110-0.672; standardized effect=0.243), whereas higher perceived diagnostic difficulty was negatively associated with trust (β=−0.474, 95% CI −0.648 to −0.301; standardized effect=−0.257). In the exploratory lagged model, trust in the preceding recommendation was associated with trust in the current recommendation (β=0.266, 95% CI 0.210-0.322; standardized effect=0.264). CONCLUSIONS Self-reported trust differentiated correct from incorrect recommendations in aggregate, while acceptance and rejection were not fully aligned with recommendation correctness. These preliminary findings identify a potential evaluation problem: lower trust in incorrect advice does not necessarily imply that users will reject it. The findings do not establish real-world clinical effects.
Authors
- Avishek Choudhury (ORCID: https://orcid.org/0000-0002-5342-0709)
- Yeganeh Shahsavar (ORCID: https://orcid.org/0000-0003-3422-7257)
- Ayşe P. Gürses (ORCID: https://orcid.org/0000-0001-7422-6852)
Publication Details
- Journal
- JMIR Human Factors
- Published
- 2026-10-08
- DOI
- https://doi.org/10.2196/97649
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- article
- Field-Weighted Citation Impact
- 0.00