When Classification Errors Preserve Correct Recommendations: The Structural Mechanism and Effects of Inadvertent Safeguards in AI Decision Support
can occur. An inadvertent safeguard refers to a counterintuitive scenario in which an AI error at the diagnosis/classification stage nevertheless produces a correct downstream decision recommendation. In this paper, we analyze the structural relationship between the AI's classification space and its recommendation outputs. We formalize the conditions under which inadvertent safeguards arise and compare the performance implications of Stage 3 action-selection support versus Stage 2 information-analysis support.BackgroundRecent research in human-automation and human-AI interaction has allowed for the structural possibility of inadvertent safeguards. However, this phenomenon has not been explicitly theorized or systematically examined.MethodWe first characterized the structural conditions under which inadvertent safeguards arise. We then conducted an experiment in which 90 participants completed a mental rotation task with a simulated AI decision aid that provided either Stage 2 information-analysis support or Stage 3 action-selection support.ResultsWhen inadvertent safeguards occur, Stage 3 action-selection support produced better performance, faster responses, and more positive trust adjustment than Stage 2 information-analysis support. Overall, Stage 3 action-selection support also produced higher perceived reliability, post-condition trust, and usability ratings.ConclusionThe impact of AI errors depends not only on classification accuracy but also on the mapping structure between classifications and actions. Many-to-one mappings, together with misclassifications between recommendation-equivalent states, can preserve correct downstream recommendations despite classification errors.ApplicationDesigners should distinguish classification accuracy from recommendation accuracy when evaluating Stage 3 action-selection support, especially in domains where multiple states can map to the same action.
Authors
- X. Jessie Yang (ORCID: https://orcid.org/0000-0001-6071-0387)
- Jin Yong Kim (ORCID: https://orcid.org/0009-0002-1626-7204)
Institutions
- University of Michigan (US)
Publication Details
- Journal
- Human Factors The Journal of the Human Factors and Ergonomics Society
- Published
- 2026-10-08
- DOI
- https://doi.org/10.1177/00187208261491447
- Primary Topic
- Human-Automation Interaction and Safety
- Type
- article
- Field-Weighted Citation Impact
- 0.00