Knowing When to Defer: Trustworthy Multimodal AI for BI-RADS-Derived Management Using Paired Mammography and Ultrasound

Background: Breast imaging ends in a management decision, from routine return to biopsy, yet most models are evaluated by accuracy alone and say nothing about when they may be wrong. We present a trust-aware multimodal model that reads paired mammography and ultrasound, recommends one of three BI-RADS-derived management actions, and returns to the radiologist the cases it cannot call. Methods: Two foundation encoders are adapted with low-rank adapters and combined by a mask-aware fusion head. A trustworthiness layer adds calibration, conformal prediction sets, selective deferral, and an atypicality flag. We evaluated our method on the Breast Cancer Multimodal Imaging Dataset (BCMID), comprising 332 cases from 323 patients at a single centre. The reference standard is a three-class management grouping we derive from the reporting radiologist’s BI-RADS assessment, so performance is in concordance with that derived label rather than with pathology, observed patient management or longitudinal clinical outcome. Results: Macro AUROC was 0.759 and balanced accuracy was 0.556. Isotonic calibration reduced calibration error from 0.088 to 0.052, and prediction sets reached an empirical coverage of 0.934 at a mean set size of 2.32. Deferring the least confident 30% by a retrospective ranking of the pooled cohort raised balanced accuracy to 0.631. Two of 63 positive-management cases were under-triaged and 15 routed to additional imaging, with a recall of 0.730; freezing the encoders left macro AUROC at 0.756 but raised the under-triage count to twelve. Conclusions: Error rate and error direction are separable properties, and neither accuracy nor macro AUROC records the direction. The system therefore pairs each recommendation with calibrated probabilities, a conformal set of plausible actions and an explicit defer option, so uncertain cases return to the radiologist. These results establish an internally validated operating profile for radiologist-facing support; clinical safety, deployment readiness and benefit to patients still require external and prospective evaluation.

Authors

Institutions

Publication Details

Journal
BioMedInformatics
Published
2026-09-15
DOI
https://doi.org/10.3390/biomedinformatics6050074
Primary Topic
AI in cancer detection
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Knowing When to Defer: Trustworthy Multimodal AI for BI-RADS-Derived Management Using Paired Mammography and Ultrasound

Ryo Haraguchi, Muhammad Nouman
BioMedInformatics
AI in cancer detection
article

Knowing When to Defer: Trustworthy Multimodal AI for BI-RADS-Derived Management Using Paired Mammography and Ultrasound

Ryo Haraguchi, Muhammad Nouman
article en

Abstract

Background: Breast imaging ends in a management decision, from routine return to biopsy, yet most models are evaluated by accuracy alone and say nothing about when they may be wrong. We present a trust-aware multimodal model that reads paired mammography and ultrasound, recommends one of three BI-RADS-derived management actions, and returns to the radiologist the cases it cannot call. Methods: Two foundation encoders are adapted with low-rank adapters and combined by a mask-aware fusion head. A trustworthiness layer adds calibration, conformal prediction sets, selective deferral, and an atypicality flag. We evaluated our method on the Breast Cancer Multimodal Imaging Dataset (BCMID), comprising 332 cases from 323 patients at a single centre. The reference standard is a three-class management grouping we derive from the reporting radiologist’s BI-RADS assessment, so performance is in concordance with that derived label rather than with pathology, observed patient management or longitudinal clinical outcome. Results: Macro AUROC was 0.759 and balanced accuracy was 0.556. Isotonic calibration reduced calibration error from 0.088 to 0.052, and prediction sets reached an empirical coverage of 0.934 at a mean set size of 2.32. Deferring the least confident 30% by a retrospective ranking of the pooled cohort raised balanced accuracy to 0.631. Two of 63 positive-management cases were under-triaged and 15 routed to additional imaging, with a recall of 0.730; freezing the encoders left macro AUROC at 0.756 but raised the under-triage count to twelve. Conclusions: Error rate and error direction are separable properties, and neither accuracy nor macro AUROC records the direction. The system therefore pairs each recommendation with calibrated probabilities, a conformal set of plausible actions and an explicit defer option, so uncertain cases return to the radiologist. These results establish an internally validated operating profile for radiologist-facing support; clinical safety, deployment readiness and benefit to patients still require external and prospective evaluation.

BioMedInformaticsVol. 6(5)
University of Hyogo (JP)
Good health and well-being
Openalex Percentile: Top 8%
AI in cancer detection
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.