Cross-modal attention fusion of ultrasound imaging and clinical tabular data for breast lesion classification: A multimodal deep learning framework

Abstract Accurate classification of breast lesions in ultrasound images is an important yet challenging task, plagued by visual ambiguity between classes and morphological variability within classes. This paper presents a multimodal deep learning framework combining EfficientNet-B4 visual features with an 11-dimensional, leakage-free clinical tabular feature vector, organized into three feature-group tokens (intensity, texture, and geometry), through a multi-token cross-modal attention fusion mechanism. A systematic audit of candidate tabular features identified that eight commonly used lesion-shape descriptors (area, width, height, aspect ratio, solidity, eccentricity, and mask-region intensity statistics) are only defined when a segmentation mask exists, and their availability, not merely their value, correlates with class label; these features are excluded from the tabular branch, retaining only whole-image intensity and texture statistics defined identically for every image. On the BUSI breast ultrasound dataset (780 samples: 437 benign, 210 malignant, 133 normal) evaluated by five-fold stratified cross-validation, the tabular-only ablation using this leakage-free feature set achieves 66.72% macro AUC and 51.28% accuracy, well below both the image-only ablation (86.92% AUC) and every multimodal fusion configuration tested, indicating limited independent diagnostic signal in the retained tabular features alone. The proposed multi-token cross-modal attention mechanism achieves 95.46% macro AUC and 88.59% accuracy, comparable to, but not exceeding, a plain concatenation baseline (96.28% AUC) and two of three additional fusion strategies evaluated (additive fusion, 96.00%; bilinear fusion, 95.82%). All differences among the fusion strategies tested fall within approximately one fold-to-fold standard deviation of one another, indicating no statistically meaningful separation on this dataset at this sample size. We report the tabular-feature leakage audit as the primary methodological contribution of this study, applicable to multimodal medical imaging pipelines more broadly, and report the fusion-strategy comparison directly rather than selectively emphasizing the proposed attention mechanism.

Authors

Publication Details

Journal
Scientific Reports
Published
2026-10-09
DOI
https://doi.org/10.1038/s41598-026-72938-5
Primary Topic
AI in cancer detection
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Cross-modal attention fusion of ultrasound imaging and clinical tabular data for breast lesion classification: A multimodal deep learning framework

Shaik Javed Parvez, RadhaRani Akula, N. Vimala, K. Rakesh et al.
Scientific Reports
AI in cancer detection
article

Cross-modal attention fusion of ultrasound imaging and clinical tabular data for breast lesion classification: A multimodal deep learning framework

Shaik Javed Parvez, RadhaRani Akula, N. Vimala, K. Rakesh, Syed Husna Mehanoor
article en

Abstract

Abstract Accurate classification of breast lesions in ultrasound images is an important yet challenging task, plagued by visual ambiguity between classes and morphological variability within classes. This paper presents a multimodal deep learning framework combining EfficientNet-B4 visual features with an 11-dimensional, leakage-free clinical tabular feature vector, organized into three feature-group tokens (intensity, texture, and geometry), through a multi-token cross-modal attention fusion mechanism. A systematic audit of candidate tabular features identified that eight commonly used lesion-shape descriptors (area, width, height, aspect ratio, solidity, eccentricity, and mask-region intensity statistics) are only defined when a segmentation mask exists, and their availability, not merely their value, correlates with class label; these features are excluded from the tabular branch, retaining only whole-image intensity and texture statistics defined identically for every image. On the BUSI breast ultrasound dataset (780 samples: 437 benign, 210 malignant, 133 normal) evaluated by five-fold stratified cross-validation, the tabular-only ablation using this leakage-free feature set achieves 66.72% macro AUC and 51.28% accuracy, well below both the image-only ablation (86.92% AUC) and every multimodal fusion configuration tested, indicating limited independent diagnostic signal in the retained tabular features alone. The proposed multi-token cross-modal attention mechanism achieves 95.46% macro AUC and 88.59% accuracy, comparable to, but not exceeding, a plain concatenation baseline (96.28% AUC) and two of three additional fusion strategies evaluated (additive fusion, 96.00%; bilinear fusion, 95.82%). All differences among the fusion strategies tested fall within approximately one fold-to-fold standard deviation of one another, indicating no statistically meaningful separation on this dataset at this sample size. We report the tabular-feature leakage audit as the primary methodological contribution of this study, applicable to multimodal medical imaging pipelines more broadly, and report the fusion-strategy comparison directly rather than selectively emphasizing the proposed attention mechanism.

Scientific Reports
Openalex Percentile: Top 12%
AI in cancer detection
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.