Transformer-Based Multimodal Fusion Integrating Computed Tomography, Pathology, and Clinical Features for Oral Squamous Cell Carcinoma Recurrence Prediction

Background Postoperative recurrence is the leading cause of treatment failure in oral squamous cell carcinoma (OSCC), yet current risk-stratification schemes offer only modest predictive performance. We developed a Transformer-based multimodal fusion framework integrating computed tomography (CT), histopathology, and clinical variables for individualized recurrence prediction. Methods A total of 134 patients with OSCC were retrospectively enrolled. Using a dual-stream strategy, deep semantic representations were derived from preoperative contrast-enhanced CT and haematoxylin and eosin-stained histopathological images via MnasNet-0.5, while quantitative radiomic descriptors were extracted using PyRadiomics. Together with structured clinical variables, these constituted five parallel feature streams, which – after dimensional harmonization – were passed through a three-layer Transformer encoder employing self-attention and learnable modality weights for nonlinear cross-modal integration. Performance was assessed by fivefold stratified cross-validation. Results The full-modality TransformerFusion model achieved an AUC of 0.909 ± 0.069, with a sensitivity of 0.791, specificity of 0.834, and precision of 0.854. Decision curve analysis demonstrated superior net clinical benefit over logistic regression and the treat-none strategy across threshold probabilities of 0.20 to 0.70. Ablation experiments revealed that the pathology dual-stream, CT dual-stream, and clinical-only submodels attained AUCs of 0.840, 0.751, and 0.773, respectively – each inferior to the integrated model. Self-attention-based fusion (AUC = 0.909) also outperformed feature concatenation (AUC = 0.854). Conclusion The proposed framework captures complementary prognostic signals across CT, histopathology, and clinical variables, outperforming single-modality counterparts in recurrence prediction. As it relies solely on routine preoperative CT and standard haematoxylin and eosin slides, it offers a clinically deployable tool for postoperative risk stratification in OSCC.

Authors

Institutions

Publication Details

Journal
International Dental Journal
Published
2026-09-18
DOI
https://doi.org/10.1016/j.identj.2026.111098
Primary Topic
Radiomics and Machine Learning in Medical Imaging
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Transformer-Based Multimodal Fusion Integrating Computed Tomography, Pathology, and Clinical Features for Oral Squamous Cell Carcinoma Recurrence Prediction

Binyang Fu, Zhuang Liang, Hui Dong, Shuwen Yang et al.
International Dental Journal
Radiomics and Machine Learning in Medical Imaging
article

Transformer-Based Multimodal Fusion Integrating Computed Tomography, Pathology, and Clinical Features for Oral Squamous Cell Carcinoma Recurrence Prediction

Binyang Fu, Zhuang Liang, Hui Dong, Shuwen Yang, Fuyong Sui, Zhen Wang, Qilin Liu
article en

Abstract

Background Postoperative recurrence is the leading cause of treatment failure in oral squamous cell carcinoma (OSCC), yet current risk-stratification schemes offer only modest predictive performance. We developed a Transformer-based multimodal fusion framework integrating computed tomography (CT), histopathology, and clinical variables for individualized recurrence prediction. Methods A total of 134 patients with OSCC were retrospectively enrolled. Using a dual-stream strategy, deep semantic representations were derived from preoperative contrast-enhanced CT and haematoxylin and eosin-stained histopathological images via MnasNet-0.5, while quantitative radiomic descriptors were extracted using PyRadiomics. Together with structured clinical variables, these constituted five parallel feature streams, which – after dimensional harmonization – were passed through a three-layer Transformer encoder employing self-attention and learnable modality weights for nonlinear cross-modal integration. Performance was assessed by fivefold stratified cross-validation. Results The full-modality TransformerFusion model achieved an AUC of 0.909 ± 0.069, with a sensitivity of 0.791, specificity of 0.834, and precision of 0.854. Decision curve analysis demonstrated superior net clinical benefit over logistic regression and the treat-none strategy across threshold probabilities of 0.20 to 0.70. Ablation experiments revealed that the pathology dual-stream, CT dual-stream, and clinical-only submodels attained AUCs of 0.840, 0.751, and 0.773, respectively – each inferior to the integrated model. Self-attention-based fusion (AUC = 0.909) also outperformed feature concatenation (AUC = 0.854). Conclusion The proposed framework captures complementary prognostic signals across CT, histopathology, and clinical variables, outperforming single-modality counterparts in recurrence prediction. As it relies solely on routine preoperative CT and standard haematoxylin and eosin slides, it offers a clinically deployable tool for postoperative risk stratification in OSCC.

International Dental JournalVol. 76(6)
Dalian Medical University (CN), Second Affiliated Hospital of Dalian Medical University (CN), First Affiliated Hospital of Dalian Medical University (CN)
Openalex Percentile: Top 11%
Radiomics and Machine Learning in Medical Imaging
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.