MT-brainVLM: a multi-task 3D vision–language model for hierarchical report generation with dual-path Shapley explainability

Hierarchically organised report generation from high-dimensional 3D inputs is a central yet underexplored capability of multi-task vision–language models (VLMs), obstructed by three problems: unifying report generation with structurally incompatible auxiliary heads under one weight-shared backbone, natively handling partially available modalities, and providing axiomatically grounded, numerically verifiable attributions for multi-output predictions. We propose MT-BrainVLM, an end-to-end multi-task 3D VLM coupling a modality-aware 3D U-Net encoder and mask-gated fusion, a BLIP-2-style Q-Former with three levels of learnable query tokens, a hierarchical autoregressive decoder, a disease head and an InfoNCE vision–text pathway under a joint objective, so that lesion segmentation across three incompatible taxonomies, three-level report generation and patient-level classification are emitted by one forward pass from any subset of the acquired sequences. Over it we develop a dual-path Shapley framework: data-prior KernelSHAP on 25 report-derived descriptors through a Random Forest surrogate, and model-prior exact ModalitySHAP enumerating all 2 M modality coalitions through a task-wise value multiplexer. On a held-out split of 92 cases spanning five cohorts the network attains a mean Dice of 0.325 with HD95 21.0 mm, both scored in the original voxel grid at native spacing, an impression-level BLEU-4 of 0.032 with clinical-efficacy F1 0.312, and a disease micro-F1 of 0.351, against fifteen baselines trained under an identical budget. The outcome is stated plainly: MT-BrainVLM leads none of the three tasks, trailing the strongest task-specialised segmentation network by 0.118 Dice, the expected cost of a shared backbone devoting 5.9% of its parameters to the visual path; the contribution claimed is the joint capability and the attribution pathway over it. The enumerator reproduces a closed-form ground truth to order 10 −16 and, on the trained network and real volumes, satisfies efficiency to 2.22 × 10 −16 over 120 modality attributions. We also quantify the framework’s own limits: an exact grouped (Owen) reformulation retains exactness at 2 G rather than 2 M cost; a surrogate’s predictive fidelity ( R 2 = 0.9995) does not imply agreement with the target model’s attributions (Spearman 0.60 against a 0.99 ceiling); and the same game is carried down to token, sentence, section and supervoxel granularity. Finally, much of the nominal performance on this corpus is acquisition regularity rather than image understanding: the modality-availability pattern alone identifies the source cohort on 0.565 of test cases and a cohort-majority rule reaches micro-F1 0.948 without reading a voxel. Every headline number is reported against that floor, and an interval-change experiment on 60 real two-timepoint studies is reported as the negative result it is.

Authors

Institutions

Publication Details

Journal
Journal of King Saud University - Computer and Information Sciences
Published
2026-09-30
DOI
https://doi.org/10.1007/s44443-026-01302-4
Primary Topic
Multimodal Machine Learning Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

MT-brainVLM: a multi-task 3D vision–language model for hierarchical report generation with dual-path Shapley explainability

Haitao Wu, Tingxuan Wang, Xiaodong Chen, Mingchen Xie et al.
Journal of King Saud University - Computer and Information Sciences
Multimodal Machine Learning Applications
article

MT-brainVLM: a multi-task 3D vision–language model for hierarchical report generation with dual-path Shapley explainability

Haitao Wu, Tingxuan Wang, Xiaodong Chen, Mingchen Xie, Jian Xu
article en

Abstract

Hierarchically organised report generation from high-dimensional 3D inputs is a central yet underexplored capability of multi-task vision–language models (VLMs), obstructed by three problems: unifying report generation with structurally incompatible auxiliary heads under one weight-shared backbone, natively handling partially available modalities, and providing axiomatically grounded, numerically verifiable attributions for multi-output predictions. We propose MT-BrainVLM, an end-to-end multi-task 3D VLM coupling a modality-aware 3D U-Net encoder and mask-gated fusion, a BLIP-2-style Q-Former with three levels of learnable query tokens, a hierarchical autoregressive decoder, a disease head and an InfoNCE vision–text pathway under a joint objective, so that lesion segmentation across three incompatible taxonomies, three-level report generation and patient-level classification are emitted by one forward pass from any subset of the acquired sequences. Over it we develop a dual-path Shapley framework: data-prior KernelSHAP on 25 report-derived descriptors through a Random Forest surrogate, and model-prior exact ModalitySHAP enumerating all 2 M modality coalitions through a task-wise value multiplexer. On a held-out split of 92 cases spanning five cohorts the network attains a mean Dice of 0.325 with HD95 21.0 mm, both scored in the original voxel grid at native spacing, an impression-level BLEU-4 of 0.032 with clinical-efficacy F1 0.312, and a disease micro-F1 of 0.351, against fifteen baselines trained under an identical budget. The outcome is stated plainly: MT-BrainVLM leads none of the three tasks, trailing the strongest task-specialised segmentation network by 0.118 Dice, the expected cost of a shared backbone devoting 5.9% of its parameters to the visual path; the contribution claimed is the joint capability and the attribution pathway over it. The enumerator reproduces a closed-form ground truth to order 10 −16 and, on the trained network and real volumes, satisfies efficiency to 2.22 × 10 −16 over 120 modality attributions. We also quantify the framework’s own limits: an exact grouped (Owen) reformulation retains exactness at 2 G rather than 2 M cost; a surrogate’s predictive fidelity ( R 2 = 0.9995) does not imply agreement with the target model’s attributions (Spearman 0.60 against a 0.99 ceiling); and the same game is carried down to token, sentence, section and supervoxel granularity. Finally, much of the nominal performance on this corpus is acquisition regularity rather than image understanding: the modality-availability pattern alone identifies the source cohort on 0.565 of test cases and a cohort-majority rule reaches micro-F1 0.948 without reading a voxel. Every headline number is reported against that floor, and an interval-change experiment on 60 real two-timepoint studies is reported as the negative result it is.

Journal of King Saud University - Computer and Information SciencesVol. 38(8)
Affiliated Hospital of Qingdao University (CN)
Openalex Percentile: Top 14%
Multimodal Machine Learning Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.