MT-brainVLM: a multi-task 3D vision–language model for hierarchical report generation with dual-path Shapley explainability
Hierarchically organised report generation from high-dimensional 3D inputs is a central yet underexplored capability of multi-task vision–language models (VLMs), obstructed by three problems: unifying report generation with structurally incompatible auxiliary heads under one weight-shared backbone, natively handling partially available modalities, and providing axiomatically grounded, numerically verifiable attributions for multi-output predictions. We propose MT-BrainVLM, an end-to-end multi-task 3D VLM coupling a modality-aware 3D U-Net encoder and mask-gated fusion, a BLIP-2-style Q-Former with three levels of learnable query tokens, a hierarchical autoregressive decoder, a disease head and an InfoNCE vision–text pathway under a joint objective, so that lesion segmentation across three incompatible taxonomies, three-level report generation and patient-level classification are emitted by one forward pass from any subset of the acquired sequences. Over it we develop a dual-path Shapley framework: data-prior KernelSHAP on 25 report-derived descriptors through a Random Forest surrogate, and model-prior exact ModalitySHAP enumerating all 2 M modality coalitions through a task-wise value multiplexer. On a held-out split of 92 cases spanning five cohorts the network attains a mean Dice of 0.325 with HD95 21.0 mm, both scored in the original voxel grid at native spacing, an impression-level BLEU-4 of 0.032 with clinical-efficacy F1 0.312, and a disease micro-F1 of 0.351, against fifteen baselines trained under an identical budget. The outcome is stated plainly: MT-BrainVLM leads none of the three tasks, trailing the strongest task-specialised segmentation network by 0.118 Dice, the expected cost of a shared backbone devoting 5.9% of its parameters to the visual path; the contribution claimed is the joint capability and the attribution pathway over it. The enumerator reproduces a closed-form ground truth to order 10 −16 and, on the trained network and real volumes, satisfies efficiency to 2.22 × 10 −16 over 120 modality attributions. We also quantify the framework’s own limits: an exact grouped (Owen) reformulation retains exactness at 2 G rather than 2 M cost; a surrogate’s predictive fidelity ( R 2 = 0.9995) does not imply agreement with the target model’s attributions (Spearman 0.60 against a 0.99 ceiling); and the same game is carried down to token, sentence, section and supervoxel granularity. Finally, much of the nominal performance on this corpus is acquisition regularity rather than image understanding: the modality-availability pattern alone identifies the source cohort on 0.565 of test cases and a cohort-majority rule reaches micro-F1 0.948 without reading a voxel. Every headline number is reported against that floor, and an interval-change experiment on 60 real two-timepoint studies is reported as the negative result it is.
Authors
- Haitao Wu (ORCID: https://orcid.org/0000-0003-4111-9211)
- Tingxuan Wang
- Xiaodong Chen
- Mingchen Xie (ORCID: https://orcid.org/0000-0002-4109-933X)
- Jian Xu
Institutions
- Affiliated Hospital of Qingdao University (CN)
Publication Details
- Journal
- Journal of King Saud University - Computer and Information Sciences
- Published
- 2026-09-30
- DOI
- https://doi.org/10.1007/s44443-026-01302-4
- Primary Topic
- Multimodal Machine Learning Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00