From visual saliency to anatomical explanation: generating textual diagnostic reports from brain MRIs with saliency-grounded LLMs

The opaque nature of deep learning models remains a significant barrier to their clinical adoption in medical imaging. This paper presents a systematic empirical benchmark of visual saliency attribution, atlas-based visualisation, and LLM-based reporting for brain tumour MRI classification, integrated into a unified pipeline, leveraging large language models (LLMs) to deliver human-interpretable diagnostic narratives. The proposed framework operates through three coupled stages. First, nine CNN architectures are extended with a dual-output hybrid formulation that simultaneously optimises a classification head and a segmentation head, enabling spatially richer feature learning. Second, visual saliency attribution methods, namely Grad-CAM, Grad-CAM++, and ScoreCAM, are applied to generate class-discriminative heatmaps, which are subsequently refined into coarse binary masks via an adaptive percentile thresholding pipeline. Third, the resulting masks are projected onto the Harvard–Oxford cortical atlas as an illustrative anatomical overlay, offering an approximate visual reference for the tumour’s cortical neighbourhood, and the extracted findings are encoded into a structured JSON file that conditions three LLMs (Grok3, Mistral, and LLaMA) to generate coherent, radiological-style diagnostic reports. Evaluated on a dataset of 4,834 contrast-enhanced T1-weighted brain MRI images spanning three tumour classes, InceptionResNetV2 achieved the highest classification performance, and Grad-CAM++ yielded the best segmentation overlap. Among the language models, Grok3 led in lexical diversity and coherence, while LLaMA achieved the highest readability score. By integrating visual, anatomical, and linguistic modalities into a unified pipeline, the framework attempts to produce technically grounded and meaningfully interpretable explanations, offering a step toward more transparent artificial intelligence-assisted brain tumour diagnosis.

Authors

Institutions

Publication Details

Journal
BioData Mining
Published
2026-09-14
DOI
https://doi.org/10.1186/s13040-026-00601-w
Primary Topic
Multimodal Machine Learning Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

From visual saliency to anatomical explanation: generating textual diagnostic reports from brain MRIs with saliency-grounded LLMs

Yusuf Brima, Marcellin Atemkeng, Elie Tagne Fute, Paul Valery Nguezet et al.
BioData Mining
Multimodal Machine Learning Applications
article

From visual saliency to anatomical explanation: generating textual diagnostic reports from brain MRIs with saliency-grounded LLMs

Yusuf Brima, Marcellin Atemkeng, Elie Tagne Fute, Paul Valery Nguezet, Benoit Martin Azanguezet
article en

Abstract

The opaque nature of deep learning models remains a significant barrier to their clinical adoption in medical imaging. This paper presents a systematic empirical benchmark of visual saliency attribution, atlas-based visualisation, and LLM-based reporting for brain tumour MRI classification, integrated into a unified pipeline, leveraging large language models (LLMs) to deliver human-interpretable diagnostic narratives. The proposed framework operates through three coupled stages. First, nine CNN architectures are extended with a dual-output hybrid formulation that simultaneously optimises a classification head and a segmentation head, enabling spatially richer feature learning. Second, visual saliency attribution methods, namely Grad-CAM, Grad-CAM++, and ScoreCAM, are applied to generate class-discriminative heatmaps, which are subsequently refined into coarse binary masks via an adaptive percentile thresholding pipeline. Third, the resulting masks are projected onto the Harvard–Oxford cortical atlas as an illustrative anatomical overlay, offering an approximate visual reference for the tumour’s cortical neighbourhood, and the extracted findings are encoded into a structured JSON file that conditions three LLMs (Grok3, Mistral, and LLaMA) to generate coherent, radiological-style diagnostic reports. Evaluated on a dataset of 4,834 contrast-enhanced T1-weighted brain MRI images spanning three tumour classes, InceptionResNetV2 achieved the highest classification performance, and Grad-CAM++ yielded the best segmentation overlap. Among the language models, Grok3 led in lexical diversity and coherence, while LLaMA achieved the highest readability score. By integrating visual, anatomical, and linguistic modalities into a unified pipeline, the framework attempts to produce technically grounded and meaningfully interpretable explanations, offering a step toward more transparent artificial intelligence-assisted brain tumour diagnosis.

BioData Mining
Osnabrück University (DE), Université de Dschang (CM), Rhodes University (ZA), National Institute for Theoretical Physics (ZA)
Openalex Percentile: Top 14%
Multimodal Machine Learning Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.