An Explainable Multimodal Deep Learning Framework for Alzheimer's Disease Classification Using MRI, PET, and Clinical Data

{"Abstract":[0],"Alzheimer's":[1,68,226,617,628,765,835,856,926],"disease":[2,227,299,328,381,618,678,766,804,836,857,904,927],"(AD)":[3,228],"often":[4,993],"remains":[5,30,235],"undetected":[6],"until":[7,254],"memory":[8],"loss":[9,376],"and":[10,26,62,79,86,144,161,171,183,205,234,250,280,290,330,352,373,411,429,521,544,565,591,623,639,670,684,727,761,780,799,822,864,889,928,956,968,1005,1097,1151,1165,1196,1214,1307],"cognitive":[11,305,408,1193,1215],"decline":[12,991],"become":[13,262,313],"clinically":[14],"apparent,":[15],"at":[16,485],"which":[17,391,1161],"point":[18],"substantial":[19],"neurological":[20],"damage":[21],"has":[22,312,458,501,847,958,1183],"already":[23],"occurred.":[24],"Early":[25],"reliable":[27,204],"diagnosis":[28],"therefore,":[29],"one":[31,314,556,1033],"of":[32,147,196,240,297,315,335,419,437,600,734,745,772,783,807,812,903,947,1034,1046,1071,1138,1181,1187,1203,1225,1289],"the":[33,67,116,128,142,154,194,200,236,273,283,294,298,350,358,363,380,385,438,464,471,487,545,560,597,627,635,650,700,708,732,738,778,813,818,850,894,900,914,936,1015,1035,1039,1065,1119,1170,1185,1290,1301,1314],"most":[34,119,533,705,1137],"pressing":[35,318],"challenges":[36],"in":[37,121,357,480,935,944,1128,1229,1287],"dementia":[38,241],"care.":[39],"This":[40],"study":[41,133,714,771,1133,1247,1316],"presents":[42,610],"a":[43,87,100,230,332,416,434,444,540,586,665,932,962,1178,1244,1256],"multimodal,":[44],"explainable":[45,220,612],"deep":[46,212,1263],"learning":[47,457,486,500,647,978],"framework":[48,615,756,1040],"for":[49,203,337,365,473,616,688,764,802,854,919,987,1023],"AD":[50,197,277,340,492,538],"classification":[51,619,709,740,820,837,1026],"that":[52,153,179,257,377,421,490,547,592,703,757,793,816,830,897,965,992,1073,1110,1212,1221,1248],"integrates":[53,758],"Magnetic":[54],"Resonance":[55],"Imaging":[56,73],"(MRI),":[57],"Positron":[58,950],"Emission":[59,951],"Tomography":[60,952],"(PET),":[61,953],"clinical":[63,96,172,184,207,216,324,430,598,624,671,689,762,800,1198,1213,1239,1309],"data":[64,149,431,785,832,1150,1164,1240],"obtained":[65],"from":[66,77,493,514,539,626,872,925,1121],"Disease":[69,629],"Neuroimaging":[70,630],"Initiative":[71,631],"(ADNI).":[72],"features":[74,660,672],"are":[75,174,307,642,661,673,1222,1266],"derived":[76],"MRI":[78,361,638,722,846,939,1004,1050,1230],"PET":[80,640,724,982,1007,1047],"scans":[81,874],"using":[82,620,645,1009],"ResNet50-based":[83],"transfer":[84,646],"learning,":[85,213,219],"patient-level":[88,666,788],"fusion":[89,156,668,791,833,1092,1113,1147],"network":[90],"combines":[91,795],"these":[92,606,1139],"imaging":[93,170,182,422,542,659,852,911,975,1075,1112,1149,1242],"representations":[94],"with":[95,164,276,281,581,840,1028,1064,1241],"variables":[97,1216],"to":[98,114,126,140,264,301,527,559,595,653,675,698,707,729,776,981,1020,1038,1049,1093,1124,1148,1169,1299,1318],"generate":[99],"final":[101,561],"diagnostic":[102,189,602,916,1078,1219],"prediction.":[103,679,805],"To":[104,680],"ensure":[105],"interpretability,":[106],"Gradient-weighted":[107,691,1278],"Class":[108,692,1279],"Activation":[109,693,1280],"Mapping":[110,694,1281],"(Grad-CAM)":[111,695,1282],"is":[112,229,362,569,696,715,1062,1227,1255],"applied":[113,980],"visualize":[115,1300],"brain":[117,497,701,814,1302],"regions":[118,702,815,1303],"influential":[120],"each":[122,148,432,555,735],"prediction,":[123],"allowing":[124],"clinicians":[125,393,580],"inspect":[127],"model's":[129,201,819],"reasoning.":[130],"An":[131],"ablation":[132,713,1132,1246],"comprising":[134],"four":[135,719,773],"experimental":[136,720,774,827],"configurations":[137,775],"was":[138,589],"conducted":[139,716],"evaluate":[141,777],"individual":[143,779,1251],"combined":[145,663,781],"contributions":[146,744,782],"modality.":[150],"Results":[151],"show":[152,401],"multimodal":[155,218,613,667,754,789,831,1111],"model":[157,682,824,1305],"consistently":[158,1108],"outperforms":[159,1114],"unimodal":[160,938,1056],"bimodal":[162],"baselines,":[163],"predictive":[165],"accuracy":[166,190,1022],"improving":[167,522,823],"progressively":[168],"as":[169,379,508,575,751,849,961,1269],"information":[173,625,763,801,970,1099],"integrated.":[175],"These":[176,658,1106,1209],"findings":[177],"demonstrate":[178],"combining":[180],"complementary":[181,796,963,1095],"signals":[185,1079,1220],"not":[186,891,971,1237],"only":[187],"enhances":[188,834],"but":[191,1233],"also":[192],"improves":[193],"interpretability":[195,1254],"classification,":[198],"supporting":[199],"potential":[202],"transparent":[206],"decision":[208],"support.":[209],"Keywords:Alzheimer's":[210],"disease,":[211,439],"MRI,":[214,427,621,759],"PET,":[215,383,428,622,760,1232],"data,":[217],"artificial":[221],"intelligence,":[222],"Grad-CAM,":[223],"ResNet50":[224,509,651],"1.Introduction":[225],"progressive":[231],"neurodegenerative":[232],"condition":[233],"single":[237,449,541,1245],"largest":[238],"cause":[239],"globally.":[242],"Its":[243],"course":[244],"typically":[245],"erodes":[246],"memory,":[247],"language,":[248],"reasoning,":[249],"decision-making":[251],"over":[252,463],"time,":[253],"everyday":[255],"tasks":[256],"once":[258],"required":[259],"no":[260],"thought":[261],"difficult":[263],"manage":[265],"without":[266,875],"help.":[267],"Rising":[268],"life":[269,336],"expectancy":[270],"worldwide":[271],"means":[272],"population":[274],"living":[275],"keeps":[278],"expanding,":[279],"it":[282,310,347,1032],"strain":[284],"placed":[285],"on":[286,343,384,396,453,649,909,973,1014,1080,1084,1153],"patients,":[287],"families,":[288],"caregivers,":[289],"health":[291],"systems.":[292],"Because":[293,426,1262],"underlying":[295],"biology":[296],"tends":[300],"shift":[302],"years":[303],"before":[304,399],"symptoms":[306],"noticeable,":[308],"catching":[309],"early":[311,901,989],"healthcare's":[316],"more":[317],"priorities.":[319],"Earlier":[320],"intervention":[321],"can":[322,748,867,1217],"slow":[323,596],"decline,":[325],"guide":[326],"better":[327],"management,":[329],"preserve":[331],"patient's":[333],"quality":[334],"longer.":[338],"Diagnosing":[339],"leans":[341],"heavily":[342],"medical":[344,413,460,529],"imaging,":[345,1177],"since":[346,859],"reveals":[348],"both":[349,1163],"structural":[351,910,974,995,1096],"functional":[353,397,967,1074,1098],"changes":[354,888],"taking":[355],"place":[356],"brain.":[359],"Structural":[360,845],"workhorse":[364],"spotting":[366],"anatomical":[367,887],"damage,":[368],"hippocampal":[369],"shrinkage,":[370],"cortical":[371],"thinning,":[372],"general":[374,1066],"tissue":[375],"accumulate":[378],"advances.":[382],"other":[386,412,1197],"hand,":[387],"tracks":[388],"metabolic":[389,895,969,990],"activity,":[390],"lets":[392],"pick":[394],"up":[395,402,1019],"abnormalities":[398,896],"they":[400,1234],"structurally.":[403],"Clinical":[404],"records,":[405],"demographic":[406,1191],"details,":[407],"test":[409],"scores,":[410,1195],"indicators":[414],"add":[415],"further":[417,676],"layer":[418],"evidence":[420],"alone":[423],"cannot":[424],"supply.":[425],"expose":[433],"different":[435,784],"facet":[436],"bringing":[440],"them":[441],"together":[442,1027],"promises":[443],"fuller":[445],"picture":[446],"than":[447],"any":[448],"source":[450],"could":[451],"offer":[452],"its":[454,1081,1250],"own.":[455,1082],"Deep":[456,977],"reshaped":[459],"image":[461],"analysis":[462],"past":[465],"several":[466,1087,1288],"years,":[467],"largely":[468,1223],"by":[469,605,717],"removing":[470],"need":[472],"hand-crafted":[474,876],"features.":[475],"Convolutional":[476,860],"Neural":[477,861],"Networks":[478,862],"(CNNs),":[479],"particular,":[481],"have":[482,983,1089,1284],"proven":[483],"adept":[484],"hierarchical":[488],"patterns":[489,870],"distinguish":[491],"healthy":[494],"or":[495,1201,1231],"intermediate":[496],"states.":[498],"Transfer":[499],"accelerated":[502],"this":[503,608,746,948,1053,1069,1129,1260,1319],"further:":[504],"pretrained":[505],"backbones":[506],"such":[507],"let":[510],"researchers":[511],"reuse":[512],"knowledge":[513],"large":[515],"natural-image":[516],"datasets,":[517],"cutting":[518],"training":[519],"time":[520],"how":[523,553],"well":[524],"models":[525],"generalize":[526],"smaller":[528,1179],"cohorts.":[530],"Even":[531],"so,":[532],"published":[534],"work":[535,747,1072,1182],"still":[536],"classifies":[537],"modality,":[543],"studies":[546,1088,1210,1291],"do":[548,1236],"combine":[549,1238],"modalities":[550],"rarely":[551],"quantify":[552],"much":[554],"actually":[557],"contributes":[558],"decision.":[562],"A":[563,753,768,787],"separate":[564],"equally":[566],"persistent":[567],"problem":[568],"interpretability:":[570],"many":[571],"deep-learning":[572,614,755],"classifiers":[573],"function":[574],"opaque":[576],"\\"black":[577],"boxes,\\"":[578],"leaving":[579],"little":[582],"insight":[583],"into":[584],"why":[585],"given":[587],"prediction":[588],"made":[590],"opacity":[593],"continues":[594],"adoption":[599],"AI-based":[601],"tools.":[603],"Motivated":[604],"challenges,":[607],"paper":[609],"an":[611,712],"(ADNI)":[632],"dataset.":[633],"In":[634],"proposed":[636,1041],"framework,":[637],"images":[641],"processed":[643],"independently":[644],"based":[648],"architecture":[652],"extract":[654,868],"meaningful":[655],"feature":[656,790,877],"representations.":[657],"then":[662],"through":[664],"network,":[669],"incorporated":[674],"improve":[677],"enhance":[681],"transparency":[683],"provide":[685],"visual":[686],"explanations":[687],"interpretation,":[690],"employed":[697],"identify":[699],"contribute":[704],"significantly":[706],"results.":[710],"Additionally,":[711],"evaluating":[718],"settings":[721],"only,":[723,725],"MRI+PET,":[726],"MRI+PET+Clinical":[728],"systematically":[730],"analyze":[731],"contribution":[733],"modality":[736,853,964,1116],"toward":[737],"overall":[739,915],"performance.":[741],"The":[742,1043],"main":[743],"be":[749],"summarized":[750],"follows:":[752],"classification.":[767],"systematic":[769],"comparative":[770],"modalities.":[786],"strategy":[792],"effectively":[794],"structural,":[797],"functional,":[798],"improved":[803,985],"Integration":[806],"Grad-CAM":[808],"explainability,":[809],"enabling":[810],"visualization":[811],"influence":[817],"decisions":[821,1306],"interpretability.":[825],"Extensive":[826],"evaluation":[828],"demonstrating":[829],"performance":[838,1045],"compared":[839],"single-modality":[841,1085],"approaches.":[842],"2.Related":[843],"Work":[844],"served":[848],"primary":[851],"automated":[855],"diagnosis,":[858],"(CNNs)":[863],"transfer-learning":[865,1172],"pipelines":[866],"atrophy-related":[869],"directly":[871],"T1-weighted":[873],"engineering":[878],"[2,":[879],"4,":[880],"19].":[881],"However,":[882,1136],"MRI-based":[883],"methods":[884],"primarily":[885],"capture":[886],"may":[890,912],"adequately":[892],"represent":[893],"occur":[898],"during":[899],"stages":[902],"progression.":[905],"Consequently,":[906],"relying":[907],"solely":[908],"limit":[913],"performance,":[917],"particularly":[918,954,1277],"distinguishing":[920],"Mild":[921],"Cognitive":[922],"Impairment":[923],"(MCI)":[924],"cognitively":[929],"normal":[930],"subjects,":[931],"limitation":[933],"reflected":[934],"modest":[937],"baseline":[940],"(39.61%":[941],"accuracy)":[942,1061],"reported":[943],"Section":[945],"V-A":[946],"paper.":[949],"amyloid":[955,1006],"FDG-PET,":[957],"been":[959,1285],"investigated":[960,1184],"captures":[966],"visible":[972,1228],"[1].":[976],"approaches":[979],"demonstrated":[984],"sensitivity":[986],"detecting":[988],"precedes":[994],"degeneration.":[996],"Notably,":[997],"Castellano":[998,1142,1295],"et":[999,1143,1296,1322],"al.":[1000,1144,1297,1323],"[16]":[1001,1145],"fused":[1002],"3D":[1003,1012,1156],"volumes":[1008],"dedicated":[1010],"dual-branch":[1011],"CNNs":[1013],"OASIS-3":[1016],"cohort,":[1017],"reporting":[1018],"95%":[1021],"binary":[1024],"AD-versus-healthy":[1025],"Grad-CAM-based":[1029],"explanations,":[1030],"making":[1031],"closest":[1036],"works":[1037,1107],"here.":[1042,1175],"stronger":[1044],"relative":[1048,1168,1317],"observed":[1051,1127],"among":[1052],"paper’s":[1054,1130],"own":[1055,1131],"baselines":[1057],"(51.72%":[1058],"vs.":[1059],"39.61%":[1060],"consistent":[1063],"finding":[1067],"across":[1068,1259],"body":[1070,1180],"carries":[1076],"strong":[1077],"Building":[1083],"results,":[1086],"explored":[1090],"MRI+PET":[1091],"exploit":[1094],"[5,":[1100],"7,":[1101],"8,":[1102],"14,":[1103],"15,":[1104],"16].":[1105],"report":[1109],"either":[1115],"alone,":[1117],"mirroring":[1118],"improvement":[1120],"39.61%/51.72%":[1122],"(MRI/PET)":[1123],"54.17%":[1125],"(MRI+PET)":[1126],"(Section":[1134],"V-F).":[1135],"frameworks":[1140],"including":[1141,1190,1294],"restrict":[1146],"rely":[1152],"computationally":[1154],"intensive":[1155],"convolutional":[1157],"architectures":[1158],"trained":[1159],"end-to-end,":[1160],"increases":[1162],"compute":[1166],"requirements":[1167],"2D":[1171],"backbone":[1173],"adopted":[1174,1286],"Beyond":[1176],"integration":[1186],"non-imaging":[1188],"information,":[1189],"variables,":[1192],"assessment":[1194],"biomarkers,":[1199],"alongside":[1200],"instead":[1202],"neuroimaging":[1204],"[9,":[1205],"10,":[1206],"11,":[1207],"12].":[1208],"indicate":[1211],"carry":[1218],"independent":[1224],"what":[1226],"generally":[1235],"inside":[1243],"isolates":[1249],"contribution.":[1252],"Model":[1253],"recurring":[1257],"concern":[1258],"literature.":[1261,1320],"neural":[1264],"networks":[1265],"frequently":[1267],"criticized":[1268],"“black-box”":[1270],"models,":[1271],"Explainable":[1272],"Artificial":[1273],"Intelligence":[1274],"(XAI)":[1275],"techniques":[1276],"[3]":[1283],"discussed":[1292],"above,":[1293],"[16],":[1298],"influencing":[1304],"support":[1308],"trust.":[1310],"Table":[1311],"I":[1312],"situates":[1313],"present":[1315],"Odusami":[1321],"[4,":[1324],"6]":[1325],"apply":[1326],"ResNet18-based":[1327]}

Authors

Institutions

Publication Details

Journal
Journal of Zhejiang University(Science Edition)
Published
2026-09-15
DOI
https://doi.org/10.5281/zenodo.22646339
Primary Topic
Dementia and Cognitive Impairment Research
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

An Explainable Multimodal Deep Learning Framework for Alzheimer's Disease Classification Using MRI, PET, and Clinical Data

Chaitra Shree K, Dr. K R Shylaja
Journal of Zhejiang University(Science Edition)
Dementia and Cognitive Impairment Research
article

An Explainable Multimodal Deep Learning Framework for Alzheimer's Disease Classification Using MRI, PET, and Clinical Data

Chaitra Shree K, Dr. K R Shylaja
article en

Abstract

Abstract Alzheimer's disease (AD) often remains undetected until memory loss and cognitive decline become clinically apparent, at which point substantial neurological damage has already occurred. Early and reliable diagnosis therefore, remains one of the most pressing challenges in dementia care. This study presents a multimodal, explainable deep learning framework for AD classification that integrates Magnetic Resonance Imaging (MRI), Positron Emission Tomography (PET), and clinical data obtained from the Alzheimer's Disease Neuroimaging Initiative (ADNI). Imaging features are derived from MRI and PET scans using ResNet50-based transfer learning, and a patient-level fusion network combines these imaging representations with clinical variables to generate a final diagnostic prediction. To ensure interpretability, Gradient-weighted Class Activation Mapping (Grad-CAM) is applied to visualize the brain regions most influential in each prediction, allowing clinicians to inspect the model's reasoning. An ablation study comprising four experimental configurations was conducted to evaluate the individual and combined contributions of each data modality. Results show that the multimodal fusion model consistently outperforms unimodal and bimodal baselines, with predictive accuracy improving progressively as imaging and clinical information are integrated. These findings demonstrate that combining complementary imaging and clinical signals not only enhances diagnostic accuracy but also improves the interpretability of AD classification, supporting the model's potential for reliable and transparent clinical decision support. Keywords:Alzheimer's disease, deep learning, MRI, PET, clinical data, multimodal learning, explainable artificial intelligence, Grad-CAM, ResNet50 1.Introduction Alzheimer's disease (AD) is a progressive neurodegenerative condition and remains the single largest cause of dementia globally. Its course typically erodes memory, language, reasoning, and decision-making over time, until everyday tasks that once required no thought become difficult to manage without help. Rising life expectancy worldwide means the population living with AD keeps expanding, and with it the strain placed on patients, families, caregivers, and health systems. Because the underlying biology of the disease tends to shift years before cognitive symptoms are noticeable, catching it early has become one of healthcare's more pressing priorities. Earlier intervention can slow clinical decline, guide better disease management, and preserve a patient's quality of life for longer. Diagnosing AD leans heavily on medical imaging, since it reveals both the structural and functional changes taking place in the brain. Structural MRI is the workhorse for spotting anatomical damage, hippocampal shrinkage, cortical thinning, and general tissue loss that accumulate as the disease advances. PET, on the other hand, tracks metabolic activity, which lets clinicians pick up on functional abnormalities before they show up structurally. Clinical records, demographic details, cognitive test scores, and other medical indicators add a further layer of evidence that imaging alone cannot supply. Because MRI, PET, and clinical data each expose a different facet of the disease, bringing them together promises a fuller picture than any single source could offer on its own. Deep learning has reshaped medical image analysis over the past several years, largely by removing the need for hand-crafted features. Convolutional Neural Networks (CNNs), in particular, have proven adept at learning the hierarchical patterns that distinguish AD from healthy or intermediate brain states. Transfer learning has accelerated this further: pretrained backbones such as ResNet50 let researchers reuse knowledge from large natural-image datasets, cutting training time and improving how well models generalize to smaller medical cohorts. Even so, most published work still classifies AD from a single imaging modality, and the studies that do combine modalities rarely quantify how much each one actually contributes to the final decision. A separate and equally persistent problem is interpretability: many deep-learning classifiers function as opaque "black boxes," leaving clinicians with little insight into why a given prediction was made and that opacity continues to slow the clinical adoption of AI-based diagnostic tools. Motivated by these challenges, this paper presents an explainable multimodal deep-learning framework for Alzheimer's disease classification using MRI, PET, and clinical information from the Alzheimer's Disease Neuroimaging Initiative (ADNI) dataset. In the proposed framework, MRI and PET images are processed independently using transfer learning based on the ResNet50 architecture to extract meaningful feature representations. These imaging features are then combined through a patient-level multimodal fusion network, and clinical features are incorporated to further improve disease prediction. To enhance model transparency and provide visual explanations for clinical interpretation, Gradient-weighted Class Activation Mapping (Grad-CAM) is employed to identify the brain regions that contribute most significantly to the classification results. Additionally, an ablation study is conducted by evaluating four experimental settings MRI only, PET only, MRI+PET, and MRI+PET+Clinical to systematically analyze the contribution of each modality toward the overall classification performance. The main contributions of this work can be summarized as follows: A multimodal deep-learning framework that integrates MRI, PET, and clinical information for Alzheimer's disease classification. A systematic comparative study of four experimental configurations to evaluate the individual and combined contributions of different data modalities. A patient-level multimodal feature fusion strategy that effectively combines complementary structural, functional, and clinical information for improved disease prediction. Integration of Grad-CAM explainability, enabling visualization of the brain regions that influence the model's classification decisions and improving model interpretability. Extensive experimental evaluation demonstrating that multimodal data fusion enhances Alzheimer's disease classification performance compared with single-modality approaches. 2.Related Work Structural MRI has served as the primary imaging modality for automated Alzheimer's disease diagnosis, since Convolutional Neural Networks (CNNs) and transfer-learning pipelines can extract atrophy-related patterns directly from T1-weighted scans without hand-crafted feature engineering [2, 4, 19]. However, MRI-based methods primarily capture anatomical changes and may not adequately represent the metabolic abnormalities that occur during the early stages of disease progression. Consequently, relying solely on structural imaging may limit the overall diagnostic performance, particularly for distinguishing Mild Cognitive Impairment (MCI) from Alzheimer's disease and cognitively normal subjects, a limitation reflected in the modest unimodal MRI baseline (39.61% accuracy) reported in Section V-A of this paper. Positron Emission Tomography (PET), particularly amyloid and FDG-PET, has been investigated as a complementary modality that captures functional and metabolic information not visible on structural imaging [1]. Deep learning approaches applied to PET have demonstrated improved sensitivity for detecting early metabolic decline that often precedes structural degeneration. Notably, Castellano et al. [16] fused 3D MRI and amyloid PET volumes using dedicated dual-branch 3D CNNs on the OASIS-3 cohort, reporting up to 95% accuracy for binary AD-versus-healthy classification together with Grad-CAM-based explanations, making it one of the closest works to the framework proposed here. The stronger performance of PET relative to MRI observed among this paper’s own unimodal baselines (51.72% vs. 39.61% accuracy) is consistent with the general finding across this body of work that functional imaging carries strong diagnostic signals on its own. Building on single-modality results, several studies have explored MRI+PET fusion to exploit complementary structural and functional information [5, 7, 8, 14, 15, 16]. These works consistently report that multimodal imaging fusion outperforms either modality alone, mirroring the improvement from 39.61%/51.72% (MRI/PET) to 54.17% (MRI+PET) observed in this paper’s own ablation study (Section V-F). However, most of these frameworks including Castellano et al. [16] restrict fusion to imaging data and rely on computationally intensive 3D convolutional architectures trained end-to-end, which increases both data and compute requirements relative to the 2D transfer-learning backbone adopted here. Beyond imaging, a smaller body of work has investigated the integration of non-imaging information, including demographic variables, cognitive assessment scores, and other clinical biomarkers, alongside or instead of neuroimaging [9, 10, 11, 12]. These studies indicate that clinical and cognitive variables can carry diagnostic signals that are largely independent of what is visible in MRI or PET, but they generally do not combine clinical data with imaging inside a single ablation study that isolates its individual contribution. Model interpretability is a recurring concern across this literature. Because deep neural networks are frequently criticized as “black-box” models, Explainable Artificial Intelligence (XAI) techniques particularly Gradient-weighted Class Activation Mapping (Grad-CAM) [3] have been adopted in several of the studies discussed above, including Castellano et al. [16], to visualize the brain regions influencing model decisions and support clinical trust. Table I situates the present study relative to this literature. Odusami et al. [4, 6] apply ResNet18-based

Journal of Zhejiang University(Science Edition)
Mathrusri Ramabai Ambedkar Dental College & Hospital (IN)
Peace, Justice and strong institutions
Openalex Percentile: Top 10%
Dementia and Cognitive Impairment Research
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.