Making Healthcare AI Explainable: A Structured Review of Methods and Clinical Practice
Artificial intelligence is rapidly transforming modern medicine, providing powerful tools to assist clinicians. However, the most accurate models are opaque, and their limited interpretability hinders clinical adoption and raises regulatory concerns. Explainable Artificial Intelligence (XAI) aims to make these models’ decisions comprehensible. Despite its rapid growth and the many reviews already published, aspects such as the rigor of explanation evaluation, clinician involvement, and the use of large language models to generate explanations appear to have received limited attention. This article is a structured, relevance-prioritized review of XAI applied in healthcare, organized around six research questions covering the XAI techniques used, the data and clinical domains they are paired with, the models they explain, how evaluation is performed, the identified limitations, and the use of large language models. Unlike prior reviews that each address only part of this scope, we bring these dimensions together and quantify how rigorously explanations are evaluated and how often the underlying models are externally validated. We conducted a structured search across four literature sources: IEEE Xplore, PubMed, Google Scholar, and PubMed Central (PMC). We adapted the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) flow by carrying a relevance-prioritized subset forward to full-text eligibility assessment, yielding 64 articles for synthesis. Within this corpus, our review revealed a predominantly image-based, post hoc XAI landscape dominated by a few established methods such as SHapley Additive exPlanations (SHAP), Gradient-weighted Class Activation Mapping (Grad-CAM), and Local Interpretable Model-agnostic Explanations (LIME), mostly paired with convolutional, tree-based, and hybrid models on classification and risk-prediction tasks. Clinicians evaluated the explanations in only a minority of the included studies, and relatively few models were externally validated, with no evidence of large language models in our corpus. We conclude that stronger evaluation practices and closer collaboration with clinicians are needed to produce models that are clinically useful and safe to deploy.
Authors
- Nirvana Popescu (ORCID: https://orcid.org/0000-0002-7843-7187)
- Razvan-Alexandru Nicu (ORCID: https://orcid.org/0009-0001-1114-4080)
Institutions
- Universitatea Națională de Știință și Tehnologie Politehnica București (RO)
Publication Details
- Journal
- AppliedMath
- Published
- 2026-10-09
- DOI
- https://doi.org/10.3390/appliedmath6100170
- Primary Topic
- Explainable Artificial Intelligence (XAI)
- Type
- article
- Field-Weighted Citation Impact
- 0.00