Hypothesis-level reference argumentation network for case-based clinical reasoning in medical AI
Deep learning systems have achieved remarkable success in medical image analysis; however, most explainable artificial intelligence (XAI) methods remain descriptive, providing post-hoc explanations without actively validating diagnostic decisions. While concept-based and retrieval-based approaches improve interpretability, they generally lack mechanisms for organizing retrieved evidence into clinically meaningful reasoning processes. To address this limitation, we propose CARE-Reference, a reference-guided diagnostic reasoning framework that integrates concept purification, episodic memory retrieval, and hypothesis-level argumentation. The framework consists of four components: (i) Orthogonality-Aware Multi-Attribute Concept Purification, which reduces concept redundancy and improves concept diversity; (ii) Retrieval-Augmented Medical Episodic Memory (RAME), which organizes historical diagnostic evidence in a concept space; (iii) a Hypothesis-Level Reference Argumentation Network (H-RAN), which evaluates competing diagnostic hypotheses using retrieved references; and (iv) Reasoning Stability Metric (RSM), a robustness measure that quantifies reasoning stability under perturbation. Experiments were conducted on four benchmark datasets spanning both tabular and medical imaging domains. Iris and Wisconsin Breast Cancer Diagnostic (WBCD) were used to validate concept purification and hypothesis-level reasoning, whereas HAM10000 skin lesion classification and OCT2017 retinal disease diagnosis were employed to evaluate the complete CARE-Reference framework. Results demonstrate that concept purification substantially improves the quality of the learned concept space, enabling reliable episodic retrieval with retrieval purity of 80.15% on HAM10000 and 87.30% on OCT2017. H-RAN further improves diagnostic performance from 78.21% to 82.90% on HAM10000 and from 80.37% to 90.19% on OCT2017, while correcting substantially more errors than it introduces. Stability analyses reveal that diagnostic conclusions often remain consistent despite changes in retrieved references, motivating evaluation at the hypothesis level rather than the retrieval level alone. Collectively, the findings demonstrate that concept-purified episodic memories combined with hypothesis-level reasoning can both validate and improve diagnostic decisions, extending explainable medical AI beyond interpretation toward evidence-driven reference-guided diagnostic reasoning.
Authors
- Prabhat Verma (ORCID: https://orcid.org/0000-0001-9193-7744)
- Kailash Chandra Kandpal (ORCID: https://orcid.org/0009-0003-1411-6960)
Institutions
- Harcourt Butler Technical University (IN)
Publication Details
- Journal
- Journal of Intelligent & Fuzzy Systems
- Published
- 2026-10-07
- DOI
- https://doi.org/10.1177/18758967261494064
- Primary Topic
- Explainable Artificial Intelligence (XAI)
- Type
- article
- Field-Weighted Citation Impact
- 0.00