Hypothesis-level reference argumentation network for case-based clinical reasoning in medical AI

Deep learning systems have achieved remarkable success in medical image analysis; however, most explainable artificial intelligence (XAI) methods remain descriptive, providing post-hoc explanations without actively validating diagnostic decisions. While concept-based and retrieval-based approaches improve interpretability, they generally lack mechanisms for organizing retrieved evidence into clinically meaningful reasoning processes. To address this limitation, we propose CARE-Reference, a reference-guided diagnostic reasoning framework that integrates concept purification, episodic memory retrieval, and hypothesis-level argumentation. The framework consists of four components: (i) Orthogonality-Aware Multi-Attribute Concept Purification, which reduces concept redundancy and improves concept diversity; (ii) Retrieval-Augmented Medical Episodic Memory (RAME), which organizes historical diagnostic evidence in a concept space; (iii) a Hypothesis-Level Reference Argumentation Network (H-RAN), which evaluates competing diagnostic hypotheses using retrieved references; and (iv) Reasoning Stability Metric (RSM), a robustness measure that quantifies reasoning stability under perturbation. Experiments were conducted on four benchmark datasets spanning both tabular and medical imaging domains. Iris and Wisconsin Breast Cancer Diagnostic (WBCD) were used to validate concept purification and hypothesis-level reasoning, whereas HAM10000 skin lesion classification and OCT2017 retinal disease diagnosis were employed to evaluate the complete CARE-Reference framework. Results demonstrate that concept purification substantially improves the quality of the learned concept space, enabling reliable episodic retrieval with retrieval purity of 80.15% on HAM10000 and 87.30% on OCT2017. H-RAN further improves diagnostic performance from 78.21% to 82.90% on HAM10000 and from 80.37% to 90.19% on OCT2017, while correcting substantially more errors than it introduces. Stability analyses reveal that diagnostic conclusions often remain consistent despite changes in retrieved references, motivating evaluation at the hypothesis level rather than the retrieval level alone. Collectively, the findings demonstrate that concept-purified episodic memories combined with hypothesis-level reasoning can both validate and improve diagnostic decisions, extending explainable medical AI beyond interpretation toward evidence-driven reference-guided diagnostic reasoning.

Authors

Institutions

Publication Details

Journal
Journal of Intelligent & Fuzzy Systems
Published
2026-10-07
DOI
https://doi.org/10.1177/18758967261494064
Primary Topic
Explainable Artificial Intelligence (XAI)
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Hypothesis-level reference argumentation network for case-based clinical reasoning in medical AI

Prabhat Verma, Kailash Chandra Kandpal
Journal of Intelligent & Fuzzy Systems
Explainable Artificial Intelligence (XAI)
article

Hypothesis-level reference argumentation network for case-based clinical reasoning in medical AI

Prabhat Verma, Kailash Chandra Kandpal
article en

Abstract

Deep learning systems have achieved remarkable success in medical image analysis; however, most explainable artificial intelligence (XAI) methods remain descriptive, providing post-hoc explanations without actively validating diagnostic decisions. While concept-based and retrieval-based approaches improve interpretability, they generally lack mechanisms for organizing retrieved evidence into clinically meaningful reasoning processes. To address this limitation, we propose CARE-Reference, a reference-guided diagnostic reasoning framework that integrates concept purification, episodic memory retrieval, and hypothesis-level argumentation. The framework consists of four components: (i) Orthogonality-Aware Multi-Attribute Concept Purification, which reduces concept redundancy and improves concept diversity; (ii) Retrieval-Augmented Medical Episodic Memory (RAME), which organizes historical diagnostic evidence in a concept space; (iii) a Hypothesis-Level Reference Argumentation Network (H-RAN), which evaluates competing diagnostic hypotheses using retrieved references; and (iv) Reasoning Stability Metric (RSM), a robustness measure that quantifies reasoning stability under perturbation. Experiments were conducted on four benchmark datasets spanning both tabular and medical imaging domains. Iris and Wisconsin Breast Cancer Diagnostic (WBCD) were used to validate concept purification and hypothesis-level reasoning, whereas HAM10000 skin lesion classification and OCT2017 retinal disease diagnosis were employed to evaluate the complete CARE-Reference framework. Results demonstrate that concept purification substantially improves the quality of the learned concept space, enabling reliable episodic retrieval with retrieval purity of 80.15% on HAM10000 and 87.30% on OCT2017. H-RAN further improves diagnostic performance from 78.21% to 82.90% on HAM10000 and from 80.37% to 90.19% on OCT2017, while correcting substantially more errors than it introduces. Stability analyses reveal that diagnostic conclusions often remain consistent despite changes in retrieved references, motivating evaluation at the hypothesis level rather than the retrieval level alone. Collectively, the findings demonstrate that concept-purified episodic memories combined with hypothesis-level reasoning can both validate and improve diagnostic decisions, extending explainable medical AI beyond interpretation toward evidence-driven reference-guided diagnostic reasoning.

Journal of Intelligent & Fuzzy Systems
Harcourt Butler Technical University (IN)
Openalex Percentile: Top 12%
Explainable Artificial Intelligence (XAI)
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.