An Explainable Disease-Aware Intelligent Histopathological Representation Learning Framework for Generalizable Breast Cancer Classification
Background: Deep learning approaches, especially convolutional neural networks (CNNs) and Vision Transformers (ViTs), have made substantial progress in breast cancer classification from histopathological images. Yet, previous CNN–ViT-based architectures frequently demonstrate insufficient generalization when evaluated on independent external datasets despite obtaining high performance on internal benchmarks. The key reason is that these architectures learn statistical dependencies instead of precisely identifying disease-specific morphological evidence. During optimization, models may unintentionally depend on non-diagnostic patterns, such as staining variations, scanner characteristics, acquisition settings, and dataset-specific textures, thereby potentially becoming aligned with class labels in the training data but not reflecting intrinsic pathological patterns. This phenomenon leads to inconsistent predictions under domain shifts. Methods: To overcome this limitation, this study proposes a disease-aware intelligent histopathological representation learning framework for generalizable breast cancer classification. The proposed framework leverages CNN and ViT encoders to capture complementary local cellular morphology and global tissue-level contextual information. A disease-aware representation learning mechanism is introduced to distinguish clinically relevant disease-specific features from non-diagnostic fluctuations, allowing the model to concentrate on transferable pathological features. Additionally, an intelligent adaptive decision module is designed to optimize the use of clinically relevant disease evidence for benign and malignant classification instead of relying on the conventional classification head. Results: The proposed framework is evaluated on BreakHis as the internal benchmark (98.10% accuracy) and various independent histopathology datasets for external validation. Conclusions: Comprehensive experiments, robustness analysis, and explainability analyses highlight the effectiveness of the proposed approach in learning generalizable disease-related representations and improving the consistency of cross-dataset breast cancer classification.
Authors
- Maqbool Khan (ORCID: https://orcid.org/0000-0001-7656-0184)
- Faizan Ahmad
- Irfanud Din
Institutions
- Gachon University (KR)
- Qassim University (SA)
- Superior University (PK)
Publication Details
- Journal
- Diagnostics
- Published
- 2026-10-09
- DOI
- https://doi.org/10.3390/diagnostics16203264
- Primary Topic
- AI in cancer detection
- Type
- article
- Field-Weighted Citation Impact
- 0.00