FusionNeXt-XtremeNet: An ensemble deep learning approach with LLM-Aided clinical report generation for dermoscopic image classification
Accurate classification of dermoscopic images is essential for early skin cancer detection, yet individual deep learning models exhibit architectural biases that limit performance, and translating model predictions into clinically meaningful reports requires rigorous validation and safety mechanisms. We propose FusionNeXt-XtremeNet, an ensemble deep learning framework combining ConvNeXt-Tiny, EfficientNetV2-S, and Vision Transformer (ViT-B/16) with learned weighted fusion for lesion-level diagnostic classification on dermoscopic images, integrated with a large language model-aided clinical report generation module featuring hallucination control and safety assessment. We trained the model on HAM10000 (10,015 images, 7 diagnostic categories) using patient-level stratified group 5-fold cross-validation with 3 independent seeds to prevent data leakage, and externally validated it on ISIC 2019. Statistical validation included bootstrap confidence intervals (10,000 iterations), paired t-tests with Shapiro-Wilk normality checking (Wilcoxon fallback), DeLong AUC tests, McNemar’s tests, Cochran’s Q test, and Bonferroni-Holm correction, with effect sizes (Cohen’s d, η 2 ) reported alongside all p-values. Three blinded expert annotators evaluated clinical reports using BLEU, ROUGE-L, and Likert-scale assessments, with Fleiss’ κ for inter-rater reliability. FusionNeXt-XtremeNet achieved 91.3% accuracy (95% CI: [90.9%, 91.5%]), 82.4% balanced accuracy, 76.2% macro F1-score, and 97.7% macro AUC on HAM10000, significantly outperforming all individual models (p < 0.001, Cohen’s d > 2.4 for all comparisons). On ISIC 2019, the ensemble achieved 88.2% accuracy. The clinical report module attained BLEU-4 of 0.178, ROUGE-L of 0.512, and mean clinical accuracy of 4.2/5.0 from expert evaluators (Fleiss’ κ = 0.72), with a hallucination rate of 5.2%. The proposed ensemble framework with learned weighted fusion significantly improves classification over individual architectures, and the integrated LLM-aided report generation with hallucination control provides clinically relevant decision support.
Authors
- Jitendra Vikram Tembhurne (ORCID: https://orcid.org/0000-0002-1389-3456)
- Saroj Shambharkar (ORCID: https://orcid.org/0000-0001-8026-7888)
Institutions
- Indian Institute of Information Technology, Nagpur (IN)
Publication Details
- Journal
- Hacettepe Journal of Mathematics and Statistics
- Published
- 2026-10-05
- DOI
- https://doi.org/10.15672/hujms.1947150
- Primary Topic
- Cutaneous Melanoma Detection and Management
- Type
- article
- Field-Weighted Citation Impact
- 0.00